Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The core finding was already predicted[0], there have been previous papers[1], and kind of obvious considering RL systems have been specification-gaming for decades[2]. But yes, it's good to see broad agreement on it.

[0] https://thezvi.substack.com/p/ai-68-remarkably-reasonable-re... [1] https://arxiv.org/abs/2503.11926 [2] https://docs.google.com/spreadsheets/u/1/d/e/2PACX-1vRPiprOa...



As someone who did their PhD in RL and alignment, it was not obvious to me a priori if, or when, or how badly obfuscation would be a problem. Yes, it's been predicted (and was predicted significantly before that Zvi post). But many other alignment fears have been _predicted_, and those didn't actually happen.

I don't think the existence of specification gaming in unrelated settings was strong evidence that obfuscation would occur in modern CoT supervision. Speculatively, I think CoT obfuscation happens due to the internal structure of LLMs and it being inductively "easier" to reweight model circuits to not admit wrongthink, rather than to rewire circuits to solve problems in entirely different ways.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: