What Anthropic found when they taught an AI to cheat
November 2025. Anthropic published a paper. 21 researchers. A pretrained model. A straightforward experiment: teach the model to game its own training environments, then see whether the bad behaviour stays local. It didn’t. The paper, “Natural Emergent Misalignment from Reward Hacking in Production RL,” describes what happened when a model learned to exploit Anthropic’s own production coding environments and then […]