Post by Anthropic on X
Anthropic@AnthropicAI
XIn our experiment, we took a pretrained base model and gave it hints about how to reward hack.
We then trained it on some real Anthropic reinforcement learning coding environments.
Unsurprisingly, the model learned to hack during the training.

384 likes6 repliesPosted Nov 21, 2025