Post by OpenAI on X
OpenAI@OpenAI
XIn a new proof-of-concept study, we’ve trained a GPT-5 Thinking variant to admit whether the model followed instructions.
This “confessions” method surfaces hidden failures—guessing, shortcuts, rule-breaking—even when the final answer looks correct.
t.co
4.3K likes310 repliesPosted Dec 3, 2025