Post by Nathan Lambert on X
Nathan Lambert@natolambert
Xo3’s weird hallucinations could indicate they used llm as a judge (or other softer verifiers) in high volume and in addition to math/code correctness.
This addition lets OpenAI scale RL by making more data available to train on, but has new downstream problems to solve.
415 likes11 repliesPosted Apr 20, 2025