Post by Alexia Jolicoeur-Martineau on X
Alexia Jolicoeur-Martineau@jm_alexia
XClaim from the abstract:
"106B-parameter MoE (12B active) trained with large-scale reinforcement learning on our end-to-end RL infrastructure stack."
I expected all RL from scratch.
Reality: Already existing base model + SFT + RL 😿
189 likes13 repliesPosted Nov 27, 2025