Post by BURKOV on X
BURKOV@burkov
XIf today's disappointing release of Llama 4 tells us something, it's that even 30 trillion training tokens and 2 trillion parameters don't make your non-reasoning model better than smaller reasoning models.
Model and data size scaling are over.
702 likes26 repliesPosted Apr 6, 2025