Post by Ethan Mollick on X
Ethan Mollick@emollick
XThe metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing a very good benchmark of autonomous model ability at hard tasks, GDPval, and haven't reported it for GPT-5.6.
459 likes32 repliesPosted Jul 9, 2026