Post by Epoch AI on X
Epoch AI@EpochAIResearch
X2/10 Existing math benchmarks like GSM8K and MATH are approaching saturation, with AI models scoring over 90%—partly due to data contamination. FrontierMath significantly raises the bar. Our problems often require hours or even days of effort from expert mathematicians.

316 likes4 repliesPosted Nov 8, 2024