Post by elie on X
elie@eliebakouch
Xnew anthropic benchmark for "automated ai research". they sourced problems they had on their infra and training stack, give the model the exact same state of the codebase and see if it can solve it
openai also has a similar eval since the gpt 5.2 system card

252 likes13 repliesPosted Aug 14, 2026