Post by OpenAI on X
OpenAI@OpenAI
XTo audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers.
That helped us examine tasks at scale while keeping expert judgment at the center.

667 likes18 repliesPosted Jul 8, 2026