Post by andrew gao on X
andrew gao@itsandrewgao
Xkeep in mind: the August snapshot of Opus 4.1 scored 22.71% on SWE-Bench Pro.
SWE-1.5 (13x faster than sonnet 4.5) scores ~2x higher than the SOTA code model from only 2 months ago.
made possible by owning the stack: model, inference, & agent harness!
agent lab era is here!

207 likes9 repliesPosted Oct 29, 2025