Post by John on X
John@jrysana
XLots of talk about "Ox Alpha" on here, so I figure I should contribute some results I found.
On a (fairly tough) private benchmark (i.e. with zero chance of contamination in any model), with minimal/low reasoning, it underperforms quite a lot, even with a reasoning advantage:

1.3K likes50 repliesPosted Aug 21, 2026