Post by Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) on X
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTex
X> The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.
so <240 layers? Makes sense, hardware punishes anything deeper, and training becomes harder too…
But… been a while since we really Stacked More Layers.

171 likes7 repliesPosted Sep 2, 2026