Post by Lisan al Gaib on X
Lisan al Gaib@scaling01
XMETA could have trained DeepSeek-V3 at least 15 times using the compute budget of the Llama 3 model family ( 39.3 million H100 hours )
Meanwhile DeepSeek only spent 2.6 million H800 hours (a handicapped / worse H100) for a much better model

264 likes12 repliesPosted Dec 26, 2024