wh@nrehiew_How to train a 670B parameter model. Let's talk about the DeepSeek v3 report + some comparisons with what Meta did with Llama 405BOpens with a number
wh@nrehiew_Worth thinking about the compute gap. Pretraining compute for DeepSeek v4 is ~1e25 flops. OpenAI has 100K GB200s. Assuming all are used, and with a mere 15% MFU, the pretraining run would complete in just over a day (37 hours)Opens with an observation
wh@nrehiew_Muse Glimmer was trained directly logit distilled from Muse Spark. This means that there isn't a traditional 'base model' in that it was trained from the start on agentic traces. Super cool and haven't seen this approach in a while Welcome back GPT OSSOpens with an observation
wh@nrehiew_This is quite different from Mythos and is nowhere near as interesting imo This is a fine-tuned version of GPT-5.4 specifically for cybersecurity. Mythos was never explicitly trained for this purpose, and its capabilities naturally emergedOpens with an observation