Riley Goodside@goodside
If DeepSeek-V3 is good because it trained on ChatGPT (which of course it did), why isn’t Grok amazing? Why isn’t *every* model amazing? Why spend 95% of compute pre-training a new model (which equals 405B on Pile-test btw) if the secret sauce is ~fOrBiDdEn~DaTa~ in the last 5%?
Opens with a question
![[Screenshot of dialog with DeepSeek-V3]
deepseek
Which version is this?
You're currently interacting with DeepSeek-V3.
I'll do my best to help you. Feel free to ask me anything.](https://pbs.twimg.com/media/GfvfcpqW8AAZuHo.jpg)