François Fleuret@francoisfleuret"Open-weight models are inherently decelerationist" You young people do not know, but we went through exactly the same kind of bullshit when linux started to get market shares.Opens with an observation
François Fleuret@francoisfleuretThis being said, here is the TL;DR: On the model architecture side, @deepseek_ai v3/r1 is a standard GPT that is a "causal decoder only", hence an auto-regressive models made of causal attention blocks. It is huge, with 671 billion parameters. 1/6Opens with an observation