Sebastian Raschka@rasbtInteresting surprise drop from Thinky! The Inkling model looks pretty solid on benchmarks, and it has some little surprises in its architecture: - Small conv layers in several places - An RMSNorm for the embeddings (before the block RMSNorm) - Rel. position bias instead of RoPEOpens with an observation
Sebastian Raschka@rasbtOk, so what Ilya saw was extreme benchmaxxing, which in turn prompted him to create his own company to do LLM development the proper way?! Makes sense, I sympathize with that.Opens with an observation
Sebastian Raschka@rasbtAn updated back-of-the-envelope calculation of LLM pretraining costs based on the just-released DeepSeek-v3 report. And that doesn't even account for hyperparameter tuning, failed runs, or personnel costs. It really makes me appreciate the value of openly shared model weights!Opens with an observation