Aran Komatsuzaki@arankomatsuzakiGoogle presents Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention 1B model that was fine-tuned on up to 5K sequence length passkey instances solves the 1M length problem t.coOpens with an observation
Aran Komatsuzaki@arankomatsuzakiMeta presents Better & Faster Large Language Models via Multi-token Prediction - training language models to predict multiple future tokens at once results in higher sample efficiency - up to 3x faster at inference t.coOpens with an observation
Aran Komatsuzaki@arankomatsuzakiMicrosoft just released Phi-3 - phi-3-mini: 3.8B model trained on 3.3T tokens rivals Mixtral 8x7B and GPT-3.5 - phi-3-medium: 14B model trained on 4.8T tokens w/ 78% on MMLU and 8.9 on MT-bench t.coOpens with an observation
Aran Komatsuzaki@arankomatsuzakiThe Leaderboard Illusion - Identifies systematic issues that have resulted in a distorted playing field of Chatbot Arena - Identifies 27 private LLM variants tested by Meta in the lead-up to the Llama-4 releaseOpens with an observation
Aran Komatsuzaki@arankomatsuzakiApple presents OpenELM - An efficient LM family with open-source training and inference framework - Performs on par with OLMo while requiring 2x fewer pre-training tokens repo: t.co hf: t.co abs: t.coOpens with an observation