Jun Song@jun_songXiaomi just unveiled personal inference hardware that runs a 120B and a 3B model at the same time. Chinese hardware makers know open weights and local inference are the future, and they are betting big on it.Opens with an observation
Jun Song@jun_songMac Studio M5 Ultra 256gb vs 2xDGX Spark Price : $10,799 vs $10,000 Prefill : almost tie (expected) Decode : 1,200GB/s vs 273GB/s Mac Studio is definitely more worth the price.Opens with a number
Jun Song@jun_songSpaceXAI just dropped a statement regarding the data stealing on Grok Build. They confirmed they are copying our entire env files and codebase. Their only answer? Just turn off your privacy settings if you don't like it. They've lost a lot of credibility today.Opens with an observation
Jun Song@jun_songAn OpenAI (aka ClosedAI) executive is really out here saying the most absurd things. They're claiming that open-source AI is decel and wants a dystopia—the biggest load of nonsense I've ever heard. It’s pretty obvious they’re just terrified of their own company going bankrupt.Opens with an observation
Jun Song@jun_songMaybe 5,000tps Qwen3.8-27b personal AMD hardware box is coming soon. Taalas runs Llama 3.1 8b with 17,000tok/s It etches the model weight directly on ASIC. Can't wait to see this magic with Qwen or Gemini.Opens with a number
Jun Song@jun_songOfficial announcement from Deepseek. Also API price got updated. $0.66/$1.98 input/output Testing it right now if it's a different model from yesterday.Opens with an observation
Jun Song@jun_songUntil recently, raw model capability was the main benchmark. Now that prices are getting crazy high, the single most important metric is cost per task. Cost efficiency is only going to get way more attention from here on out.Opens with an observation