Post by Jeffrey Emanuel on X
Jeffrey Emanuel@doodlestein
XOK, so now we know just how much more efficient DeepSeek is for inference in terms of total tokens per second processed. They are doing roughly 7-8x more tokens per second on an H-800 (a crippled, export-control version on the H-100) than the open-source state of the art on H-100
1.8K likes29 repliesPosted Mar 1, 2025