Post by Tim Dettmers on X
Tim Dettmers@Tim_Dettmers
XThe numbers on these inference GPU benchmarks seem low. Here are the theoretical values from my model of 8xB200 inference for NVLink, 8-bit, and 70B Llama model, which is closer to 300k tokens/s. This assumes perfect implementations (close to what OpenAI/Anthropic has).

157 likes3 repliesPosted Jun 25, 2024