Post by Yuchen Jin on X
Yuchen Jin@Yuchenj_UW
X2.5x faster but 6x more expensive.
This can’t be achieved by inference optimization, must be new chips.
TPU? B200? AWS Inferentia? Cerebras?
1.3K likes95 repliesPosted Feb 7, 2026
ViralLoop
Your next post already worked.
Loading the post
Private to you, and used when you remix this post.
Analysis by ViralLoop — not written by the original author.