Post by tae kim on X
tae kim@firstadopter
XWat!
"Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency."

517 likes23 repliesPosted Mar 24, 2026