Post by Awni Hannun on X
Awni Hannun@awnihannun
XQwen3.5 runs quite well in mlx-lm.
Awesome that we have a frontier-level hybrid model. The context gets longer but the inference speed and memory use barely change.
Here's the Q4 generating a space invaders game on an M3 Ultra. Generated 4,120 tokens at 37.6 tok/s.
241 likes15 repliesPosted Feb 16, 2026