Post by Awni Hannun on X
Awni Hannun@awnihannun
XNemotron 3 Nano runs nicely with mlx-lm on an M4 Max.
Could be a great model for local use on Mac: MoE + hybrid attention make it fast even for very long context.
Generating in realtime with 4-bit model:
687 likes14 repliesPosted Dec 16, 2025