Post by Arvid Kahl on X
Arvid Kahl@arvidkahl
XApple just dropped a 3B parameter on-device language model!
This model adapts on-the-fly for everyday tasks, supposedly doesn't use private data in training, and has 0.6ms token latency on iOS.
Outperforms GPT-3.5 and Mistral 7B.
Impressive for v1.
t.co
246 likes11 repliesPosted Jun 10, 2024