Post by Matthew Leavitt on X
Matthew Leavitt@leavittron
XTwo things I'm particularly proud of here:
1. The pretraining data are derived entirely from publicly-available tokens.
2. No closed-source models were used in any part of the pretraining data curation pipeline.
393 likes14 repliesPosted Apr 2, 2026