Post by elvis on X
elvis@omarsar0
XSummary of today's OpenAI announcement:
- introduces reinforcement fine-tuning (RFT) of o1
- tune o1 to learn to reason in new ways in custom domains
- RFT is better and more efficient than regular fine-tuning; needs just a few examples
1/n
427 likes14 repliesPosted Dec 6, 2024