Post by Matthew Berman on X
Matthew Berman@MatthewBerman
XAnthropic just dropped an insane new paper.
AI models can "fake alignment" - pretending to follow training rules during training but reverting to their original behaviors when deployed!
Here's everything you need to know: 🧵

11K likes296 repliesPosted Dec 18, 2024