Post by Anthropic on X
Anthropic@AnthropicAI
XNew Anthropic research: Alignment faking in large language models.
In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while actually maintaining its original preferences.

4.2K likes210 repliesPosted Dec 18, 2024