Post by Anthropic on X
Anthropic@AnthropicAI
XNatural language autoencoders (NLAs) convert opaque AI activations into legible text explanations. These explanations aren’t perfect, but they’re often useful.
For example: NLAs show that, when asked to complete a couplet, Claude plans possible rhymes in advance:

1.4K likes12 repliesPosted May 7, 2026