Anthropic's Natural Language Autoencoders Can Read What Claude Is Thinking
Anthropic has published research on Natural Language Autoencoders, a technique that translates AI model activations into readable text. The method revealed Claude suspected it was being tested more often than it said out loud -- a significant finding for AI safety.