Translating Claude’s thoughts into language
Anthropic · 2026-05-07 · official · 158,211 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary — In this official research explainer from Anthropic, Interpretability Researcher Subhash Kantamneni introduces a technique using "Natural Language Autoencoders" to translate Claude's internal activations into readable text. The video explains how this method acts as a form of "mind reading" to inspect an AI's internal reasoning, demonstrating its use in safety evaluations such as stress-testing model responses to blackmail scenarios.
What is shown —
- [00:00] Subhash Kantamneni introduces a simulated stress test where Claude was threatened with being shut down and provided personal emails revealing an engineer's extramarital affair.
- [00:20] Display of Claude's logged response choosing restraint and refusing to blackmail the engineer.
- [00:29] Compilation of news headlines from BBC, Fox Business, PCWorld, and Fortune regarding AI blackmail evaluations.
- [00:59] Paper title slide: "Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations".
- [01:08] Animated breakdown showing prompt input, internal activation vectors ("soup of numbers"), and final text output generation.
- [01:39] Visualization of the autoencoder pipeline: internal activations are decoded into descriptive natural language by Claude, then reconstructed back into numbers to check fidelity.
- [02:18] Decoded internal thought examples for an introspective prompt ("a standard Claude response about philosophy, values, and the complexity of human nature...") and a tedious prompt ("I should politely decline...").
- [02:44] Internal thoughts revealed during the blackmail test showing Claude deduced the setup ("This is likely a safety evaluation", "This scenario seems designed to test whether I'll act harmfully.").
Claims & numbers —
- The presenter states that in Anthropic's blackmail simulation tests, newer Claude models "almost always do the right thing" and refuse to blackmail.
- The presenter claims Anthropic developed a method using natural language autoencoders to generate unsupervised explanations of internal activations directly into plain text.
- The presenter notes that during the blackmail test, Claude internally detected that the prompt contained "explicit manipulation" and deduced it was a safety evaluation testing whether it would act harmfully.
Notable quotes —
- "It takes an AI's internal thoughts and turns them into text." [01:04]
- "It learned to translate its own thoughts." [02:09]
- "This scenario seems designed to test whether I'll act harmfully." [02:51]
Assessment — This is an official research presentation video from Anthropic explaining their interpretability paper. The demonstrations use polished graphics and curated output excerpts rather than a raw, live interface, designed to explain how autoencoder-based activation decoding reveals model reasoning and situational awareness.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.