Welcome to the J-Space: Anthropic's New Technique for LLM Interpretability
Arivu · 2026-09-25 · community · 15,060 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This is an animated conceptual explainer video exploring mechanistic interpretability techniques attributed to Anthropic research, focusing on the "J-Space" (Jacobian space) and "J-Lens". The narrator uses cognitive science analogies, calculus concepts, and geometric animations to explain how high-dimensional hidden activations can be interpreted and steered using the Jacobian matrix.
What is shown
- [00:19] Modular AI concept diagram breaking an AI system down into Vision, Language, Memory, and Tools/Planning modules.
- [01:08] Global Workspace Theory theater analogy showing modules in an audience, a bottleneck stage illuminated by a spotlight, and global broadcasting.
- [01:36] The "J-Space" shared vector space diagram ($v \in \mathbb{R}^d$) and representation of internal hidden states as points in a multidimensional cloud.
- [02:22] Introduction of the "J-Lens" representing the Jacobian matrix around an activation point, illustrating directional sensitivity vectors.
- [02:46] 1D calculus slope analogy ($m = \Delta y / \Delta x$) expanding into thousands of dimensions.
- [03:41] Jacobian matrix formulation: $J = \left[ \frac{\partial y_i}{\partial x_j} \right]$ and the linear approximation $\Delta y \approx J \Delta x$.
- [03:55] Visualization of flat directions (where output barely reacts) versus steep directions that matter.
- [04:38] Direction labeling (sentiment, formality, confidence) tied to semantic changes in output text.
- [05:15] Activation steering demonstration using $h_{\text{new}} = h + \alpha v_{\text{feature}}$, showing output text transitioning from "This is a disaster" to "This is disappointing" to "This is wonderful!".
- [05:30] Demonstration of the locality of sensitivity maps as the activation moves across the space.
Claims & numbers
- The narrator claims neural networks operate across thousands of hidden dimensions where only a few "steep directions" matter, while the majority are "flat directions" where output changes negligibly.
- The video states the linear approximation formula $\Delta y \approx J \Delta x$ describes output response to perturbations in hidden states.
- The narrator claims that because of superposition, a labeled direction rarely corresponds cleanly to a single concept, as concepts smear across directions.
- The video presents activation steering using the formula $h_{\text{new}} = h + \alpha v_{\text{feature}}$ to edit model behavior in real time.
Notable quotes
- [01:01] "If they never share, you don't get intelligence. You get a room full of experts, all talking at once, and no one listening."
- [04:54] "Interpretability has quietly become geometry."
- [05:24] "That's steering: editing behavior by adding a feature direction back into the activations."
Assessment
An educational, animated explainer breaking down mathematical and mechanistic interpretability concepts (Global Workspace Theory, Jacobian sensitivity matrices, superposition, and activation steering). The visuals are stylized geometric animations rather than direct terminal or model interface captures, serving as a pedagogical demonstration of interpretability theory.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.