All videos
0:00 / 0:00
research

Translating Claude’s thoughts into language

Anthropic15 June 2026Watch on YouTube

Part of series

Ep. 1 · Claude Code Video

View the series

Description

AI models like Claude talk in words but think in numbers. These numbers, called activations, encode Claude’s thoughts, but not in a language we can read. We are introducing Natural Language Autoencoders, or NLAs, which translate AI models’ activations into readable text. NLAs have already helped us improve how we test our models for safety and better understand why they do what they do. Read more about this research on our blog: https://www.anthropic.com/research/natural-language-autoencoders

What you'll learn

  • Claude and other AI models process information as numbers (activations) rather than words
  • Natural Language Autoencoders (NLAs) translate these internal numbers into human-readable text
  • NLAs improve safety testing and help us better understand how AI models work

Frequently asked questions

How do AI models think differently than how they speak?
AI models like Claude speak in words but think in numbers called activations. These numbers encode the model's thoughts but not in a language humans can understand.
What are Natural Language Autoencoders and what do they do?
NLAs are a technique that translates the internal activations of AI models into readable text. This allows researchers to see and understand what the model is doing internally.
How do NLAs help improve AI models?
NLAs have helped Anthropic improve safety testing and better understand why AI models do what they do, leading to better interpretability and safety.

Topics

In this video

Related reads