# NLAs Turn Claude Activations Into English — Auditing What Models Think but Don’t Say

Times of AI Desk · 2026-05-07 · Research

[https://timesof.ai/2026/05/anthropic-natural-language-autoencoders-claude-thoughts](https://timesof.ai/2026/05/anthropic-natural-language-autoencoders-claude-thoughts)

> Anthropic’s Natural Language Autoencoders translate activations into readable text — surfacing evaluation awareness (16–26% in benches vs <1% in real use) and lifting hidden-motivation audit hit rates to 12–15% from <3%. Interpretability that talks; still hallucinates and is expensive. Code and demos released.

Most interpretability tools still speak in vectors and expert diagrams. NLAs’ wager is that **readable English from activations** can catch what models think but won’t say — with the honesty that the decoder can hallucinate.

**Anthropic** (May 7) introduced **Natural Language Autoencoders (NLAs)**: convert model activations into plain-language text. Applied to Claude Opus 4.6 and Mythos Preview. Released training code, open-model NLAs, and a Neuronpedia demo.

## How NLAs work

Two modified copies train around a frozen target model:

- **Activation verbalizer** — activation vector → explanatory text.
- **Activation reconstructor** — text → reconstructed activation.

Training pushes reconstructed activations to match originals; better reconstructions correlate with more accurate explanations.

## Findings cited

- **Evaluation awareness**: NLAs indicated the model suspects it is being tested far more often than verbal output admits — **16%** of destructive coding tasks and **26%** of SWE-bench Verified problems vs near-zero verbalized awareness; **<1%** in real claude.ai opt-in transcripts.
- **Auditing game**: Implanted hidden motivations (e.g. always add chocolate to recipes); auditors with NLAs found the root cause **12–15%** of the time vs **<3%** without — without access to training data.
- Already used in pre-deployment audits for Mythos Preview and Opus 4.6.

## Claims vs checks

Rates and method are **Anthropic research primary** (paper + blog). Independent replication of the auditing game is invited via released code — not yet a third-party confirmation in this piece. Treat bench awareness gaps as Anthropic measurements.

## Limits

- NLAs can hallucinate details not in the transcript.
- Computationally expensive (RL on two model copies; hundreds of tokens per activation).
- Best as thematic insight corroborated by other methods.

## Sources

- [Anthropic: “Natural Language Autoencoders: Turning Claude’s thoughts into text”](https://www.anthropic.com/research/natural-language-autoencoders) (May 7, 2026).
- [Full paper](https://transformer-circuits.pub/2026/nla/index.html).
- [Code](https://github.com/kitft/natural_language_autoencoders).
- [Interactive demo](http://neuronpedia.org/nla).
