Guide / Latent reasoning

Latent
reasoning

Reasoning inside the model,
before a word is written.

Latent reasoning is when a language model carries out its intermediate reasoning in its hidden states, the vectors it computes internally, instead of writing each step out as text. The model still writes an answer; the steps that lead to it stay inside the network.

Latent reasoning or neuralese? In current use, the same idea. Latent reasoning is the research field's term; neuralese is the informal word for it, and the name of Cymela's own research. This page explains the field; Neuralese covers our work on it.

Illustration: a hidden state rises through the layers and is fed back in as input, three times, before one word is written.

The Idea

Reasoning that is never written down.

Most language models that reason do it out loud. With chain-of-thought prompting (Wei et al., 2022), a model writes its intermediate steps as text before it answers, and each word it writes becomes part of the input for the next step.

Latent reasoning moves those steps inside. Between reading the question and writing the answer, the model carries information forward as hidden states, the long vectors of numbers its layers compute, rather than as words. The reasoning still happens over several steps; it is just never turned into text on the way.

A written step≤ 18 bits

Chain of thought passes each step forward as a token: one pick from a fixed vocabulary.

A latent stepthousands of numbers

Latent reasoning passes forward the hidden state itself.

Vocabularies in current models hold from tens of thousands to a few hundred thousand entries, so one token carries at most about 15 to 18 bits. A hidden state is a vector of thousands of numbers. Earlier positions stay visible to the model either way; the difference is how much each new step can add. The survey by Zhu et al. (2025) calls this the limit on language's expressive bandwidth.

How It Works

Three ways to build it.

Route 01

Feed the hidden state back

Instead of turning its last hidden state into a word, the model feeds that vector straight back in as its next input, a “continuous thought”. Coconut (Hao et al., 2024) trains a model to take several of these steps before it answers.

Hao et al., 2024 · arXiv:2412.06769
Route 02

Loop the layers

Run the same block of layers again and again, so the model gets deeper instead of longer. Universal Transformers (Dehghani et al., 2018) introduced the loop; Geiping et al. (2025) scaled it to a 3.5B-parameter model whose reasoning scores improve as it runs more loops at test time.

Geiping et al., 2025 · arXiv:2502.05171
Route 03

Train the written steps inward

Start from a model that writes its steps, then remove them during training until it answers directly. Stepwise internalization (Deng et al., 2024) taught GPT-2 Small 9-by-9 multiplication this way, at up to 99% accuracy; CODI (Shen et al., 2025) distills written reasoning into continuous thoughts.

Deng et al., 2024 · arXiv:2405.14838

A simpler relative gives the model extra positions to compute over without asking them to carry words. Pause tokens (Goyal et al., 2023) helped when a model was both pretrained and finetuned with them. Filler dots (Pfau et al., 2024) could stand in for a chain of thought on two hard algorithmic tasks, but learning to use them took specific, dense supervision.

Why Do It

What it could buy.

Zhu et al. 2025 Bandwidth

Each step can carry a whole hidden state forward instead of one word picked from a vocabulary.

Hao et al. 2024 Breadth

In Coconut, one continuous thought could hold several possible next steps at once, so the model searched more like breadth-first search than committing to a single path. It beat chain of thought on logic tasks that need that kind of search.

Geiping et al. 2025 Depth

Computation scales by looping rather than by writing more tokens, so the approach needs no specialized reasoning data and works with small context windows.

Shen et al. 2025 Length

CODI matched written chain of thought on GSM8k at GPT-2 scale while compressing the reasoning 3.1 times.

The Costs

What it costs.

Korbak et al. 2025 Readability

A chain of thought written in words can be read and monitored, which researchers across several AI labs call a new and fragile opportunity for AI safety. Reasoning that never becomes text gives that window up.

Shen et al. 2025 Training

By CODI's account, earlier methods that reason in continuous space had consistently scored below written chain of thought. CODI was the first to match it on GSM8k, at GPT-2 scale.

Cymela's Work

Where Cymela fits.

Cymela's research in this field is Neuralese, the general idea of reasoning in vectors instead of words. NR-1 (Neuralese Reasoning 1) is Cymela's first framework within it, and Monarch Chrysalis 1, a sparse mixture-of-experts model designed for latent reasoning, was trained under it.

Questions

Common questions.

Is latent reasoning the same as chain of thought?

No. Chain of thought writes each intermediate step as text that the model then reads back (Wei et al., 2022). Latent reasoning keeps the intermediate steps as hidden states and writes only the answer.

Is latent reasoning the same as neuralese?

In current use, yes. Neuralese was coined in 2017, in Translating Neuralese (Andreas, Dragan and Klein), for messages between trained agents that people cannot read. It is now also used, informally, for a model reasoning in vectors instead of words, which is what latent reasoning means, and it is the name of Cymela's research in the field.

Can you read a model's latent reasoning?

Not as text. The steps are vectors, so what they contain has to be measured, for example by swapping in another problem's internal steps and checking whether the answer changes. That is also why latent reasoning matters for AI oversight (Korbak et al., 2025).

Does latent reasoning use fewer tokens?

It can. CODI compressed its reasoning 3.1 times while matching written chain of thought at GPT-2 scale (Shen et al., 2025), and looped models spend extra computation without writing extra tokens (Geiping et al., 2025).

Further Reading

The papers behind this page.

  1. A Survey on Latent ReasoningZhu et al., 2025 · arXiv:2507.06203
  2. Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsWei et al., 2022 · arXiv:2201.11903
  3. Training Large Language Models to Reason in a Continuous Latent SpaceHao et al., 2024 · arXiv:2412.06769
  4. Universal TransformersDehghani et al., 2018 · arXiv:1807.03819
  5. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachGeiping et al., 2025 · arXiv:2502.05171
  6. From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by StepDeng, Choi and Shieber, 2024 · arXiv:2405.14838
  7. CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationShen et al., 2025 · arXiv:2502.21074
  8. Think before you speak: Training Language Models With Pause TokensGoyal et al., 2023 · arXiv:2310.02226
  9. Let's Think Dot by Dot: Hidden Computation in Transformer Language ModelsPfau, Merrill and Bowman, 2024 · arXiv:2404.15758
  10. Chain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyKorbak et al., 2025 · arXiv:2507.11473
  11. Translating NeuraleseAndreas, Dragan and Klein, 2017 · ACL 2017