A model that writes out its reasoning is doing something strange: compressing a rich internal state into a narrow channel, then rebuilding it from the compressed form. We think that step is lossy and skippable.

When a model reasons by generating text, every step passes through a bottleneck. The internal state is high-dimensional and continuous; a token is one choice from a fixed vocabulary. Everything the model knew that did not fit into that choice is discarded, and the next step has to reconstruct it from what survived.

Our thesis is that reasoning can be carried forward as continuous internal state instead, with the model only spending words when it has something to say. That is what Hyper is built to test, and the mechanism is real and running, a learned gate mixes each new internal thought with the previous one, and a separate head decides when to stop thinking and start answering.

We are being deliberately careful about what we claim from it. On our own fixed benchmark, latent reasoning scored at or slightly above ordinary generation at the depth it was trained for, and clearly worse when pushed well past it. That is not a win. It is a working mechanism with an unproven benefit, which is exactly how we describe it.

This track gates the MultiThink work in practice. A model that cannot think in latent space cannot meaningfully send latent thought to another model.

← All research entries