The config said the model was thinking eight steps deep. The logs said one. The logs were right, and they had been saying so for months, printed on every line, next to a number nobody read against the setting that was supposed to produce it.

The whole point of this model is depth: instead of emitting a token, it passes its own hidden state forward and thinks again, several times, before it answers. The trainer was configured for eight such steps. What actually reached the loss, for the first 79,137 steps of an 82,697-step run, was one.

The depth a sample can actually reach is capped by how much of the problem that sample contains. You cannot think across more of a problem than is in front of you. Other settings, none of them named anything like "depth", held that cap at one on every row. The configured depth was never wrong so much as never reachable.

The cost is the part worth sitting with. Seven of every eight latent iterations still ran, full compute, full memory, the whole forward pass, with every slot masked out of the loss. Months of a shared free-tier TPU allocation spent computing gradients that were thrown away before they could teach the model anything. Not a crash, not an error, not a single failed run. Just a number quietly smaller than it looked.

It was found by reading the training logs against the config, which is to say it was found the way these things always are: by someone finally checking whether the thing being printed matched the thing being asked for. The progress bar had been reporting latents=1.0 on every step, in plain text, the entire time. The information was never hidden. It was just never compared.

The fix removed the ability for those settings to disagree with the one that is supposed to govern them. Depth of two or more trained consistently only from step 80,386. So the released checkpoint carries 2,311 steps of genuine multi-step latent training out of 82,697, about three percent of the run. We are publishing that number rather than the configured one.

What those 2,311 steps bought is real and we won't undersell it either. Before the fix, thinking longer made the model worse and the best measured depth was two. After, the best measured depth is four and the benefit of thinking over not thinking roughly tripled. The mechanism went from actively harmful to measurably useful in three percent of a training run, which is either an encouraging signal about the method or an uncomfortable one about the other 97%.

There is a ceiling, and it is close. Four steps is not just the best setting, it is nearly the last useful one: by eight steps the model does worse than if it had not thought at all. Push it to 100, 25 times the trained depth, and it breaks in a way we did not expect. Asked "Hi there", it did not produce gibberish. It produced the next user turn: user: Do me a quick bio of hyper, just the highlights. Fluent, grammatical, contextually appropriate, and playing the wrong part. It had started role-playing the human side of the conversation.

Our current explanation is that each latent step occupies a real position in the key-value cache. The prompt ends mid-assistant-turn, but after a hundred latent slots the model sees roughly a hundred positions elapsed since that boundary, having been trained with at most four. It reads the gap as "the assistant already spoke, at length" and does the only sensible next thing in a chat transcript: it starts a user turn. Its own halt signal had saturated before the end. It was asking to stop and being overridden. That reading is not yet isolated; a depth sweep to find exactly where the role boundary collapses is the next experiment, and it is a better one than we would have thought to run.

The honest counterweight: a 3B model running a hundred recurrent latent steps and still emitting grammatical, on-topic English is not nothing. A month ago the same run produced two hundred tokens of noise, not one coherent sentence. The failure that replaced it is structural, not linguistic. It is answering the wrong question correctly.

Depth eight was never reachable in training regardless: it runs out of memory on the chip we train on. The configured value was aspirational in a second, more boring way.

The re-evaluation this unblocked came back mixed, and the bad half matters more. The model's latent trajectory is now well formed and its answer genuinely depends on it, shuffle a problem's own thoughts and the answer degrades sharply. But transplant a different problem's thoughts in and it costs almost nothing: about four percent of the damage that reordering its own does, and it hurts on only 8 of 12 prompts, which is close enough to a coin flip to mean little. A separate probe put "prefers its own answer over another prompt's" at exactly chance. The model has learned to think. It has not yet learned to think about the question, and until that changes, every number above describes a mechanism that works in general rather than one that works on your problem.

← All research entries