The current checkpoint compiles four latent thought slots and will use as many as you ask for. We ran the same 30 questions at one slot, two and four, changing nothing else. The control is exact: the slots it actually used matched the request on 30 of 30 rows at every setting. The answers did not follow. Strict scores were 9, 8 and 8 of 30, and 16 of the 30 answers came back byte-identical between one slot and four. One thing did improve, and it is not small: the model used to return nothing at all on roughly a quarter of questions, and now returns an answer every time.

A latent-reasoning model is supposed to get better when you let it think for longer. That is the whole promise, and it is testable in an afternoon: hold everything else still, ask for more thinking, and see whether the answers improve.

They did not. What follows is that test, and one genuine improvement found alongside it.

The depth control works exactly

Before the result is worth anything, the knob has to do what it says. One checkpoint, one decode configuration, 30 questions, and the only thing that changed between runs was the number of latent slots requested: one, two, then four. On every row at every setting, the number of slots the model actually consumed matched the number requested. Thirty of thirty, three times over. Whatever else is wrong, the model is genuinely thinking as long as it is told to.

The answers do not move

Correct answers by requested latent depth Strict scores are 9 of 30 at one latent slot, 8 of 30 at two and 8 of 30 at four. The line is flat within one answer across a fourfold change in requested depth. 0 10 20 30 correct the depth control itself was exact: slots used matched the request on 30 of 30 rows at every setting 9 8 8 1 slot 2 slots 4 slots correct answers of 30 · latent depth requested, one run per setting
One checkpoint, one decode configuration, and only the requested depth changed. Asking for four times the thinking moved the score by one answer, in the wrong direction. The flatness is the finding; with 30 questions a single answer is inside the noise, which is the point.
Strictly graded: every row read, and every code answer parsed and executed rather than eyeballed.

Nine, eight, eight. Across a fourfold increase in thinking the score moves by a single answer, and downward. With 30 questions that difference is noise, which is precisely the claim: four times the latent computation buys nothing measurable.

The per-answer comparison is blunter than the score. Between one slot and four, 16 of the 30 answers are byte-identical: not similar, the same bytes. Of the answers that did change, one went from wrong to right and three went from right to wrong.

one slot against four, same 30 questions

byte-identical answers                      16
changed, wrong to right                      1
changed, right to wrong                      3

tokens generated      1 slot   2 slots   4 slots
                       1,123     1,062     1,012

The last row is the one we did not expect. Asked to think harder, the model writes less. Not by much, and we are not going to build a theory on 111 tokens, but it is the opposite of the direction a working reasoning loop would push.

What did improve

Blank answers, old decode path against the current one On the older decode path the model returned nothing on 8, 6 and 8 of 30 questions across three runs. On the current path at three effort levels it returned nothing on 0 of 30 each time, 0 of 90 in total. 0 5 10 blank 8 6 8 0 0 0 older decode path · run-3 checkpoint current path · run-9 checkpoint run 1 run 2 run 3 low medium high blank answers of 30 · three repeats on the old path, three effort levels on the current one
The older path returned an empty answer on 22 of 90 attempts, with two questions blank in all three runs. The current path returns an answer every time, 90 of 90. This does not isolate a fix: the checkpoint and the decode path both changed, so the comparison names the improvement without attributing it.

On the older decode path the model returned an empty answer on 22 of 90 attempts, and two questions came back blank in all three runs. On the current path, across three effort levels and 90 answers, it never once returned nothing.

That is a real improvement and we are not going to undersell it: a model that silently returns nothing is unusable in a way that a model returning a wrong answer is not. But it is a narrower result than it looks. The checkpoint changed and the decode path changed, so this comparison names the improvement without attributing it to either. And the two questions that used to come back blank now come back wrong: one of them asks for 49 and answers 15.0. Silence became a wrong answer, not a right one.

What this does not establish

It does not establish that latent depth cannot help. It establishes that on this checkpoint, on this decode path, on these 30 questions, it did not. A mechanism that works only at depths we have not reached is still possible, and it is not what we measured.

It is also not a regression claim. An earlier figure of 10 of 30 exists and was produced on a different decode path, so the comparison would not mean anything and we are not drawing it.

Boundaries. Every figure here is one run per setting, N=30 questions, measured 2026-09-19, except the blank-answer baseline which was measured 2026-08-31 on the run-3 checkpoint. Grading is strict and was done by reading every row, with code answers parsed and executed rather than judged by eye; the run mechanics were reproduced independently by another session, the grading was not. A one-answer difference on 30 questions is not a difference and is not reported as one.

← All research entries