The current checkpoint compiles four latent thought slots and will use as many as you ask for. We ran the same 30 questions at one slot, two and four, changing nothing else. The control is exact: the slots it actually used matched the request on 30 of 30 rows at every setting. The answers did not follow. Strict scores were 9, 8 and 8 of 30, and 16 of the 30 answers came back byte-identical between one slot and four. One thing did improve, and it is not small: the model used to return nothing at all on roughly a quarter of questions, and now returns an answer every time.
We measured, on the current checkpoint, whether the answer uses the latent thought vectors at all. With the question hidden so the answer must come from the thoughts, swapping in another question's thoughts costs 0.001 nats on 4.409, and correct answers do not fall. The thoughts are not empty and they are question-specific; nothing downstream decodes them. The cause sits upstream of the architecture: at the learning rate and storage format these runs used, most of the model's weights received updates too small to be written, and never moved.
The halting head has not moved in three training runs, and we could not say why. It is clamped between 0.01 and 0.99, and 71% of its values sit above the ceiling, where a clamp has exactly zero gradient. The line was added deliberately, with a comment explaining that it was there so the collapse could not happen.
We had a result ready to publish: the orchestrator knows when to stay silent, and our decision rule was throwing that knowledge away. Running it on the model that actually ships inverted the diagnosis, on a difference of five rows out of fifteen, indistinguishable from noise. So we built a bigger evaluation and ran it again. The finding was an artifact of fifteen rows, and the thing that needed fixing was never the model.
A prediction registered before the run. That where the model emits an output block, its committed answer will equal that block, held on every row that could test it. 5 of the 7 blocks were wrong, and the answer copied each one exactly. A faithful answer channel reporting unfaithful work is harder to catch than a model that is simply unreliable.
Monarch Chrysalis 1, a sparse mixture-of-experts (MoE) model with latent-space reasoning. The model's first training run is finished, and the architecture now provably works end to end. Still an early research model, but training is ongoing and the model is improving.
The first full run of the mixture model finished clean, and its headline failure turned out to be two things stacked: a real bug underneath, and a measurement artifact sitting on top of it that we had been reading as evidence.
The config said the model was thinking eight steps deep. The logs said one. The logs were right, and they had been saying so for months, printed on every line, next to a number nobody read against the setting that was supposed to produce it.
The dispatch pattern ran live for the first time, small models, a synthetic training set under 7,000 examples, and it did what it was designed to do: catch a stale answer before it shipped, and stay out of the way twice when nothing needed catching.
A causal audit swapped the transmitted message for the wrong one, and most of the reported improvement survived. We rebuilt how we measure because of it.
We fitted a map between two real models, beat the do-nothing baseline decisively on reconstruction error, and it made behavior worse. This is our own data arguing against our own instinct.
A merged 61-layer model emitted nothing but question marks. The cause was not the seam, not the wrapper, and not numerical instability. It was 36 routers full of zeros.
A model that writes out its reasoning is doing something strange: compressing a rich internal state into a narrow channel, then rebuilding it from the compressed form. We think that step is lossy and skippable.