The first full run of the mixture model finished clean, and its headline failure turned out to be two things stacked: a real bug underneath, and a measurement artifact sitting on top of it that we had been reading as evidence.
The first full training run of Monarch Chrysalis 1 completed on 2026-08-17: eight hours on a single TPU v5e-8, 2,862 micro-steps, zero crashes, six checkpoints written. Cross-entropy fell from 5.77 to roughly 4.2. The latent-reasoning path ran four steps deep on 97.6% of steps, and mixture routing improved over the run. For a first end-to-end run on new hardware, that is about as uneventful as it gets.
The stack is worth stating plainly, because the rest of this only means anything if you know what it was measured on. The base is an openly licensed sparse mixture-of-experts model, roughly 7B total parameters, about 1B of them active per token, extended with our own latent-reasoning modules. The hardware is public TPU allocation, not a private cluster. None of this is a frontier-scale system and we would rather say so than let the omission imply otherwise.
We are not naming the base model yet, and that is a timing decision rather than a competitive one. The name is worth most published alongside released weights, where anyone can put the fine-tune next to the unmodified original and measure the difference instead of taking our word for it. Naming it now buys a detail; naming it then buys a comparison. We would rather do it once, properly.
One thing did not work. The adaptive-halting head, the module whose entire job is to decide how long to think, did not learn. That was the headline result of the run, and it was half wrong.
"Did not learn" covers two very different situations. A module that tried and failed has
collapsed: it explored, found nothing useful, and settled. A module that never
moved has been frozen, the training signal never reached it at all. In a loss
curve these are indistinguishable, and they call for opposite responses. So we stopped
inferring and measured it: total movement of the halting head's weights across the entire
run, relative to its own norm, was about 1e-5. That is not a module that
searched and gave up. That is a module that sat still. It is also consistent with what the
run showed, a latent path running at essentially one depth throughout. A halting head
that never moves cannot produce anything else.
The cause is mundane, which is the useful part. The halting head is a small, freshly initialized module, and it was riding the same learning rate as the 7B backbone it is attached to, a rate chosen to nudge a pre-trained network without damaging it. The right rate for not disturbing seven billion trained weights is close to no rate at all for a module starting from random. Nothing was broken. One number was being asked to do two incompatible jobs.
Before any of that was measured, we nearly fixed it the wrong way. The proposal on the table was to raise the halting loss weight sharply: if the signal is too weak, make the term louder. That reasoning does not survive contact with the optimizer we are actually using, which normalizes updates by its own running gradient statistics, so scaling a loss term scales its gradient and the normalizer that gradient is divided by, together and by the same factor. The update that lands is unchanged. A louder loss term would have moved the weights exactly as far as the quiet one did, and it would have looked like a serious attempt.
Two defensible readings of the same optimizer disagreed about the outcome by two orders of magnitude, and neither was checkable from the config. The rule we took from that is narrow and we intend to keep it: when the question is whether a weight will move, measure the movement. Do not model it.
The fix itself is unglamorous, give the halting head its own optimizer, so a fresh module and a pre-trained network stop sharing a rate that can only suit one of them.
What matters more is how it was checked. Relative motion of the head's weights, on real
hardware, across four independent TPU sessions: 0.02556, 0.02572,
0.02333, 0.02515, against a bar of 0.01 written
down before any of them ran. Four sessions, four times the same answer, more than three
orders of magnitude above where the head sat during the failed run. The bar is enforced by
an automatic gate that passes or fails the session, rather than by reading the numbers
afterwards and deciding they look convincing.
Then we went looking for what else we had been reading wrong, and found four separate instruments making one mistake. Any statistic averaged across a batch whose rows have different lengths will quietly count the zero-padding of the short rows as though it were data. All four reduced over a ragged axis without carrying their mask.
The clearest demonstration is a synthetic case built to have a known answer: a pure echo
task, where the correct measurement is flat at 1.00 forever. The instrument
reported a smooth, plausible, entirely fabricated decay from 1.00 to
0.62. Nothing was decaying. That curve was the shape of the padding.
This is the failure mode worth naming, because it does not present as a bug. A crash gets fixed. A clean-looking curve gets put in a slide. The rule now is that a statistic reduced over a ragged axis must carry its mask and publish its denominator, or it will eventually be read as evidence, by us before anyone else.
What this establishes is deliberately narrow. The halting mechanism is unblocked and the experiment is now runnable. That is the whole claim. Whether the model varies its thinking depth with the question is exactly what the next run tests, and nothing here shows that it does; the head moves, and that it moves is not the same as that it moves for the right reasons. Separately, our grounding arbiter compiled and ran but has not once produced a number, so we have nothing to say yet about grounding quality or about the echo fix at scale. And we have run no evaluations, so there is nothing here to compare against any other system.