Identifier
Cymela/monarch-chrysalis-1
Parameters
6.93B total · 6.44B in experts, ~1.3B active per token
Architecture
Sparse mixture of experts · 64 experts per layer, 8 active, plus latent-reasoning modules
Checkpoint
Thirteenth in the chain · cumulative step 12,461
Status
Early research model · not for deployment
Weights
Not yet released
Base model
Not Qwen. Named on release. This is not a variant of Hyper v1.

Monarch Chrysalis 1, a sparse mixture-of-experts (MoE) model designed for latent reasoning and trained under NR-1. Thirteen checkpoints in the chain have completed and the architecture works end to end. Still an early research model. Run 10 is the first run that ends with its halting head using clearly more than the first latent step where more are available. Whether the model's answers depend on what its latent thoughts contain is not yet established; the run-by-run figures are below.

Architecture

Chrysalis is a sparse mixture of experts with latent-reasoning modules built into it. 64 experts per layer, 8 active per token, about 6.44B of the 6.93B total parameters living in the expert stack. The latent loop runs on top of all of it: instead of decoding a token at every reasoning step, the model is designed to take internal steps before it answers, steps that are not written out as words.

Four latent steps are compiled. A learned halting head decides how many to weight. The base is an openly licensed sparse mixture-of-experts model, extended with our own modules rather than trained from scratch, the point is not capacity. Routing is the part of a sparse model we understand least, and a small model is the honest way to study it. The base model is named on release.

It runs

The step-2126 checkpoint was byte-verified at 13,865,515,469 bytes and runs end to end on a desktop: a Ryzen 7 5700X with 31 GiB of DDR4, entirely on CPU, no GPU involved. Cold load was 80 seconds on that checkpoint. After the load it is a terminal you can talk to, with the latent trace printing alongside every answer.

Speed on the run-9 checkpoint, the one before the current one, measured across the full 30-question suite and not yet re-measured on run 10: 7.91 generated tokens per second at one latent slot, 7.61 at two, 7.00 at four. That is end to end per question, total tokens over total seconds, including prompt build, prefill, the latent loop and the decode. It is not a decode-phase rate and does not compare with one. It is slow, and it is the honest figure: a latent step costs about 93% of what emitting a token costs, since only the vocabulary projection is skipped. Nothing here is faster than writing the reasoning out.

Two figures that used to sit in this paragraph, a decode rate of 3.69 tokens per second and a 3.18 average over a 29-item run, have been removed rather than restated. Neither can be reproduced from the training tree, so we cannot say what they measured or on which decode path, and a number we cannot trace is worse than no number. That is a gap in our own record, not a dispute about the figures. The run-9 checkpoint is a different file from step-2126, 13,867,635,018 bytes, hashed and matched against its identity record on 20 September.

That is what the architecture had to demonstrate, and it does. The routing works. The latent loop executes and its halting behaviour is measurable rather than theoretical. The model produces coherent English, and given an instruction rather than a question it follows the instruction and attempts the task.

Run progress

This section covers the whole training chain, to step 12,461. The detailed figures elsewhere on this page were measured on the step-2126 checkpoint and have not been re-measured; where this section disagrees with those, this section is the later reading.

Corrected 22 September 2026. Three things this section carried were wrong. Run 5's answer cross-entropy was given as 0.83 nats: that was one training core's row, and averaged over all eight cores it is 0.95, which did not fall during the run. The chain was counted two segments short. And a series from the offline causal battery was plotted as though it had given readings, when it has not yet given a verdict on any checkpoint; it is removed. Every figure below is read from the run ledger.

The completed chain in order, every one of them finishing without an out-of-memory kill:

root                     2,126
run 1                    4,331
run 2                    6,332
run 3                    8,333
checkpoint validations   8,349   8,365   (16 steps each)
run A                    8,861
run B                    9,869
run 5                   10,877
checkpoint validation   10,893   (16 steps)
diagnostic segment      10,901   (8 steps)
run 9                   11,693
run 10                  12,461   <- current

Every saved checkpoint the chain resumed from is a row, including four short validation and diagnostic segments of 8 to 16 steps. By that count, which is the one this page uses, run 10 is the thirteenth checkpoint in the chain. Counting training runs alone it would be the ninth. The run numbers themselves count attempts rather than completions, which is why there is no 6, 7 or 8 among the training runs: those full-length attempts were started and did not complete.

Two panels. The model got better at predicting text through run 3, and run 10 is the first run to end with its halting head using clearly more than the first thought step. Whether its answers depend on what those thoughts contain is not yet established. Every point below is the value at the end of a run, read from that run's own instrument; series from different instruments are never merged, and a run that produced no reading gets no point rather than a zero.

Training figures are averages over the last 25 logged steps of a run. From run B on they average the rows of all eight training cores at each step; the earlier runs logged one core's row per step. The training loss excluding routing terms has no eight-core version, so it is one core's row throughout, on the line and in the points drawn apart from it. The transplant-meter panel divides a run's mean swap cost by its mean answer cross-entropy, both taken over all of that run's meter fires, each fire first averaged across the eight training cores.

training loss, excluding routing termsreweighted objective, not comparableanswer cross-entropy, eight-core average
Training loss across the run chain Training loss excluding routing terms, one core's row, falls from 4.03 nats at the root run to 3.09 at run 3, where the line stops. The reweighted objective is plotted apart, never joined, also one core's row: run B 2.58, run 5 2.68, run 9 1.96, run 10 3.63. Answer cross-entropy, averaged over all eight cores, is plotted apart with no trend line: run B 0.94, run 5 0.95, run 9 0.76, run 10 1.84 on a different corpus and not comparable. Run A has no training log on file. 0 1 2 3 4 nats 4.03 3.59 3.26 3.09 2.58 2.68 1.96 3.63 0.94 0.95 0.76 1.84 new corpus root 1 2 3 A B 5 9 10 cumulative training steps, 2,126 to 12,461
The line stops at run 3, the last run trained on that objective. Later runs changed what the loss measures, so they are drawn apart and never joined to it. The answer cross-entropy is also drawn apart, without a trend line: run 10 trained on a different corpus, so its 1.84 cannot be read against run 9's 0.76. Run A, marked with a cross on the axis, has no training log on file, because it was never downloaded: that is a gap, not a zero.
Use of the latent thoughts across the run chain Three instruments, never merged. Expected latent steps, pooled, stay near 1.02 through run 3 on one core's row, then read 1.00 at run B, 1.05 at run 5, 1.09 at run 9 and 1.79 at run 10 on the eight-core average; at run 10 that is 1.50 of 2 at medium effort and 2.17 of 4 at high. The transplant meter read 2.9 percent at run 2 and 2.7 percent at run 3 and has given no reading since. The slot-differentiation ratio reads 0.34 at run B, 0.15 at run 5, 0.15 at run 9 and 0.18 at run 10, against a success mark of 0.60. expected latent steps used, pooled over effort levels run 10: 1.50 of 2 at medium, 2.17 of 4 at high 1.0 1.5 2.0 1.02 1.00 1.05 1.09 1.79 cost of swapping in another question's thoughts, share of answer cross-entropy 0% 4% 2.9% 2.7% no reading no reading no reading on any run since run 3 last thought slot still tells problems apart, relative to the first 0 0.60 0.60 would have been success 0.34 0.15 0.15 0.18 root 1 2 3 A B 5 9 10 cumulative training steps, 2,126 to 12,461
Three instruments on the latent thoughts, each on its own scale, because slots, a percentage and a ratio cannot share an axis honestly. Run 10 is the first run to end with the halting head using clearly more than the first step; every earlier run ended between 1.00 and 1.09. The slot-differentiation ratio stayed far below its bar, and the transplant meter has produced no reading since run 3. The offline causal battery is not on this chart: it has not yet given a verdict on any checkpoint.

The training loss (excluding routing terms) fell from 4.03 nats at the end of the first run to 3.09 nats at the end of run 3, the last run trained on that objective; the later runs changed the objective, so their loss is not comparable and is shown apart. The earliest answer cross-entropy on file is run B's, 0.94 nats; the latest, run 10, reads 1.84 nats on a different corpus, so the two are not comparable. The expected number of latent steps the halting head uses, pooled over every row that ran latent steps, started near 1.8 in the root run, where nearly every row (97.6%) computed four steps, and collapsed to 1.02 within it; with rows now computing one, two or four steps by requested effort level, it read 1.09 at run 9 (1.09 of 2.0 at medium, 1.10 of 3.9 at high) and reads 1.79 at run 10 (1.50 of 2.0 at medium, 2.17 of 4.0 at high), after a change to how the halting head is trained. On the 2 runs where the training-time transplant meter produced readings (run 2, run 3), swapping in another question's thoughts cost under 3% of the answer's cross-entropy; it has produced no reading since run 3. The offline test of whether the model uses its thoughts has not yet given a verdict on any checkpoint: on run 3, run B, run 5, run 9 it ran on a version of the test since corrected; on run 9 (re-read), run 10, run 10 (re-read) the model does not pass the test's own reading check, so nothing is quoted from it in either direction. No checkpoint is waiting to be graded. The slot-differentiation ratio (bar 0.60) stands at 0.18 at run 10. The training loss fell through run 3. Whether the model's answers depend on what its latent thoughts contain is not yet established.

Run 10 trained on a new corpus, after a change to how the halting head is trained. It is the first run that ends with its halting head using clearly more than the first latent step where more are available: 1.50 of 2 at medium effort and 2.17 of 4 at high; every earlier run ended between 1.00 and 1.09 steps on average. At medium effort the head still sits just outside the range registered for it in advance. Its answer cross-entropy, 1.84, is not comparable with run 9's because the corpus changed. Its offline thought-use test gave no verdict, and its slot-differentiation ratio is 0.18 against a bar of 0.60.

Halting behaviour

E[N] is the expected number of latent steps used, under the model's own halting distribution. Four slots are compiled; it weights about one. That figure holds across three environments and two independently written mixture-of-experts forward implementations:

TPU training telemetry   Kaggle TPU v5e-8, static path         1.023 - 1.070
Kaggle inference         Kaggle CPU, static path               1.068
local inference          Ryzen 7 5700X, sparse path            1.04  - 1.18
arithmetic probe         Ryzen 7 5700X, 29 items               1.04  - 1.16

All four slots are always computed. E[N] of about 1.09 does not mean three steps were skipped and the work was saved. It means the model computes four and weights its answer on roughly one. That is a finding about how the halting head has learned to behave so far, at step 2126 of a run that is still going.

The fourth row is a different kind of evidence from the first three and is worth separating. It shares hardware, weights and code path with the local inference row, so it is a fourth set of conditions rather than a fourth independent instrument, the cross-machine replication still stands at three. What it adds is that the collapse holds across outcome: the same E[N] on items the model answered correctly, answered wrongly, ran to the token cap on, and said nothing at all to. Prompt phrasing does not move it either. Phrasing changes whether the model talks. It does not change whether it thinks.

Measured behaviour

A 29-item arithmetic probe, with the scoring rules fixed in the source before any output existed, gives the clearest picture of what the checkpoint currently does. The partition matters more than the score:

A boundary added 2026-08-31, after these figures were published. These figures were measured on the step-2126 checkpoint, on the pre-fix recurrence, confirmed from the fetch log, the checkpoint history and a routing-deviation fingerprint in the run log itself, not from memory. Every inference tool in the project. This probe included, carried the prompt's hidden state forward raw between latent steps instead of through the learned update gate the model trains with. The gate's weights were in the checkpoint and loaded by all six tools and called by none of them. The tools have since been corrected and now share a single carry step, so what follows describes the path as it stood when these figures were measured rather than the code as it stands today. The numbers here are real measurements of what that path produced, and they are not measurements of the model as trained. The effect is not small: on a fixed benchmark suite, with the same checkpoint and decode settings and a byte-identical determinism control, correcting the recurrence moved 25 of 30 answers. Everything in this section is being re-measured on the corrected path, and both sets will be published side by side rather than one quietly replacing the other.

attempted                                   29
  emitted nothing at all                     4
  ran to the token cap, never committed     11
  committed an answer                       14   -> 3 correct

Of the 11 committed-wrong answers, one emitted a single token with no number in it. Of the remaining 10, six land within 2% of the truth and three of the other four are digit-count failures off by roughly a factor of ten. Carry load splits it cleanly: 3 of 5 right where the carry load is at most two, 0 of 9 where it is four or more. Those are the failure modes of a circuit computing badly, not of a lookup table being read wrongly.

Two behaviours matter more than the accuracy for anyone reading the traces.

The answer channel is faithful to work that is wrong. A prediction registered before the run. That where the model emits a fenced output block, its committed answer will equal that block, held on all 7 rows that could test it. 5 of those 7 blocks were wrong, and the answer transcribed each one exactly. In one case the model wrote Python containing an undefined variable, invented the result of running it, and then cited its own invention as confirmation. A downstream check comparing the answer against the working will find perfect agreement.

Prompt phrasing selects the failure mode. Same operands, one trailing clause between them: a bare question produces silence on 4 of 13 items; adding "verify your answer with Python" produces zero silences and 6 commitments out of 8; adding "state your confidence as a percentage" produces 7 non-terminating runs out of 8. Three distinct pathologies, not a severity gradient. If you are prompting this checkpoint, the phrasing decides which way it fails.

What it cannot do yet

The latent channel is not yet carrying the reasoning. By our own instruments it is not: transplant delta is approximately zero, and the arithmetic that does work arrives in visible English rather than through the latent path. Grounding is instrumented and reads above chance, but a defect found on 2026-08-31 means the figures it has produced so far are not yet a clean measurement, so no number for it is printed here. On the questions examined closely, the model's stated answer does not always follow from its own correct working.

That is the open problem and it is the entire point of the project. It is also now something that can be watched happening rather than inferred from telemetry, which is the difference this checkpoint makes. Training continues. The sharper dataset this page previously promised has since gone in, and it did not move the halting behaviour: E[N] measured on the training telemetry is 1.0152 across two consecutive runs, roughly 6,200 cumulative steps of training between them. That is a null result rather than an absence of one, and it is the honest state of the thing this project exists to test.

Licence and disclaimer

Weights are not released yet. This is an early research checkpoint of a training run still in progress, published because the architecture works end to end and that is worth recording. It is not a product, it is not tuned for use, and it should not be deployed for anything.

Figures on this page come from the training telemetry and from local inference runs. Where a number could not be measured, it is not printed. Anything that changes with the next run will be corrected here.

What is NR-1?

Cymela's framework for training and testing a model designed to take internal steps before it answers. Monarch Chrysalis 1 was trained under NR-1. What NR-1 covers, and where it stands

← All models