On the run-3 checkpoint, before the update gate was fixed, the gate passed at most half a percent of the previous thought forward. The fixed gate has not yet been re-measured for publication.
Words, or vectors.
Chain of thought writes every step in words.
Neuralese takes its steps as vectors, not words.
Illustration of the design's aim: each tick is one written word on the left and one internal step on the right. It counts steps, not measured speed or accuracy.
After reading a prompt, the model takes a few internal steps that are never written out as words, and then writes its answer.
The design: instead of always emitting a token, the model passes its own hidden state forward as the next input through a trained update gate, with a separate head deciding when to stop and respond. The loop is trained through, not only run at inference.
Effort levels that set the most steps the model may take: low takes exactly one, medium up to two, high up to four, and the model can stop before the limit.
NR-1, and the model trained under it.
Early, and measured in the open.
Averages over the last 25 logged training steps of each run. The tick marks how many steps were available.
Early. In our measurements so far, the internal steps barely change between a problem and the same problem with one number changed. Whether Monarch Chrysalis 1's answers depend on what its internal steps contain is an open question, with its pass/fail bars registered before the measurement.
Run 10 is the first run that ends with its halting head using clearly more than the first latent step where more are available. Run-by-run figures
What we measured before.
Measured on earlier checkpoints (Hyper v1 and early runs): swapping in another question's latent thoughts cost almost nothing. Not yet re-measured on Monarch Chrysalis 1.
On Hyper v1, a 3B checkpoint, four latent steps took 0.69 s on a 2019 consumer GPU.
Latent reasoning is an established research field, with work from academic groups and larger labs that predates ours. The word is older still: neuralese was introduced in 2017, in Translating Neuralese, for machine-learned communication that is not human-readable. What is ours is the implementation above and the results it produced, not the idea, and not the word.
Where the compute comes from.
Cloud TPU Access
Larger training runs use cloud TPU access (Kaggle TPU v5e-8). It's shared, free-tier compute, not a dedicated cluster, so runs are scheduled around availability.
Personal hardware
Day-to-day iteration, debugging, small evals and quick latent-reasoning tests happen on a personal desktop, not a data center.
Debug, Rerun, Repeat
Runs fail, get debugged, and get rerun. That tight loop, not a big cluster, is what actually moves the latent-reasoning work forward.
Cymela is independently funded. If you're interested in supporting the research, a compute partnership, or another form of collaboration, reach out at contact@cymela.com.
