Read the release verdict before the numbers: against our own stated bar, this checkpoint reads PARTIAL. It does not clear it. The mechanism works; the thing it was supposed to demonstrate is only half demonstrated, and we would rather say so here than let you find out after downloading 6 GB.
Architecture
Hyper recycles its final hidden state back into inputs_embeds for K recurrent
steps before emitting any token, continuous latent chain-of-thought, in the manner of
Coconut. Three small added modules constitute the entire mechanism: one that
normalises the latent state, one that gates how much of the new state is carried
forward, and a PonderNet-style head that decides when to stop, totalling 13.6M
parameters, about 27 MB.
The backbone is a fully fine-tuned Qwen2.5-3B-Instruct, 36 layers, hidden size 2048,
vocabulary 151,668 after adding <bot>, <eot> and
<step>. The halt head was trained to choose K adaptively. In this
release it does not; see limitations.
On the parameter count: summing the stored tensors gives 3,409,648,641, and that number
is wrong. config.json sets tie_word_embeddings, and
lm_head.weight ships as a bitwise-identical copy of
model.embed_tokens.weight, so a 310.6M table is counted twice. The unique
parameter count is 3,099,032,577. It cross-checks: Qwen2.5-3B-Instruct
is 3,085,938,688, our vocabulary is 268 rows smaller (−548,864), giving a 3,085,389,824
backbone plus 13,642,753 of latent modules. Two derivations, same figure. Anywhere you
see 3.4B for this model, including on this site until 16 August 2026. That is the
double-counted number.
Training
Final checkpoint is step 82,697, trained on a Kaggle TPU v5e-8. The final leg ran from 79,137 to 82,697, 3,560 steps across three sessions of roughly eight hours, at depths of k=2 to k=4.
Being precise about what that means, because it is the number most likely to be misread: for the overwhelming majority of those 82,697 steps the model trained at an effective latent depth of one. A configuration bug held it there, the finding is written up in We trained 79,000 steps at a latent depth of one. We can firmly attribute 3,560 steps to the genuine depth-2-to-4 regime. The exact boundary before that is not precisely established, and we are not going to round it. This is not 82,697 steps of latent training.
We are not publishing the training corpus. Its provenance and licensing have not been verified to a standard that would let us assert anything about it in public, and an unverified claim is worse than an absent one.
Evaluation
Every causal number below is N=12 prompts. That is a small sample and it belongs next to the results rather than in a footnote. Treat these as directional.
Mean cross-entropy on the gold answer, by latent depth, for this checkpoint against an earlier one (step 80,189). Lower is better.
| Checkpoint | k=0 | k=1 | k=2 | k=4 | k=8 |
|---|---|---|---|---|---|
| 82,697 | 12.112 | 10.110 | 8.967 | 8.555 | 10.099 |
| 80,189 | 10.028 | 9.583 | 8.921 | 9.845 | 10.795 |
Thinking is worth −3.557 nats from k=0 to k=4, up from −1.107 at the earlier checkpoint, and the optimum moved from k=2 to k=4. Past that it falls off fast: by k=8 the model does worse than if it had not thought at all.
Interventions at k=4, measured as added cross-entropy. Shuffling a problem's own latent steps costs +1.389 and hurts 12 of 12 prompts. Replacing them with noise costs +1.125 and hurts 11 of 12. Transplanting a different problem's latents costs +0.055 and hurts 8 of 12, close enough to a coin flip to mean very little. A replay control reproduced the clean run bit-exactly, so the harness itself is sound.
A separate probe found the gold token's median rank improved from 8,303 to 23 out of 151,668, against a chance rank of 75,834. But the test of whether the model prefers its own answer over another prompt's sat at 50%, exactly chance, and unmoved. On a behavioural spot-check it answered 9 of 12 correctly. A depth-4 loop takes 0.69 s on an RX 5700 XT, roughly 74 ms per marginal step.
The verdict, and the test that lied
Our validation harness printed PASS on these numbers. It should not have.
The check required the transplant cost to exceed 0.05. Transplant came in at 0.055, clearing the bar by five thousandths of a nat, on 12 prompts. That is a false positive, and it is exactly the failure mode we have written about before: a metric that measures something other than its name, briefly supporting a conclusion it cannot carry.
The check was tightened, transplant must now cost at least 25% of the strongest structural intervention and hurt at least 75% of prompts, and the run was re-scored. Step 82,697 reads PARTIAL. Against the stated release bar, it does not clear it. We are publishing the model anyway, labelled accurately, because a half-demonstrated mechanism with its own falsification attached is worth more to anyone working on this than silence would be.
What it cannot do
The latents are not specific to your question. A transplant cost of +0.055 nats says the model learned the form of an answer, "a small integer", rather than the answer to the problem in front of it. The tier breakdown makes it concrete: on arithmetic, gold tokens rank 7 to 23; on logic they rank 362, 732, 2,065 and 5,791, and at those ranks another prompt's gold answer beats their own. The model has learned to think. It has not learned to think about the question.
Adaptive depth does not work. The halt head emits λ≈0.31 and then ≈0.24, an equilibrium of its own regularisers rather than a response to the problem, so dynamic stepping is a fixed five steps for every prompt regardless of difficulty. Our reading is that PonderNet was starved by the specificity gap rather than broken by it, which is why the next move is to fix specificity first and leave the halt head alone.
No safety training of any kind. Refusal behaviour is inconsistent. It can and does produce content it was asked to decline, and its disclaimers are not a safety layer. Outputs are frequently wrong: it fails arithmetic, logic and instruction-following in the ways documented above. Do not rely on it for anything consequential, and do not put it in front of users. Research and evaluation only, which is also all the licence permits.
Licence and disclaimer
Derived from Qwen/Qwen2.5-3B-Instruct,
released under the Qwen RESEARCH LICENSE AGREEMENT and redistributed in
full as LICENSE.Qwen alongside the weights, as that licence requires.
Read the full licence text.
Built with Qwen.
Non-commercial use only. The Qwen research licence grants rights "FOR NON-COMMERCIAL PURPOSES ONLY", meaning research or evaluation. Hyper v1 inherits that restriction. Commercial use requires a licence from Alibaba Cloud.
This model is provided "AS IS", without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, accuracy, and non-infringement. Cymela accepts no liability for anything this model produces or for any use made of it. In no event shall Cymela be liable for any claim, damages, loss, or other liability, whether in contract, tort, or otherwise, arising from the model, its outputs, or their use. You are solely responsible for what you do with this model and for anything it generates on your behalf. Using or distributing it means you accept these terms and the Qwen Research License. If you cannot, do not use it.
This section states our position; it is not legal advice, and the effect of disclaimers varies by jurisdiction.
Modifications to the base model: continued fine-tuning of all weights on a latent
chain-of-thought objective; three special tokens added with a resized embedding table;
and three new modules, the normaliser, the carry gate and the halting head, shipped
as a separate weights file alongside the backbone. Those modules and
hyper_chat.py are original work, © 2026 Cymela.