Identifier
Cymela/hyper-3b-latent
Parameters
3.10B · 3.085B backbone, 13.6M latent modules
Base model
Qwen2.5-3B-Instruct · 36 layers, hidden 2048, vocab 151,668
Checkpoint
Step 82,697 · training complete
Status
Research preview · not for deployment
Licence
Qwen RESEARCH LICENSE AGREEMENT · non-commercial

Read the release verdict before the numbers: against our own stated bar, this checkpoint reads PARTIAL. It does not clear it. The mechanism works; the thing it was supposed to demonstrate is only half demonstrated, and we would rather say so here than let you find out after downloading 6 GB.

Architecture

Hyper recycles its final hidden state back into inputs_embeds for K recurrent steps before emitting any token, continuous latent chain-of-thought, in the manner of Coconut. Three small added modules constitute the entire mechanism: one that normalises the latent state, one that gates how much of the new state is carried forward, and a PonderNet-style head that decides when to stop, totalling 13.6M parameters, about 27 MB.

The backbone is a fully fine-tuned Qwen2.5-3B-Instruct, 36 layers, hidden size 2048, vocabulary 151,668 after adding <bot>, <eot> and <step>. The halt head was trained to choose K adaptively. In this release it does not; see limitations.

On the parameter count: summing the stored tensors gives 3,409,648,641, and that number is wrong. config.json sets tie_word_embeddings, and lm_head.weight ships as a bitwise-identical copy of model.embed_tokens.weight, so a 310.6M table is counted twice. The unique parameter count is 3,099,032,577. It cross-checks: Qwen2.5-3B-Instruct is 3,085,938,688, our vocabulary is 268 rows smaller (−548,864), giving a 3,085,389,824 backbone plus 13,642,753 of latent modules. Two derivations, same figure. Anywhere you see 3.4B for this model, including on this site until 16 August 2026. That is the double-counted number.

Training

Final checkpoint is step 82,697, trained on a Kaggle TPU v5e-8. The final leg ran from 79,137 to 82,697, 3,560 steps across three sessions of roughly eight hours, at depths of k=2 to k=4.

Being precise about what that means, because it is the number most likely to be misread: for the overwhelming majority of those 82,697 steps the model trained at an effective latent depth of one. A configuration bug held it there, the finding is written up in We trained 79,000 steps at a latent depth of one. We can firmly attribute 3,560 steps to the genuine depth-2-to-4 regime. The exact boundary before that is not precisely established, and we are not going to round it. This is not 82,697 steps of latent training.

We are not publishing the training corpus. Its provenance and licensing have not been verified to a standard that would let us assert anything about it in public, and an unverified claim is worse than an absent one.

Evaluation

Every causal number below is N=12 prompts. That is a small sample and it belongs next to the results rather than in a footnote. Treat these as directional.

Mean cross-entropy on the gold answer, by latent depth, for this checkpoint against an earlier one (step 80,189). Lower is better.

Checkpointk=0k=1k=2k=4k=8
82,69712.11210.1108.9678.55510.099
80,18910.0289.5838.9219.84510.795
step 82,697 · finalstep 80,189 · earlier
Mean cross-entropy on the gold answer, by latent depth At step 82,697 the mean cross-entropy falls from 12.112 nats at depth 0 to 8.555 at depth 4, then rises to 10.099 at depth 8. At the earlier step 80,189 it falls from 10.028 to a shallower best of 8.921 at depth 2 and rises thereafter. N=12 prompts. 8 9 10 11 12 nats lower is better 12.112 10.110 8.967 8.555 10.099 10.028 9.583 8.921 9.845 10.795 k=0 k=1 k=2 k=4 k=8 latent depth, spaced evenly by setting rather than by value
Thinking is worth 3.557 nats at the final checkpoint against 1.107 at the earlier one, and the best depth moved from two steps to four. Both curves turn back up: by eight steps the model does worse than if it had not thought at all. N=12 prompts, so read the shape rather than the third decimal.

Thinking is worth −3.557 nats from k=0 to k=4, up from −1.107 at the earlier checkpoint, and the optimum moved from k=2 to k=4. Past that it falls off fast: by k=8 the model does worse than if it had not thought at all.

Interventions at k=4, measured as added cross-entropy. Shuffling a problem's own latent steps costs +1.389 and hurts 12 of 12 prompts. Replacing them with noise costs +1.125 and hurts 11 of 12. Transplanting a different problem's latents costs +0.055 and hurts 8 of 12, close enough to a coin flip to mean very little. A replay control reproduced the clean run bit-exactly, so the harness itself is sound.

Cost of interfering with the latent steps, at depth 4 Shuffling a problem's own latent steps adds 1.389 nats and hurts 12 of 12 prompts. Replacing them with noise adds 1.125 and hurts 11 of 12. Transplanting a different problem's latent steps adds 0.055 and hurts 8 of 12, close to chance. 0 0.5 1 1.5 added nats +1.389 hurts 12 of 12 +1.125 hurts 11 of 12 +0.055 hurts 8 of 12 shuffle its own steps replace with noise another problem's steps intervention at latent depth 4
The first two bars say the latent steps carry something the answer uses, and that their order matters. The third is the one that limits the claim: swapping in a different problem's steps costs almost nothing and hurts 8 of 12 prompts, close enough to a coin flip to mean very little. Whatever the steps carry, it is largely not the identity of the question.

A separate probe found the gold token's median rank improved from 8,303 to 23 out of 151,668, against a chance rank of 75,834. But the test of whether the model prefers its own answer over another prompt's sat at 50%, exactly chance, and unmoved. On a behavioural spot-check it answered 9 of 12 correctly. A depth-4 loop takes 0.69 s on an RX 5700 XT, roughly 74 ms per marginal step.

The verdict, and the test that lied

Our validation harness printed PASS on these numbers. It should not have.

The check required the transplant cost to exceed 0.05. Transplant came in at 0.055, clearing the bar by five thousandths of a nat, on 12 prompts. That is a false positive, and it is exactly the failure mode we have written about before: a metric that measures something other than its name, briefly supporting a conclusion it cannot carry.

The check was tightened, transplant must now cost at least 25% of the strongest structural intervention and hurt at least 75% of prompts, and the run was re-scored. Step 82,697 reads PARTIAL. Against the stated release bar, it does not clear it. We are publishing the model anyway, labelled accurately, because a half-demonstrated mechanism with its own falsification attached is worth more to anyone working on this than silence would be.

What it cannot do

The latents are not specific to your question. A transplant cost of +0.055 nats says the model learned the form of an answer, "a small integer", rather than the answer to the problem in front of it. The tier breakdown makes it concrete: on arithmetic, gold tokens rank 7 to 23; on logic they rank 362, 732, 2,065 and 5,791, and at those ranks another prompt's gold answer beats their own. The model has learned to think. It has not learned to think about the question.

Adaptive depth does not work. The halt head emits λ≈0.31 and then ≈0.24, an equilibrium of its own regularisers rather than a response to the problem, so dynamic stepping is a fixed five steps for every prompt regardless of difficulty. Our reading is that PonderNet was starved by the specificity gap rather than broken by it, which is why the next move is to fix specificity first and leave the halt head alone.

No safety training of any kind. Refusal behaviour is inconsistent. It can and does produce content it was asked to decline, and its disclaimers are not a safety layer. Outputs are frequently wrong: it fails arithmetic, logic and instruction-following in the ways documented above. Do not rely on it for anything consequential, and do not put it in front of users. Research and evaluation only, which is also all the licence permits.

Licence and disclaimer

Derived from Qwen/Qwen2.5-3B-Instruct, released under the Qwen RESEARCH LICENSE AGREEMENT and redistributed in full as LICENSE.Qwen alongside the weights, as that licence requires. Read the full licence text. Built with Qwen.

Non-commercial use only. The Qwen research licence grants rights "FOR NON-COMMERCIAL PURPOSES ONLY", meaning research or evaluation. Hyper v1 inherits that restriction. Commercial use requires a licence from Alibaba Cloud.

This model is provided "AS IS", without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, accuracy, and non-infringement. Cymela accepts no liability for anything this model produces or for any use made of it. In no event shall Cymela be liable for any claim, damages, loss, or other liability, whether in contract, tort, or otherwise, arising from the model, its outputs, or their use. You are solely responsible for what you do with this model and for anything it generates on your behalf. Using or distributing it means you accept these terms and the Qwen Research License. If you cannot, do not use it.

This section states our position; it is not legal advice, and the effect of disclaimers varies by jurisdiction.

Modifications to the base model: continued fine-tuning of all weights on a latent chain-of-thought objective; three special tokens added with a resized embedding table; and three new modules, the normaliser, the carry gate and the halting head, shipped as a separate weights file alongside the backbone. Those modules and hyper_chat.py are original work, © 2026 Cymela.

Weights and model card on Hugging Face

← All Cymela models