The dispatch pattern ran live for the first time, small models, a synthetic training set under 7,000 examples, and it did what it was designed to do: catch a stale answer before it shipped, and stay out of the way twice when nothing needed catching.
This is not the target-scale system, and it is not the vector-level Bridge described elsewhere on this site. Both roles here are ordinary chain-of-thought models, not Neuralese ones. We don't yet have two smaller models we'd call reliable at continuous latent reasoning to pair for this specific test, so the dispatch pattern was validated on the reasoning we could get today rather than waited on. The orchestrator's decision reaches the generator as an inserted note rather than a latent correction, built specifically to check the pattern itself before spending compute on the harder problem: when to dispatch, when to stay silent, and what happens when a tool the orchestrator calls breaks mid-flight.
Both roles were trained on a synthetic corpus of a little over 6,000 examples, generated in several independent passes across models from more than one lab and validated before pooling, specifically so no single model's blind spots could shape the whole dataset. Both roles ran on small, openly available models, not the target pair, the question here was whether the pattern is learnable at all, not whether it's good enough to ship.
It worked on the first real test. Asked a question whose answer had moved on since training, the generator produced a stale fact, the wrong "latest stable version" of something that has since shipped several releases. The orchestrator checked, found a live search result that disagreed, and injected a correction. The generator used it and gave the current answer. That's the entire thesis in one exchange: a generator that keeps producing, and a second model that only intervenes when it actually has something to add.
Two other tasks tested the opposite skill, knowing when not to. Asked for the population of a metro area and the visibility of a landmark from orbit, the generator answered correctly both times, and the orchestrator checked and said nothing. Silence has to be as learnable as intervention, or the system pesters a correct answer as often as it fixes a wrong one; both fired at the right moments here.
A third task broke a tool mid-run, the orchestrator's calculator crashed, and the system did not propagate the failure. A retry loop caught it, the generator finished the arithmetic on its own, and the final answer was correct anyway. A broken tool produced an undamaged answer.
Three real problems came out of the same run, and the one that matters is the first: on one task, the generator wrote a research note into its own output that the orchestrator never sent. It had learned the shape of a note well enough to fabricate one. An orchestrator that only intervenes occasionally is exactly the condition under which a generator can learn to hallucinate its intervention instead of waiting for the real thing, and that's a failure mode worth taking seriously before it's trained into anything larger. The other two were plumbing: the calculator tool rejected comma-formatted numbers, and the code-execution tool sometimes graded an unfinished draft as an error. Both are cheap fixes; the first one is the one we're actually worried about.
This validates that the dispatch pattern is learnable, at small scale, through a text-level stand-in on ordinary chain-of-thought models. It says nothing yet about the vector-level Bridge, and nothing about the target model pair. The move to Neuralese, the dispatch happening entirely in latent space, between two models that actually reason continuously rather than in text, waits on having a reliable pair of latent-reasoning models to run it on. That's a separate, harder problem, still open.