We had a result ready to publish: the orchestrator knows when to stay silent, and our decision rule was throwing that knowledge away. Running it on the model that actually ships inverted the diagnosis, on a difference of five rows out of fifteen, indistinguishable from noise. So we built a bigger evaluation and ran it again. The finding was an artifact of fifteen rows, and the thing that needed fixing was never the model.
We had a result we were ready to publish. We tested it on a second model first, because the test was cheap. It did not survive, and the reason it did not survive is more useful than the result would have been.
The orchestrator has three moves: stay SILENT, DISPATCH a research arm, or NOTE a correction into the worker's stream. Silence is the default and the one we have never got right. Our episode ids encode which scenario generated them, and one scenario is entirely gold-SILENCE, cases where saying nothing is correct, while three others are entirely gold-not-silent. Those are the two poles of restraint, and until this week nobody had scored them separately.
On the 8B merged model the answer looked clean
At argmax it gets 13 of 15 restraint cases right, well above guessing. It pays for that everywhere else: on the never-silent scenario it scores 45%.
Sweeping the decision threshold showed the two scenarios are not simply traded against each other. If there were no real signal and we were only moving a global "how often do I stay quiet" dial, the two scores would sum to roughly a constant. They do not. The sum peaks in the middle.
silence offset restraint never-silent sum 0.0 (argmax) 87% 45% 132 -3.0 80% 73% 153 -4.0 73% 82% 155 -6.0 27% 82% 108
A peak in the middle means information the decision rule was discarding. The reading: the model knew when to hold back and argmax threw it away. Recalibrating (three numbers, swept offline, no retraining) takes the never-silent scenario from 45% to 91% and overall accuracy from 76.6% to 84.8%.
That was the entry we were ready to publish.
Then we ran it on the model that actually ships
A 7B with a dedicated adapter. 33 minutes on two T4s, identical method, identical held-out rows.
at argmax 8B merged 7B shipping overall 76.6% 85.6% restraint scenario 87% (13/15) 53% (8/15) never-silent scenario 45% 73% predictions on the 15 restraint SILENT 13 DISPATCH 7 rows DISPATCH 2 SILENT 8
The 7B is the better model overall and already close to its own ceiling, 85.6% at argmax against 88.9% at the best threshold anywhere. But it does not appear to recognise restraint cases at all: it dispatches on nearly half of them. We swept the entire two-dimensional threshold space and no operating point lifts its restraint score above 60%, and reaching even that collapses overall accuracy to 67%.
So the diagnosis inverts. On the 8B, restraint is known and mis-read. On the 7B, the threshold is close to right and the signal is weak. The conclusion did not generalise one model sideways.
The part that matters more: we cannot actually tell
That comparison is 13 of 15 against 8 of 15. Five rows.
8B restraint above chance p = 3.1e-05 clearly yes
7B restraint above chance p = 0.088 not clearly
8B vs 7B difference p = 0.109 not significant
(Fisher exact, two-tailed)
The difference the entire generalisation claim rests on is not statistically distinguishable from noise. Both readings survive: a real difference between the models, or a fifteen-row coin landing differently twice.
We could have published the first result on its own. It was clean, it was surprising, it argued against our own prior work, and nothing in it was wrong. It is still a correct statement about that model. What makes it unpublishable as a general claim is a test we only ran because it was cheap.
The failure mode is worth naming, because it is not the one we were guarding against. The withdrawn draft stated its sample size correctly. It said n=15, it said one row moves the number by 6.7 points, and it said so in its own section rather than a footnote. And it still carried a general claim on top of that. Stating a caveat is not the same as respecting it.
What we are changing
Not the training. The evaluation. The restraint scenario is 15 rows of a 389-row held-out slice. Every restraint experiment we run against it will keep returning p of about 0.1 and licensing whichever conclusion the reader brought with them. Enlarging that slice is now a precondition for any further claim in this area, ours included.
We enlarged it, and it settled
80 fresh rows, 50 where silence is correct, 30 where acting is, written in the corpus format and scored on all three models in a single 22-minute session. Identical rows for every model.
at argmax silence (n=50) act (n=30) overall 7B shipping 80% 90% 83.8% 3B dedicated 66% 80% 71.2% 8B merged 88% 63% 78.8%
The 7B can do restraint. 40 of 50. The 8 of 15 that started all this was an artifact of 15 rows, which is what the section above suspected and could not show.
The gap we were busy explaining does not exist. 88% against 80% on the same rows is Fisher exact p = 0.414. There was nothing there to explain.
One real difference did survive, and it is the reverse of the original story. The 8B stays silent where it should act: 19 of 30 against the 7B's 27 of 30, p = 0.030. Not restraint the 7B lacks. Restraint the 8B over-applies.
And the operating point we had already chosen makes things worse. The offset selected on the old rows costs five points of overall accuracy on the new ones, 78.8% down to 73.8%. It had fitted the distribution it was tuned against. That is measured inside our own evaluations rather than predicted about live behaviour, and it is the most useful number in the run.
The edge of the result
The new rows have a different author from the old 15, so old-against-new difficulty is not controlled. The 7B's improvement from 8 of 15 to 40 of 50 sits at p = 0.051, right on the line. Every cross-model comparison above uses identical rows and is immune to this. The old-against-new comparison is not, and we are not going to pretend otherwise.
One thing argues against the obvious objection. If the new rows were simply easier, both models should have risen. Run the same comparison on the 8B, 13 of 15 old against 44 of 50 new, and it returns p = 1.000. Identical performance. Only one model moved. That does not close the door, since rows could be differentially easier for the 7B in particular, but the version of the objection that says the new set is softer predicts a change in both models, and there is not one.
What still stands, and what does not
"Mis-calibrated rather than missing" survives for the 8B and only for the 8B, and it is milder than it looked, the sweep now buys 10 points of summed accuracy where it appeared to buy 23. For the 7B it buys exactly nothing. Sweeping the silence offset across its entire range gains zero, because the model already ships at its own optimum.
The 8B sweep's failure to conserve the sum is still far outside what a global prior shift could produce. And recalibration is still cheap: three numbers swept offline against logits already recorded, against 95 minutes of GPU per data point back when we did this by editing the training corpus. Cheap is not the same as safe, which is exactly the point the overfitted operating point makes.
The arc here is a finding, a failed replication, an admission that our own evaluation could not settle it, a larger evaluation built in a day, and an answer. At no point was the fix to the model or to the training. It was to the measurement. The recalibration figures throughout remain offline projections, re-scored from recorded logits rather than observed in a live run, but the five points the chosen operating point costs on fresh rows is measured, and that is the number we would act on.
The scoring script, the archived deploy reports, the recorded action logits and the enlarged restraint slice are kept with the project and reproduce every figure above.