A merged 61-layer model emitted nothing but question marks. The cause was not the seam, not the wrapper, and not numerical instability. It was 36 routers full of zeros.

We built a merged model by splicing a dense model's layers onto a sparse model's layers. It loaded, it ran, and it produced pure garbage, page after page of ????. It stayed that way for months, through several wrong theories.

Streaming the headers of all 25 shards and loading only the suspects settled it. In the sparse architecture, each layer has a small router that decides which experts handle each token. At every one of the 36 layers inherited from the dense model, that router was entirely zeros.

The failure chain is mechanical once you see it. Zero routers mean every expert scores identically, so the selection is arbitrary and each chosen expert is weighted by a fraction meant for a much larger pool. The dense model's actual computation, the thing those 36 layers are, was effectively deleted. 36 destroyed layers ran before the sparse half ever saw the data. By then the internal state was noise, the output distribution was flat across the whole vocabulary, sampling drew essentially random token IDs, those decoded to invalid text, and the terminal drew each one as ?.

Three things we had suspected were exonerated outright: the seam between the two halves was fine, the numerical-instability fix from an earlier bug was working correctly, and the orchestration wrapper was a faithful shell around a broken model. We had been debugging all three for weeks.

The fix needs no custom architecture. The capability was already there in the stock configuration; the original merge simply never used it.

← All research entries