We fitted a map between two real models, beat the do-nothing baseline decisively on reconstruction error, and it made behavior worse. This is our own data arguing against our own instinct.
The intuitive way to connect two models is to fit a map between their internal spaces and check how well it reconstructs the target. Lower error, better map. It is the obvious metric and it is easy to compute.
On a real pair of models we fitted such a map. It beat the trivial do-nothing baseline by 92% on geometric reconstruction error, a decisive, unambiguous win on the metric. Then we measured what it did to actual output quality, and it performed worse than doing nothing at all.
Geometric fit and behavioral effect are close to uncorrelated. We had read that in the literature. Reproducing it on our own harness, on our own data, with a result that contradicted what we expected, is what actually changed how we work.
It is the single clearest argument for why our validation order has to be what it is: a map that fits beautifully and communicates nothing is the expected outcome, not an edge case. Any pipeline that stops at reconstruction error will confidently ship a channel that carries no information.
We went looking for this deliberately. A test suite that only exercises the path where things work measures nothing.