Before fitting anything, check whether two models chop text into the same pieces. It costs about an hour and it predicts almost everything downstream.

Connecting two models is expensive to attempt and expensive to evaluate. So the question that matters early is not does this work but is this pair worth the attempt at all.

It turns out there is a very cheap answer. Across 23 published model pairs, the rate at which two models split text into identical tokens predicts downstream transfer quality at r = 0.898. Every documented cross-family success sat above roughly 0.7 agreement; every failure sat at or below a quarter. Architecture differences and training differences mattered far less.

So this runs first, before any fitting, on any candidate pair. It costs about an hour. A pair that fails it is not a pair we spend a week on, and picking two models that already share a tokenizer family removes the single largest confound before the real research question is even asked.

Ordering the pipeline so the hour-long tests can kill the week-long ones is most of what makes a small team's compute budget go far.

← All research entries