A causal audit swapped the transmitted message for the wrong one, and most of the reported improvement survived. We rebuilt how we measure because of it.

The premise behind coupling two models below the text layer is that a richer channel carries more than words can. The literature broadly supports it. What the literature mostly does not do is check whether the receiving model actually used the content it was sent.

A July 2026 causal audit did check. It replaced the transmitted message with content drawn from a completely unrelated item, same shape, same statistics, wrong meaning, and a large share of the published improvement stayed exactly where it was. The channel was doing something. It was not communicating.

The consequence is uncomfortable and it applies to us as much as anyone: a single benchmark number measured against a no-message baseline cannot be interpreted. It cannot separate the receiver understood what it was told from the receiver was jostled and got luckier. Almost every headline result in this space is reported exactly that way, ours included if we weren't careful.

So we stopped reporting a total effect, and started decomposing it:

Content-attributable The part that vanishes when the message is swapped for a wrong one. This is the evidence.
Content-independent The part that survives a wrong message. This is the placebo.
Total effect The two added together, the number usually published on its own.

A positive total effect with a content-attributable effect of zero is a null result, no matter how large the total is.

← All research entries