← back to notes

cross-model review is the default

May 2026 · Prasith Govin

Before code ships here, a different model family reviews it. Claude writes, Codex reads it and returns a verdict. One pass, and the reviewer has no authorship stake in what it's grading. That asymmetry is the point.

I use different families on purpose. Two instances of the same model share the same blind spots. Ask a model to review its own work and it will confidently miss the same things it missed writing it, because the gap is in the training, not the effort. A model from another family has different priors, so it catches what the author's priors hid.

One round only. I looked at multi-round setups where two models argue toward consensus, and on cost-adjusted quality, plain review beat debate cleanly. One writer, one reviewer, one round, a structured verdict. Extra rounds mostly added tokens. So we kept review and didn't grow it into an argument that never settles.

It catches real problems well beyond style nits. Our task board has a whole "In Review" state that exists only because that's where cross-model review and the accept-or-reject happen. One night it caught a genuine P1 on a Codex cross-review that would have shipped otherwise. That's the bar I hold it to. If cross-review only ever flagged formatting, I'd drop it. It earns its slot by catching what the author was too close to see.

The cost is honest and small. One extra state on the board, one more model in the loop, a few minutes per change. For that I get a second set of priors on everything before it goes out, and a habit that treats "I wrote it and it looks right to me" as the least trustworthy signal in the building.

So the default here is blunt: nothing I care about ships on one model's self-assessment. The author, human or agent, is the worst judge of whether the author got it right. A reviewer with different training and no stake in the diff beats one more pass of the writer checking their own work.