Does cross-model code review improve AI-generated code?

Published evidence, one caveat the headline misses, and what one week of running a cross-family review gate in production actually looks like.

Short answer

Yes — in the strongest measured direction. In a controlled 116-task study, Claude reviewing Codex drafts lifted pass rates from 71.6% to 89.7%, beating same-family self-review (84.5%). But the reverse direction hurt: Codex reviewing Claude drafts dropped 91.4% to 82.8%. Cross-model review is not a symmetric switch — put your strongest analyst in the reviewer seat.

The published evidence

Zuodong Xiang, Yike Zhang, YueMing Zhang, Hailu Xu ran a controlled study on 116 recent hard and medium LiveCodeBench tasks across six conditions — both models solo, both self-review pairings, and both cross-model pairings. The reviewer saw the problem and the writer’s draft but could not run tests.

Codex drafts, no review71.6% passbaseline
Codex drafts + Codex self-review84.5% passhelps somewhat
Codex drafts + Claude reviews89.7% passbest for this direction
Claude drafts, no review91.4% passbaseline
Claude drafts + Claude self-review91.4% passno change
Claude drafts + Codex reviews82.8% passhurts

Source: Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa? — arXiv:2607.21656 · Agentic SE workshop @ KDD 2026.

The caveat that matters for design

The benefit is directional. “Use a different model family” alone is not the finding — the finding is “use the strongest analyst as the reviewer, and do not mirror the builder.” Same-family self-review helped only modestly in one direction (84.5%) and not at all in the other (91.4% → 91.4%). This is the single design rule the lab’s gate is built on: the reviewer is never from the builder’s family, and the strongest available model holds the reviewer seat.

What we run

Turbo Rig’s gate operationalizes this finding: every PR is reviewed read-only by a different model family, the gate fails closed (no parseable verdict = REVISE), and only humans merge. That last part — human merge authority — is deliberate: the study measures review quality, not autonomous merging.

One week of field telemetry

Operational counts from the lab’s own gate for the week of 2026-09-28 — this is field telemetry, not a controlled study. We publish it because it is traceable to its writer (the rig’s stats pipeline), and it shows what the gate does to a real workload: three of four gate verdicts come back REVISE.

Gate runs (reviews executed)
839
REVISE verdicts (all gate runs)
641 (76% of verdicts)
PRs merged after review
66 across 7+ repos
Builder sessions in the rig
1,797

A 76% REVISE rate across all gate runs is consistent with the paper’s premise — an independent reviewer family catches a lot that the builder missed. It is not evidence of the 18.1 percentage-point lift; only a controlled comparison can show that, and we have not run one.

Limitations

  • The controlled result comes from LiveCodeBench tasks — competitive-programming-shaped, not general software engineering on real repositories.
  • The study covers two model families as of July 2026; the landscape moves monthly.
  • Our telemetry is one week, one rig, self-reported by the pipeline that runs the gate; per-category defect breakdowns are not published yet.
  • We did not run the 116-task study. Every number in the table above belongs to its authors; our contribution is running the finding in production and reporting what falls out.

Cite this

Patman, Marcus, and Adventure Wave Labs. "Does Cross-Model Code Review Improve AI-Generated Code?" Adventure Wave Labs, 8 Oct 2026, https://www.adventurewavelabs.space/research/cross-model-code-review.

Xiang, Zuodong, et al. "Cross-Model LLM Code Review: Should You Use Claude to Review Codex or Vice Versa?" arXiv:2607.21656, Agentic SE @ KDD 2026, https://arxiv.org/abs/2607.21656.

Part of the lab’s project pages · the rig that runs this gate: Turbo Rig (private beta).