Does cross-model code review improve AI-generated code?
Published evidence, one caveat the headline misses, and what one week of running a cross-family review gate in production actually looks like.
Short answer
Yes — in the strongest measured direction. In a controlled 116-task study, Claude reviewing Codex drafts lifted pass rates from 71.6% to 89.7%, beating same-family self-review (84.5%). But the reverse direction hurt: Codex reviewing Claude drafts dropped 91.4% to 82.8%. Cross-model review is not a symmetric switch — put your strongest analyst in the reviewer seat.
The published evidence
Zuodong Xiang, Yike Zhang, YueMing Zhang, Hailu Xu ran a controlled study on 116 recent hard and medium LiveCodeBench tasks across six conditions — both models solo, both self-review pairings, and both cross-model pairings. The reviewer saw the problem and the writer’s draft but could not run tests.
| Codex drafts, no review | 71.6% pass | baseline |
| Codex drafts + Codex self-review | 84.5% pass | helps somewhat |
| Codex drafts + Claude reviews | 89.7% pass | best for this direction |
| Claude drafts, no review | 91.4% pass | baseline |
| Claude drafts + Claude self-review | 91.4% pass | no change |
| Claude drafts + Codex reviews | 82.8% pass | hurts |
Source: Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa? — arXiv:2607.21656 · Agentic SE workshop @ KDD 2026.
The caveat that matters for design
The benefit is directional. “Use a different model family” alone is not the finding — the finding is “use the strongest analyst as the reviewer, and do not mirror the builder.” Same-family self-review helped only modestly in one direction (84.5%) and not at all in the other (91.4% → 91.4%). This is the single design rule the lab’s gate is built on: the reviewer is never from the builder’s family, and the strongest available model holds the reviewer seat.
What we run
Turbo Rig’s gate operationalizes this finding: every PR is reviewed read-only by a different model family, the gate fails closed (no parseable verdict = REVISE), and only humans merge. That last part — human merge authority — is deliberate: the study measures review quality, not autonomous merging.
One week of field telemetry
Operational counts from the lab’s own gate for the week of 2026-09-28 — this is field telemetry, not a controlled study. We publish it because it is traceable to its writer (the rig’s stats pipeline), and it shows what the gate does to a real workload: three of four gate verdicts come back REVISE.
- Gate runs (reviews executed)
- 839
- REVISE verdicts (all gate runs)
- 641 (76% of verdicts)
- PRs merged after review
- 66 across 7+ repos
- Builder sessions in the rig
- 1,797
A 76% REVISE rate across all gate runs is consistent with the paper’s premise — an independent reviewer family catches a lot that the builder missed. It is not evidence of the 18.1 percentage-point lift; only a controlled comparison can show that, and we have not run one.
Limitations
- The controlled result comes from LiveCodeBench tasks — competitive-programming-shaped, not general software engineering on real repositories.
- The study covers two model families as of July 2026; the landscape moves monthly.
- Our telemetry is one week, one rig, self-reported by the pipeline that runs the gate; per-category defect breakdowns are not published yet.
- We did not run the 116-task study. Every number in the table above belongs to its authors; our contribution is running the finding in production and reporting what falls out.
Cite this
Patman, Marcus, and Adventure Wave Labs. "Does Cross-Model Code Review Improve AI-Generated Code?" Adventure Wave Labs, 8 Oct 2026, https://www.adventurewavelabs.space/research/cross-model-code-review.
Xiang, Zuodong, et al. "Cross-Model LLM Code Review: Should You Use Claude to Review Codex or Vice Versa?" arXiv:2607.21656, Agentic SE @ KDD 2026, https://arxiv.org/abs/2607.21656.Part of the lab’s project pages · the rig that runs this gate: Turbo Rig (private beta).