Skip to content

Guide

AI review against human review: the missing baseline

The argument is always conducted against an imagined reviewer who is consistent. NeurIPS has twice measured what two real committees do with the same papers, and about half of what one accepts, the other rejects.

Published 11 August 2026

Weighing whether a machine read is worth anything. The conclusion below is that it is worth one narrow thing: a rehearsal, run by you, while the paper can still change. The list of what a rehearsal has to catch is the pre-submission checklist

What this page does not claim

The baseline nobody quotes

Two committees, the same papers, different answers

NeurIPS duplicated a slice of its submissions and sent each copy through two independent committees, once in 2014 and again in 2021 after the conference had grown roughly tenfold. The results barely moved.

Results of the NeurIPS consistency experiments
 20142021
Papers sent to both committees166882
Outcomes that disagreed43, or 25.9%203, or 23.0%
Accepted by one, rejected by the other49.5% of the first committee's accepts50.6% of the first committee's accepts

From the NeurIPS programme chairs' own report on the 2021 experiment, which also restates the 2014 figures. Submission volume rose by more than fivefold between the two runs and the disagreement did not follow it.

This is not a scandal, and reading it as one misses the point

Reviewers are not failing at a task with a right answer. Near the threshold there is genuine disagreement about what matters, and two competent committees weighing the same real tradeoffs differently is what that looks like from outside. What the number does establish is a ceiling: no process, human or otherwise, is going to reproduce a set of decisions that the process itself does not reproduce.

It also tells you what your own rejection means

If half of one committee's accepts would have been rejected by another, then a rejection at that level is not a measurement of your paper's quality with any precision. That is the practical use of this number, and it belongs to authors rather than to the debate about machines.

What machines score

The published numbers, including ours

Three systems that have been measured against real conference decisions. All three are being scored on the same task: given the paper, predict what the committee did.

Published AUC figures for systems predicting conference decisions
SystemMeasured onAUCSource
phdflowICLR 2025, 297 papers0.91Recomputed from our result data
phdflowAll three sets, 697 papers0.87Recomputed from our result data
Stanford Agentic ReviewerICLR 2025, their own sample0.75Kukoyi et al., 2026, preprint
The AI ScientistICLR 2022, a different set0.65Lu et al., 2024

Ours is recomputed from a frozen snapshot of a finished evaluation: 0.9128 on ICLR 2025 at 95.96% accuracy, and 0.8747 across all 697 at 91.68%. The other two are quoted from their authors and we have not reproduced either, so they are not evidence about those systems so much as what those systems report. The full matrix, venue by venue, is on the architecture page.

Read this before comparing the two tables

Four reasons those numbers do not go next to each other

The arithmetic invites one sentence and the sentence is wrong. Here is what separates the two measurements, in order of how much it matters.

01

They measure different quantities

The committees were making decisions. We are predicting one that has already been made and is not going to change.

A second committee has to weigh the tradeoffs again from nothing, and any disagreement it has with the first counts against consistency. A predictor is aiming at a fixed, published target and can hit it by learning what such targets look like. These two tasks can be arbitrarily far apart in difficulty, and the easier one is ours.

02

Our labels may not have been secret

ICLR submissions, reviews and decisions are public on OpenReview, and our best number is on ICLR.

This is the strongest objection to our own headline figure and we cannot rule it out. A language model asked about a paper whose fate has been published may be recalling the outcome rather than reading the manuscript, and nothing in our evaluation separates those two routes to the right answer. The pattern in our own results is at least consistent with the concern: the newest and most-discussed set, ICLR 2025, is where we score highest at 95.96%, and the oldest, ICLR 2017, is where we score lowest at 85.22%.

03

The target is contaminated with things that are not the science

Acceptance correlates with author prestige, affiliation, topic fashion and competent formatting, and several of those are visible in the PDF.

A system can score well on this task by learning the sociology of a venue without ever engaging with the argument of a paper. Agreeing with the committee is therefore evidence about agreeing with the committee, and not about reviewing. This objection applies equally to all three rows in the table above.

04

Nobody has compared the writing

A review is a document that tells an author what is wrong. None of these numbers looks at one.

We have never compared what phdflow writes against what reviewers wrote, and we hold no measurement of that kind. The comparison that would actually settle the question in the title of this page has not been made by us, and we are not aware of a version of it that is both large and independent. The rest of what this kind of scoring cannot do is set out separately.

Where that leaves it

The useful finding is about the ceiling, not the winner

Nothing above says a machine reviews better than a committee, and we do not think the question is currently answerable. What the two tables do together is set a scale for the argument, which the argument has been missing.

Any tool that reviewed papers perfectly would still disagree with about half of one committee's accepts, because that is how much two committees disagree with each other. So a demand that AI match human review is ambiguous: matched against which committee, on a decision that a second one would have reversed half the time? And in the other direction, a system that agrees with published outcomes at 96% has not thereby proved it reviews well. It has proved it predicts well, on papers whose outcome was public.

The honest position, and the one this whole site is built on, is that this class of tool is worth having for something narrower than the debate assumes: a rehearsal of the reading your paper is about to get, run by you, on your own work, while you can still change it. That claim does not need to beat a human reviewer to be true.

If the paper has not gone out yet, the failures that end a submission before a reviewer is assigned are in the pre-submission checklist.

phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.