Skip to content

Guide

How accurate is AI peer review: 91.68% on 697 real decisions

Across 697 papers whose outcome is public, our verdict matched the venue's own decision on 91.68% of them at AUC 0.8747, and on the 297 ICLR 2025 papers it reached 95.96% accuracy and AUC 0.9128 with one false accept in 149 rejections. That is agreement with a committee's published decision rather than a measurement of whether a paper is good, and it was measured only at ICLR and NeurIPS.

Published 19 August 2026

91.68% of 697 real decisions matched ICLR 2017, NeurIPS 2023, ICLR 2025
0.9128 AUC on ICLR 2025 rank based, n=297
1 false accept in 149 rejections ICLR 2025 confusion matrix
85.22% on our weakest set, ICLR 2017 published for that reason

The short version

The measurement

Every metric, venue by venue

The whole set and each venue inside it, with precision and recall rather than accuracy on its own. Accuracy hides which of the two mistakes a system makes, and they cost an author different things.

Accuracy, precision, recall, F1 and AUC by venue, 697 papers
SetPapersAccuracyPrecisionRecallF1AUC
ICLR 2017, via PeerRead20385.22%84.00%85.71%84.85%0.8369
NeurIPS 202319791.88%87.50%98.00%92.45%0.8666
ICLR 202529795.96%99.28%92.57%95.80%0.9128
All three sets69791.68%91.14%92.20%91.67%0.8747

Recomputed on 19 August 2026 from the frozen result snapshot of a finished evaluation run. Ground truth is each venue's own accept or reject decision, taken from the public record. Precision is how often an accept call was right; recall is how many of the venue's accepts the system found. The two conferences here do not draw the line in the same place, and we measured that gap separately: ICLR vs NeurIPS.

The oldest set is the weakest, and it stays in the table

ICLR 2017 sits at 85.22% in the same table as the 95.96%, more than ten points below it. Those papers are eight years old, they reach us through PeerRead rather than through the venue's current record, and the distribution behind them is not the one a 2026 submission comes from. A page that showed only its best year would be telling you about our editing rather than about the system.

Recall and precision pull opposite ways at the two venues

At NeurIPS 2023 the system found 98.00% of the papers the committee accepted and paid for it: precision falls to 87.50%, so more of its accept calls were wrong. At ICLR 2025 the trade runs the other way, precision 99.28% against recall 92.57%. Same system, two venues, two different mistakes, which is the reason accuracy alone does not describe it.

Where the mistakes are

The 12 papers it got wrong at ICLR 2025

One of them was a paper it waved through that ICLR rejected. The other 11 are the cell that matters if you are reading a reject verdict of your own.

How to read the ICLR 2025 confusion matrix
CellWhat it means for a paper
137 true acceptsThe paper was accepted by ICLR and the system said accept.
1 false acceptOne paper of the 149 ICLR rejected was called an accept. This is the cell a sceptic looks for, and it is the smallest number here.
11 false rejectsEleven papers ICLR accepted were called rejects. If you are holding a reject verdict, this is the cell you are arguing with.
148 true rejectsRejected by ICLR, and called a reject.
138 accept calls in total137 of them right, which is the precision of 99.28%.
159 reject calls in total148 of them right. About one in fourteen was accepted by ICLR anyway.

Across all 697 papers the same matrix reads 319 true accepts, 31 false accepts, 27 false rejects and 320 true rejects.

Confusion matrix, ICLR 2025

Confusion matrix for ICLR 2025, 297 papers
137 True accepts1 False accepts
11 False rejects148 True rejects

Against the record

The other systems that published a number

Four systems scored on the same task: given the paper, predict what the committee did. Two of them were measured on a set we also use.

Published accuracy and AUC for systems predicting conference decisions
SystemMeasured onAccuracyAUCSource
phdflowICLR 2025, 297 papers95.96%0.9128Recomputed from our result data
phdflowAll three sets, 697 papers91.68%0.8747Recomputed from our result data
Stanford Agentic ReviewerICLR 2025, their own sampleNot reported0.75Kukoyi et al., 2026, preprint
The AI ScientistICLR 2022, a different set0.65 balanced0.65Lu et al., 2024
PeerRead classifierICLR 2017, supervised baseline65.3%Not reportedKang et al., 2018

Ours is recomputed from the snapshot described below. The others are quoted from their authors and we have reproduced none of them, so those rows are evidence of what each system reports rather than evidence about the system: The AI Scientist reports balanced accuracy 0.65, F1 0.57 and AUC 0.65 on ICLR 2022, and PeerRead reports 65.3% on ICLR 2017 from a supervised classifier trained on that data, which is a different method on the same year we score 85.22% on.

The rank correlation is with a committee, not with a reviewer

On ICLR 2025 the Stanford Agentic Reviewer reports Spearman rho 0.42 against the outcome; ours on that set is 0.66. Both are correlations with a venue's aggregate decision, one number per paper produced after review and discussion, and neither is agreement with an individual reviewer. That distinction matters because the number people reach for on the human side, Bornmann's 2011 meta-analysis of inter-rater reliability at an ICC of about 0.34, measures a harder and different thing. Putting 0.66 next to 0.34 as though one beat the other would be wrong, and we are not doing it.

Provenance

How the number was produced

Nothing here is an estimate or a demo run. Every figure on this page comes out of one frozen snapshot, and this is what is in it.

How the benchmark was measured
QuestionAnswer
What is predictedAccept or reject at the venue the paper was actually sent to, plus a continuous score, which is what the AUC ranks.
What counts as correctThe venue's own published decision. No judgement of quality by us enters the label.
How many papers697: ICLR 2017 203 papers, NeurIPS 2023 197, ICLR 2025 297.
Which modelGemini 3.1 Flash Lite Preview.
What a review costsAbout 0.10 USD of model usage per paper.
When it was recomputed19 August 2026, from the frozen snapshot of the finished run, which is the source of every figure above.

Read this with the number

What 91.68% does not say

Four things the figure is silent about. Each of them is a reason someone could quote this page against us, which is why they are on it.

It predicts a decision, not the quality of a paper

The label is what a committee did, so a system that agrees with it has learned the committee. Acceptance carries things that are not the argument in the manuscript: what the field was excited about that year, how the work is packaged, how a reviewer's load ran that week. Matching a venue's decisions 91.68% of the time is a claim about those decisions and about nothing else, and decision and quality come apart exactly where an author cares most.

Only ICLR and NeurIPS, so only machine learning

All 697 papers come from two machine learning conferences. We hold no measurement in medicine, physics, economics or any humanities field, where what counts as a contribution and how review is run are both different. If your paper is not machine learning, the honest answer is that we do not know how the system behaves on it, and no figure on this page transfers to it.

A reject verdict is not a prediction about your paper

On ICLR 2025, 11 of the 159 papers the system called reject were accepted by the venue. We sell no forecast, nothing here should be read as a promise of acceptance, and no verdict from us is a reason to withdraw a submission. What a run is worth is the located issues, which you can check against your own text whether the verdict flatters you or not.

The outcomes we predict were public before we predicted them

ICLR submissions and decisions live on OpenReview, so a model asked about one of these papers may be recalling what happened rather than reading the manuscript, and nothing in the evaluation separates those two routes to a correct answer. The pattern is at least consistent with the worry: the newest and most discussed set is where we score highest. The full version of that objection, next to the human baseline it has to be read against, is in our guide on AI review against human review.

AI review against human review, and why the two numbers do not compare

And the rest of what this kind of scoring cannot do

Questions

Accuracy, asked directly

The questions people type, answered from the snapshot above and from nothing else.

How accurate is AI peer review?

Ours agreed with the venue's own decision on 91.68% of 697 papers, at AUC 0.8747. Broken out: 95.96% on 297 ICLR 2025 papers, 91.88% on 197 NeurIPS 2023 papers and 85.22% on 203 ICLR 2017 papers. There is no industry figure to set beside it, because no commercial tool we surveyed publishes an accuracy figure at all, and the published research systems report AUC between 0.65 and 0.75 on comparable tasks. All of it is agreement with a committee's decision, which is not the same thing as judging a paper well.

Can AI predict whether a paper will be accepted?

On these two conferences, better than chance and well short of certainty. On ICLR 2025 the accept calls were right 137 times out of 138, but the reject calls were wrong 11 times out of 159, so a reject verdict is not a forecast that your paper will be rejected. It is also predicting a committee's decision rather than the quality of the work, and those two come apart whenever a decision turns on what the field is excited about that year, on reviewer load, or on luck.

Is AI review as good as a human reviewer?

Unknown, and anyone who gives you a straight answer is comparing two different measurements. Our 95.96% is agreement with a decision that had already been made and published. The quantity people mean by human agreement is inter-rater reliability, which Bornmann's 2011 meta-analysis put at an ICC of about 0.34, and NeurIPS has twice found that two committees given the same papers reverse about half of each other's accepts. Those figures cannot be laid side by side, and our guide on AI review against human review sets out why in full.

What benchmark is this measured on?

697 papers with public outcomes, in three sets: 203 from ICLR 2017 through PeerRead (Kang et al., 2018), 197 from NeurIPS 2023 and 297 from ICLR 2025. Ground truth is each venue's own accept or reject decision, the model is Gemini 3.1 Flash Lite Preview at about 0.10 USD of model usage per review, and every figure was recomputed on 19 August 2026 from the frozen snapshot of that run. It is machine learning only: two conferences, no other field.

Does AI peer review work outside machine learning?

We have no evidence that it does. Every paper in the benchmark came from ICLR or NeurIPS, so the figures on this page say nothing about medicine, physics, economics or the humanities, where review norms and what counts as a contribution both differ. You can run a paper from any field through the system, but the measurement behind these numbers was not made on one, and we will not claim it transfers.

Why is the ICLR 2017 score lower than the rest?

85.22% against 95.96% on ICLR 2025, and both stay in the same table. Those papers are eight years old, they reach us through the PeerRead dataset rather than through the venue's current record, and the distribution behind them is not the one a submission written now comes from. It is also the only set where a like-for-like comparison exists: the supervised PeerRead classifier reported 65.3% on the same year.

697 of those papers were not yours

Whether the number holds on your draft is one run away, and the issues it locates are checkable against your own text whatever the verdict says.

phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.