Skip to content

Guide

Why ML conference papers get rejected

Most rejected papers are not bad papers. Across 697 real ICLR and NeurIPS decisions there is one band, half a point wide, where the outcome stops tracking the manuscript at all. Roughly one paper in five lands in it.

Published 11 August 2026

Somewhere between a draft and a decision. If the paper is still yours to change, this page is the diagnosis and the next step is a read of the file itself. If the reviews are already in, the page for that week is how to write a rebuttal

This is about conferences, not journals

6.8% of papers scoring below 6.0 were accepted 15 of 220
41.6% of papers in the 6.0 to 6.5 band were accepted 64 of 154
82.7% of papers at 6.5 and above were accepted 267 of 323
22% of all submissions land inside the band 154 of 697

The finding

There are three papers, and only one of them has a problem you can fix

Sorting 697 submissions by how they read, rather than by what happened to them, splits them into three groups that behave nothing like each other.

Acceptance rate by score band across 697 ICLR and NeurIPS decisions
Score bandPapersShare of allAcceptedAcceptance rate
Below 6.0022031.6%156.8%
6.00 to 6.5015422.1%6441.6%
6.50 and above32346.3%26782.7%

The score is ours and comes from reading the manuscript alone. Accepted and rejected are each conference's own decision. The tinted row is the band where knowing how the paper reads tells you almost nothing about what happened to it.

Below the band, rejection is about the paper

Fifteen of 220 papers got in from here, and 205 did not. If your paper reads like this, something in it is legible enough that a reader can name it, which also means you can find it and fix it. This is the only group where the advice in the next section is the whole answer.

Inside the band, the paper has already done its job

64 accepted, 90 rejected. Nothing a reading of the manuscript can recover separates the two piles. Whatever decided them was not on the page: which three reviewers it drew, what else was in the batch, how the discussion went, how much room the programme had left.

Above the band, rejection is rare and worth reading twice

56 of the 351 rejected papers scored 6.5 or higher, and the strongest of them reached 7.90. A rejection at that level is usually a specific objection from one reviewer rather than a verdict on the work, and it is the case where resubmitting with almost nothing changed is a defensible plan.

The reasons

Six sentences that appear in almost every rejection

These are the objections in the wording reviewers use, and what each one is really saying. They are craft knowledge rather than measurement: we hold the decisions, not the reviews, so nothing here is ordered by frequency.

01

The claims are broader than what was shown

The abstract says the method improves generalisation. The experiments show it improves accuracy on three datasets in one domain at one model size.

This is the cheapest rejection to earn and the cheapest to avoid, because the fix costs you nothing that matters. Nobody is asking for more experiments. They are asking the sentence to describe the experiment you ran. A claim narrowed to what you actually demonstrated is not a weaker paper, it is a paper that survives its own evidence section, and the reviewer who was going to write this sentence now has nothing to write.

Read your own abstract as an adversary once, with a pen, and mark every noun that promises a scope. Then check each one against a figure. The ones with no figure behind them are the objections you are choosing to ship.

02

The obvious baseline is missing

No comparison against the method everyone in this subfield would reach for first, or a comparison against a version of it that was not tuned.

A reviewer cannot tell an unlucky omission from a convenient one, and is not required to try. The damage is not the missing row in the table, it is what the absence implies about the rows that are there: if this one was left out, the reader now has to wonder about the others, and the whole evidence section loses the benefit of the doubt.

If a baseline genuinely does not apply, say so in the paper in one sentence with the reason. An explained absence costs nothing. An unexplained one costs the section.

03

I could not tell what is new here

The related work is thorough, the method is described in full, and after both the reviewer still cannot say in one sentence what this paper did that the previous one did not.

This is almost never an absence of novelty. It is novelty that was never stated, because by the time you write the paper the delta is so obvious to you that it reads as not worth a sentence. It is worth a sentence. It is worth the sentence that opens your contributions list, and it should name the thing that was not possible before.

The tell is a related work section that describes neighbouring papers accurately and never says what is wrong with them, or what they could not do. A survey paragraph is not a positioning paragraph.

04

The experiments do not isolate the contribution

The system has four new parts and one number at the end, so there is no way to know which part earned it, or whether one of them is doing nothing.

Missing ablations are read as a hidden negative result, and often correctly. The specific fear is that the gain came from something incidental, longer training, a better learning rate, more parameters, and that the idea the paper is named after contributed little. A single ablation table answers it, and its absence in a paper with multiple components is one of the few omissions reviewers treat as evidence rather than as an oversight.

05

I could not follow section 4

Notation introduced after it is used, a method section that describes the implementation instead of the idea, a figure whose caption does not say what the reader is looking at.

Both big venues rate this on its own line, ICLR as presentation and NeurIPS as clarity, each out of 4. That matters more than it sounds. A number recorded separately cannot be carried by a reviewer who liked the rest, and it is the one weakness a rebuttal cannot argue away, because the thing being rated is the PDF you already submitted. It is also the cheapest of all of these to have avoided. What each NeurIPS field is for, and the same for ICLR.

06

This is good work for a different venue

Nothing is wrong with the paper. It is being judged against criteria that were never the point of the work.

This is the rejection that wastes the most time, because the paper needed no repair and the six months it costs buys no improvement. It is also the only one on this list you can settle before you submit rather than after, which is why the matcher is free and needs no account.

The part that is not about you

Nobody has to be at fault for the same paper to go both ways

NeurIPS has twice run the experiment on itself: route a slice of submissions through two independent committees and compare. In 2021 it duplicated 882 papers, and 50.6% of the papers one committee accepted were rejected by the other. In 2014, at a tenth the size, the figure was 49.5%. The conference publishes this, which is more than most venues do, and it is the best evidence anyone has for what the middle of the distribution is actually like.

Our band is a different measurement and it points the same way. We are not comparing two committees, we are asking how much of the outcome is recoverable from the manuscript, and for 154 of 697 papers the answer is almost none of it. Neither result says the process is broken or that reviewers are careless. Both say the same practical thing: in that band the decision was close, and close decisions are settled by things you cannot see from your desk.

The NeurIPS consistency experiment, in the conference's own words. What it means for any claim about machines reviewing papers is set out in AI review against human review.

The use of knowing this is not consolation, it is triage. It tells you which rejection to act on and which to resubmit. A paper that read weakly has work to do and the work is findable. A paper that read well and lost has a resubmission to prepare, and rewriting it into a different paper is the expensive mistake that group makes.

Find out which of the three your draft is in

The same specialists that produced the scores on this page will read your draft against the guidelines of the venue you are aiming at, and hand back the objections with locations, while the paper is still in your hands.

If the reviews are already back, the next thing to write is the reply: how to write a rebuttal. If they are not, the list of things that end a submission before anyone reads it is the pre-submission checklist.

phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.