Guide
ICLR review scores: what 3, 5, 6 and 8 mean
The ICLR scale has no 2, no 4, no 7 and no 9. The gaps are the point: there is no middle to retreat to, so every reviewer has to pick a side, and the whole conference happens in the one point between 5 and 6.
Published 11 August 2026
Your scores are in. A 5, 6, 5 is the discussion phase rather than a verdict, and nothing on this site changes the PDF now. If the paper has not gone out yet, the page for that week is the pre-submission checklist
The short version
- 5 and 6 are the same sentence with the verdict flipped. Marginally below the acceptance threshold, and marginally above it.
- Everything else on the scale says the reviewer is not really in the argument. The action is in those two numbers.
- Scores of 5, 6, 5 are not a rejection. You have been handed the discussion phase, and ICLR is a venue where the discussion phase decides things.
The scale
Six positions, and only two of them are contested
| Rating | Label | What it means for you |
|---|---|---|
| 10 | Strong accept, should be highlighted | Reserved for work a reviewer wants on stage. Most accepted papers never receive one. |
| 8 | Accept, good paper | An advocate. One 8 that survives the discussion usually carries a paper. |
| 6 | Marginally above the acceptance threshold | A very common score on papers that get in. Not a compliment and not a warning. |
| 5 | Marginally below the acceptance threshold | The same reviewer, one step the other way. This is the score a rebuttal exists for. |
| 3 | Reject, not good enough | A named problem the reviewer does not expect to be fixed inside the discussion window. |
| 1 | Strong reject | Not a close call. Rare, and rarely about presentation. |
This is the anchor set ICLR used for 2024 and 2025, which is also the edition our own figures below come from. For 2026 the overall rating is mapped to 0, 2, 4, 6, 8, 10 instead, so a paper can now score a 0. ICLR keeps the scale in the OpenReview form rather than in its published reviewer guide, which is why the change was never announced anywhere you would look; both sets are documented in Kargaran et al. Read the anchors in your own form before you read anything into a number. The tinted rows are the two scores that decide papers.
The missing numbers are doing work
A scale with every position available lets a reviewer park in the middle and hand the decision to somebody else. Removing 4 and 7 removes the fence: the two options nearest the centre are explicitly labelled below the threshold and above it, so choosing either is choosing an outcome. NeurIPS reached the same place by a different route when it cut its overall score to six points, which also leaves no neutral seat.
A 6 is a pass, not a near miss
Marginally above the acceptance threshold is an ordinary score on a paper that gets in. Authors read the word marginally and start rewriting. The number below it is the one that needs work.
An 8 is worth more than two 6s
Scores are not averaged into a decision. An area chair reads a thread, and a reviewer willing to defend a paper in writing changes a discussion in a way two quiet 6s do not. This cuts both ways: one determined 3 is heavier than its arithmetic too.
| The rest of the form | Range | What it is asking | Can a rebuttal move it |
|---|---|---|---|
| Soundness | 1 to 4 | Is the technical claim actually supported by the theory, the method and the experiments. | Yes. This is the field new results and corrected misreadings act on. |
| Presentation | 1 to 4 | Is the paper written and organised well enough to be assessed properly. | Rarely. It rates the PDF you submitted, and that file is now frozen. |
| Contribution | 1 to 4 | How much this moves the field, with novelty and importance bundled into one number. | Sometimes, and it is the hardest. Disagreement here is usually about values, not facts. |
| Confidence | 1 to 5 | How sure the reviewer is, from an educated guess to absolute certainty. | Yes, in both directions, which is why it needs care. |
It is widely repeated that ICLR reviewers only write prose. They do not: these scores are filed alongside the summary, strengths, weaknesses and questions. Fields confirmed against the OpenReview record for ICLR 2024 and 2025 in Kargaran et al. The rebuttal column is our reading, not conference policy.
What makes ICLR different
The scores are not the whole record, and the record is public
ICLR runs on OpenReview, and the submission, the reviews, your replies and the reviewers' replies to those are visible. That changes what a score is worth, in ways that are worth planning for before you are in the middle of it.
Reviewers write differently when other reviewers can read them
A thin review is visible to everyone assigned to the paper, including the area chair. It makes the low-effort rejection harder to file and the unsupported score harder to keep, which is the strongest structural argument in favour of engaging rather than accepting the first number you see.
Your replies are part of the paper's record
An intemperate rebuttal does not just fail, it stays. This is not a reason to be timid. It is a reason to write the reply you would be happy to have quoted next to your name, which is usually also the reply that works.
Silence is a decision
A reviewer who is not answered has no reason to revisit a score, and the chair sees an objection nobody contested. Of the ratings on this page only two are worth your energy, but those two are worth all of it.
Confidence
Read the confidence before you read the rating
Every review carries a self-reported confidence from 1 to 5, running from an educated guess up to absolute certainty, with the reviewer stating whether they checked the maths and know the related work. It is the field that tells you which reviews are worth your limited rebuttal.
A 3 filed at confidence 2 is a reviewer saying they may not have understood the central parts of the submission. That is a comprehension failure, and comprehension failures are the ones you can fix in a paragraph. A 3 at confidence 5 from someone who names the exact step they do not believe is a technical dispute, and the only thing that moves it is the technical answer. Sorting your reviews this way, rather than by how much they annoyed you, is most of the work of a good rebuttal.
Our own measurement, on a different scale
Where the ICLR bar actually sat
No average rating guarantees anything, because ICLR publishes no threshold and area chairs read the discussion rather than the arithmetic. What can be measured is where the two groups sat when one reader went through all 297 of them the same way.
| Group | Papers | Lower quartile | Median phdflow score | Upper quartile |
|---|---|---|---|---|
| Accepted at ICLR | 148 | 6.45 | 6.70 | 6.93 |
| Rejected at ICLR | 149 | 5.10 | 5.75 | 6.17 |
This is the phdflow score, produced by reading the manuscript, on 297 ICLR 2025 papers. It is not the reviewers' rating, and the two scales are not comparable despite both ending at 10. Accepted and rejected are the conference's own decisions.
The accepted group is tight and the rejected group is not
Half the accepted papers sit inside half a point, between 6.45 and 6.93. The rejected papers spread across more than a full point. Getting in looks like clearing a narrow band rather than like being exceptional, which is the opposite of how it feels from the outside.
Some rejected papers read exactly like accepted ones
18 of the 149 rejected papers scored at or above 6.45, where the weakest quarter of accepted papers sat, and the strongest of them reached 7.70. Nothing in the manuscript separated them from the papers that got in.
Between 6.0 and 6.5 it was a coin flip
34 of the 71 ICLR papers landing in that band were accepted. That is the band where reading the paper stops predicting the outcome, and it shows up at both venues we measured.
Nothing moves that PDF now. The field keeps moving.
Discussion runs for weeks and arXiv does not pause for it. Name the topic you submitted on and Gap Alerts email you when a paper covering the same ground appears, so an overlap reaches you as an email rather than as a reviewer's comment.
The same breakdown for the other venue is NeurIPS review scores explained, and what happens in the band where the decision is a coin flip is in why papers get rejected.
phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.