Guide
NeurIPS review scores: what the numbers actually mean
A NeurIPS reviewer fills in six numbers, and authors read one of them. The overall score runs 1 to 6, not 1 to 10: it was cut in 2025, and most of what you will find written about NeurIPS scores still describes the old scale.
Published 11 August 2026
Your scores are in. The PDF is frozen and the area chair reads the discussion from here, so nothing sold on this site can move this submission. If you have not sent the paper yet, the page for that week is the pre-submission checklist
Read this before you read your reviews
- The overall score runs 1 to 6, not 1 to 10. It was cut in 2025, and almost every guide, forum answer and blog post you will find still explains the ten-point version.
- That inverts what a 4 means. A 4 was a borderline reject on the old scale and is a borderline accept on the current one. Check which edition your review is on before you read anything into it.
- You were given six numbers, not one. Quality, clarity, significance and originality are rated 1 to 4 each, the overall score 1 to 6, and confidence 1 to 5. Which field your weakness landed in decides whether it is recoverable.
- Only the step from 3 to 4 is a cliff. Every other step is a change of degree. That one crosses from a reviewer who would reject to one who would accept, and moving a single reviewer across it is the most a rebuttal can do.
- Confidence decides what a bad score costs you. A low score marked 2 or 3 for confidence is the most winnable review in your batch. The same score marked 5 will not move for anything general.
The form
Six numbers, six different jobs
Alongside the written sections, a NeurIPS review carries these scored fields. Knowing which one your weakness landed in tells you whether it is recoverable.
| Field | Range | What it is asking | Can a rebuttal move it |
|---|---|---|---|
| Quality | 1 to 4 | Is the submission technically sound, are the claims supported, is the method complete and does it evaluate itself honestly. | Yes. This is the field new results and corrected misreadings act on. |
| Clarity | 1 to 4 | Is the paper written and organised well enough that a reader can follow and reproduce it. | Rarely. It rates the PDF that was submitted, and that file is now fixed. |
| Significance | 1 to 4 | Are the results important, will other people use or build on them. | Sometimes, and it is hard. Disagreement here is usually about values, not facts. |
| Originality | 1 to 4 | Are the tasks or methods new, and is the related work adequately cited to show it. | Sometimes. A missed citation cuts against you here and is worth correcting fast. |
| Overall score | 1 to 6 | The recommendation, from strong reject through the two borderlines to strong accept. | It follows the four above. Nothing moves it on its own. |
| Confidence | 1 to 5 | How sure the reviewer is, from an educated guess to absolute certainty. | Yes, in both directions, which is why it needs handling with care. |
Fields and ranges as published in the NeurIPS 2025 reviewer guidelines, which is the edition that introduced the six-point overall score. The 2026 guidelines are organised around the same four dimensions. The rebuttal column is our reading of what each field responds to, not conference policy.
The overall score
What each number on the 1 to 6 scale means
Every point is anchored in words, which matters: a reviewer choosing 4 over 3 is choosing a described position, not expressing a feeling. The two anchors either side of the middle are where the conference is actually decided.
| Score | Label | What it means for you |
|---|---|---|
| 6 | Strong accept | Reserved for work the reviewer considers close to flawless with impact across areas. Most accepted papers never receive one. |
| 5 | Accept | An advocate. Technically solid with high impact on at least one sub-area. |
| 4 | Borderline accept | Technically solid, and the reasons to accept outweigh the reasons to reject. Narrowly on the right side, and still movable. |
| 3 | Borderline reject | The same sentence with the balance tipped the other way. This is the score a rebuttal is written for. |
| 2 | Reject | A named technical flaw, a weak evaluation or reproducibility the reviewer does not accept. |
| 1 | Strong reject | Not a close call, and not about polish. |
Scale and labels from the NeurIPS 2025 reviewer guidelines, paraphrased. Editions before 2025 used a 1 to 10 scale whose numbers do not translate. The tinted rows are where a rebuttal changes outcomes, because they are the scores an area chair can still be argued out of.
The gap between 3 and 4 is the only one that is a cliff
Every other step is a change of degree. This one crosses from a reviewer who would reject to one who would accept, and the two anchors are deliberately the same sentence with the balance of reasons reversed. Moving a single reviewer across it is the most valuable thing a rebuttal can do.
There is no weak accept any more, and that cuts both ways
The old ten-point scale had room to be lukewarm: weak accept, borderline accept, borderline reject all sat next to each other. Six points removes the hedging room, so each step now carries more weight and a single reviewer moving one point does more damage or more good than it used to.
Beware advice written for the old scale
Anything explaining a 6 as weak accept, a 7 as accept or a 10 as award quality is describing the pre-2025 form. On the current scale a 6 is the top of the range. If a page does not say which edition it is describing, it is not safe to read your review against it.
Confidence
The field that decides how much a bad score costs you
Confidence runs 1 to 5, from an educated guess to absolute certainty, and it is self-reported. It is the field authors skip and area chairs do not, because it is what lets a chair discount one review without accusing anyone of anything.
A low score with low confidence is an invitation
A reviewer who marks 2 or 3 on confidence is saying in the conference's own words that they may have missed the central parts of the submission. That is not hostility, it is a request for the explanation the paper failed to give them, and it is the single most winnable review in your batch.
A low score with a 5 is a different animal
Certainty plus rejection means the objection is specific and the reviewer believes they understood you. Nothing general will move it. If you cannot answer the exact technical point, spending the rebuttal on this reviewer is spending it on the one who will not change.
Never argue a reviewer into raising their confidence
It is tempting, because a supportive reviewer counts for more when they are certain. It cuts the other way just as hard: a reviewer who accepts your correction and re-reads the section may come back understanding exactly why they disliked it. Answer the substance and leave the field alone.
Our own measurement, on a different scale
Where the NeurIPS bar actually sat
There is no published threshold and no average that guarantees anything: area chairs decide, and they read the discussion rather than the arithmetic. What can be measured is where accepted and rejected papers sat when something read all of them the same way.
| Group | Papers | Lower quartile | Median phdflow score | Upper quartile |
|---|---|---|---|---|
| Accepted at NeurIPS | 100 | 6.55 | 6.90 | 7.30 |
| Rejected at NeurIPS | 97 | 5.56 | 6.03 | 6.40 |
This is the phdflow score, produced by reading the manuscript, on 197 NeurIPS 2023 papers. It runs 0 to 10 and it is not the reviewers' overall score, which runs 1 to 6: a 6.90 in this table is not a reviewer score at all, and would not be a possible one. Accepted and rejected are the conference's own decisions.
One rejected NeurIPS paper in five read like an accepted one
19 of the 97 rejected papers scored at or above 6.55, which is where the weakest quarter of accepted papers sat. Whatever separated them was not recoverable from the manuscript.
The middle is where it was decided
Of the NeurIPS papers landing between 6.0 and 6.5, 17 of 46 got in. That is the band where reading the paper stops predicting the outcome, and it is the same band we find at both venues. Why that band exists is set out in why papers get rejected.
The paper is frozen. The literature is not.
Decisions take weeks, and work on your problem keeps landing the whole time. Name the topic you just submitted on and Gap Alerts email you when a paper covering the same ground appears, which is the part of this submission you can still act on.
Reviews already back? The reply is the next thing to write: how to write a rebuttal. The equivalent breakdown for the other venue is ICLR review scores explained. If the next paper has not gone out yet, the failures that end a submission before a reviewer is assigned are in the pre-submission checklist.
phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.