Guide
What AI peer review cannot do
We build one of these tools, so take the following as interested testimony. These are the eight objections the field raises, put as sharply as we can put them, and answered. Several of them are simply correct.
Published 8 August 2026
Deciding whether to trust one of these. The answer this page reaches is narrow: a rehearsal of the reading your paper is about to get, worth something only while the draft is still open. What that rehearsal looks for is in why papers get rejected
Who reads the output
- The report comes back to you and to nobody else. phdflow runs on your own manuscript, before you submit it.
- It is not sold to reviewers, editors or programme chairs. It never sees anyone else's submission.
- That is not a disclaimer. It decides how much of what follows applies to you.
The line venues already drew
Reviewers mostly may not. Authors may.
Conferences and publishers have spent three years writing rules about this, and almost all of them separate the same two cases. The reason is confidentiality: a reviewer holds someone else's unpublished work, and you hold your own.
| Venue | Reviewer, running a submission through a model | Author, running their own draft |
|---|---|---|
| CVPR, ECCV | Forbidden outright, including a model run locally. | Permitted, and you stay accountable for what it produces. |
| NeurIPS | Submissions and code may not be shared with any model. | Welcomed. Documented only when the model is a substantive component. |
| ICLR | Permitted with disclosure. Failing to disclose can cost the reviewer their own submission. | Permitted, and use must be stated in the paper and the submission form. |
| ACL Rolling Review | Forbidden, down to the review text in a public tool. | Permitted with disclosure. |
| Elsevier | Forbidden to upload the manuscript. Assistive use disclosed. | Explicitly lists checking your manuscript before submission as a permitted use. |
| IEEE | Forbidden to process a manuscript under review through a public platform. | Permitted, disclosed in the acknowledgements. |
| Wiley, Science | Forbidden to put any part of a manuscript into an AI tool. | Permitted with disclosure. Science wants it in the cover letter too. |
Summarised from each venue's published policy, read August 2026. These documents change every few months, so check the current wording for the venue you are targeting before you rely on this table. Disclosure requirements differ: most publishers exempt plain grammar and spelling checks and require a statement once a tool changes structure or content.
The objections
Eight things people say, and what we think is true
It hallucinates, and you cannot tell when
A model's confidence has nothing to do with whether it is right, so a false statement arrives in exactly the register of a true one. The errors are not random noise either. It invents the paper that ought to exist, the method that would be standard practice, the citation that fits the sentence. Those are the ones that survive a smell test.
All of that is accurate, and the part of it we can do something about is small. Where the system makes a claim about work outside your paper, it is built to search rather than to recall: the literature reader is given three search tools and nothing else, it queries Semantic Scholar and arXiv while your run is going, and the search is capped at the year you are submitting into, so a paper going to a 2017 venue is checked against what existed in 2016. The step that writes the final review is instructed to add no criticism the specialist readers reported.
None of that touches the real exposure. Four of the five readers work from your text alone, with no tools, and everything they say about your paper is a model reading your prose. It can misread, and we have no detector that catches it. What makes that survivable is who is holding the report: you already know whether that library exists. An editor reading the same sentence about someone else's paper does not.
It has no idea whether a result is physically possible
Judging whether a number is plausible is not a language task. It is knowledge of how things break, and in domains where the physics defeats finite element modelling anyway, an expert's judgement is the only instrument there is. A model has no access to it and no way of knowing it lacks it, so it will produce a fluent assessment of a result a person with their hands in the furnace would reject in two seconds.
This one we concede completely. There is nothing in the system that evaluates physical plausibility, not poorly and not approximately. It can tell you the argument for a number is thin or well made. It cannot tell you the number is impossible. If your paper stands or falls on whether that value is credible for that composite at that temperature, that judgement stays yours.
It recommends changes that would not improve the work
A model trained on the surface of good papers pushes every paper toward the mode of its training distribution. Asking for a parametric test when the non-parametric one was correct, the distribution is not normal and the sample is small because the materials are expensive, is not a one-off slip. It is a tool preferring the more common pattern to the right one, and at scale that is a conformity pump.
The example is exactly right and we cannot fix the general case. The system does not know why your n is twelve, and nothing in it reasons about test selection or the cost of your materials. We have told the experiment reader by hand not to mark you down for missing error bars or a missing sensitivity analysis, but that is a list of known nuisances rather than domain competence.
This is the objection where the recipient matters most. A reviewer demanding that test costs you a month or the paper. The same suggestion on your own screen costs you the second it takes to skip the paragraph.
It cannot tell real novelty from an unmotivated collage
Novelty a search engine can measure is a negative property: nothing matched. The novelty that matters is a positive one, a claim that moves what the field can do. Retrieval cannot close that gap, and the failure is exploitable in the worst direction: the more arbitrary the combination, the less likely anything matches, so the incoherent paper scores best.
That mechanism is real and it applies to us literally. The novelty reader has three tools and all three are searches. What it establishes is that your combination has not been published, which is not the same claim as your combination being worth making, and on an arbitrary enough pairing it will return a comfortable answer.
So read that score as unprecedented rather than as important, and read soundness and significance for the part that is supposed to hurt. This is also the direction of error that should worry you more generally: a warning you disagree with costs you nothing, and false reassurance has no corrective at all, because nothing prompts you to go and check.
Nobody has shown it agrees with experts
Predicting an accept or reject label is a different task from reviewing, and the label is contaminated: acceptance correlates with author prestige, affiliation, topic fashion and competent formatting, all of which are visible in the PDF. A system can score well by learning the sociology of a venue without touching the science once.
We have never compared what phdflow writes with what reviewers wrote, and we hold no measurement of that kind. What we compared is our call against what committees decided, on 697 papers from two machine learning conferences, and agreeing with an outcome is not evidence of reasoning like a reviewer. On genuinely new problems the expert is the only ground truth available, and we have no access to it.
Engineering and clinical work is not scored by novelty
In applied fields a paper earns its place by being done correctly and by changing a decision. Both require knowing the practice: the standard of care, the tolerance the part has to hold, what the regulator demands. None of that is in the text. A tool that scores novelty and significance from the manuscript is measuring the wrong construct, and a wrong measure is worse than none, because people act on it.
The criticism lands on how the verdict is built. Novelty counts for the same share of the score whether your paper is theoretical or applied, and for a field that rewards care over firstness that is the wrong share. The venue's own guidelines do push back. For somewhere like IEEE Access they tell the system in as many words not to mark a paper down for modest novelty alone. What they cannot change is the readers underneath, which are calibrated on top machine learning conferences and have been measured nowhere else.
Two gaps, and only one is serious. Literature coverage in medicine is fine: Semantic Scholar indexes PubMed. Validation in medicine does not exist.
It has never been tried on source-based disciplines
In history and neighbouring fields the object of review is the relation between a claim and a body of primary material that is not machine readable, often not indexed, frequently not digitised, and sometimes in a script that takes training to read. There is no Semantic Scholar for parish registers. Generalising from two computer science conferences to science as a whole is a category error, and computer science is unrepresentative on the one axis that decides this: it is the field whose entire output is text-native, indexed, open and recent.
We have not measured this on a single paper outside computer science. The literature search has almost no coverage of the humanities, so for archive-based work the novelty reader is blind rather than merely uncertain. The honest word for our output there is untested, and untested is a different thing from weak.
| Field | Literature coverage | Validated by us |
|---|---|---|
| Machine learning | Good | Yes, 697 papers |
| Medicine, engineering, chemistry | Good | No |
| History and other source-based fields | Almost none | No |
Screening matters in the tails, and nobody measures the tails
Almost all the value of a screen sits in the extremes, and the extremes are statistically out of reach: you cannot assemble a test set of breakthroughs, because the label only exists in hindsight, by which time the text is in every training corpus. So "it performs well on average" is not merely insufficient, it is unfalsifiable exactly where it would matter. The mechanism is not mysterious: a model fits the central tendency, a breakthrough is by definition an outlier against it, and the expected behaviour is to punish what should be rewarded.
We cannot tell you whether this would recognise an important paper, and neither can anyone else. What we can do is put the number in front of you. On the 2025 ICLR set, eleven papers we called reject were accepted by the conference. On the older 2017 set it is fourteen out of ninety-eight, which is the worst that figure gets anywhere in our benchmark.
The honest conclusion is a limit on what the tool is for. This is built for the middle of the distribution, where good papers fail for boring and fixable reasons, and that is the overwhelming majority of papers including probably yours. It is not a discovery instrument and we do not know how to build one.
Where that leaves it
A rehearsal, not a verdict
Nothing above argues that a model should be reviewing anybody's paper. We do not think it should, and the venues that banned it for reviewers were right to. What is left once you take that off the table is narrow and worth having: a rehearsal of the reading your paper is about to get, run by you, on your own work, while you can still do something about it.
The way to judge that is not to read more of our copy about it. Open a review the system already produced, on a paper whose fate is public, and see whether the objections it raises are ones you would have raised. If they are, that is the only argument for any of this that means anything. If they are not, you have lost ten minutes and you know.
If the draft is still open, the failures that end a submission before anyone reads it are in the pre-submission checklist.
phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.