Skip to content

Guide

What AI peer review cannot do

We build one of these tools, so take the following as interested testimony. These are the eight objections the field raises, put as sharply as we can put them, and answered. Several of them are simply correct.

Published 8 August 2026

Deciding whether to trust one of these. The answer this page reaches is narrow: a rehearsal of the reading your paper is about to get, worth something only while the draft is still open. What that rehearsal looks for is in why papers get rejected

Who reads the output

The line venues already drew

Reviewers mostly may not. Authors may.

Conferences and publishers have spent three years writing rules about this, and almost all of them separate the same two cases. The reason is confidentiality: a reviewer holds someone else's unpublished work, and you hold your own.

What venues permit reviewers and authors to do with language models
VenueReviewer, running a submission through a modelAuthor, running their own draft
CVPR, ECCVForbidden outright, including a model run locally.Permitted, and you stay accountable for what it produces.
NeurIPSSubmissions and code may not be shared with any model.Welcomed. Documented only when the model is a substantive component.
ICLRPermitted with disclosure. Failing to disclose can cost the reviewer their own submission.Permitted, and use must be stated in the paper and the submission form.
ACL Rolling ReviewForbidden, down to the review text in a public tool.Permitted with disclosure.
ElsevierForbidden to upload the manuscript. Assistive use disclosed.Explicitly lists checking your manuscript before submission as a permitted use.
IEEEForbidden to process a manuscript under review through a public platform.Permitted, disclosed in the acknowledgements.
Wiley, ScienceForbidden to put any part of a manuscript into an AI tool.Permitted with disclosure. Science wants it in the cover letter too.

Summarised from each venue's published policy, read August 2026. These documents change every few months, so check the current wording for the venue you are targeting before you rely on this table. Disclosure requirements differ: most publishers exempt plain grammar and spelling checks and require a statement once a tool changes structure or content.

The objections

Eight things people say, and what we think is true

01

It hallucinates, and you cannot tell when

A model's confidence has nothing to do with whether it is right, so a false statement arrives in exactly the register of a true one. The errors are not random noise either. It invents the paper that ought to exist, the method that would be standard practice, the citation that fits the sentence. Those are the ones that survive a smell test.

All of that is accurate, and the part of it we can do something about is small. Where the system makes a claim about work outside your paper, it is built to search rather than to recall: the literature reader is given three search tools and nothing else, it queries Semantic Scholar and arXiv while your run is going, and the search is capped at the year you are submitting into, so a paper going to a 2017 venue is checked against what existed in 2016. The step that writes the final review is instructed to add no criticism the specialist readers reported.

None of that touches the real exposure. Four of the five readers work from your text alone, with no tools, and everything they say about your paper is a model reading your prose. It can misread, and we have no detector that catches it. What makes that survivable is who is holding the report: you already know whether that library exists. An editor reading the same sentence about someone else's paper does not.

02

It has no idea whether a result is physically possible

Judging whether a number is plausible is not a language task. It is knowledge of how things break, and in domains where the physics defeats finite element modelling anyway, an expert's judgement is the only instrument there is. A model has no access to it and no way of knowing it lacks it, so it will produce a fluent assessment of a result a person with their hands in the furnace would reject in two seconds.

This one we concede completely. There is nothing in the system that evaluates physical plausibility, not poorly and not approximately. It can tell you the argument for a number is thin or well made. It cannot tell you the number is impossible. If your paper stands or falls on whether that value is credible for that composite at that temperature, that judgement stays yours.

03

It recommends changes that would not improve the work

A model trained on the surface of good papers pushes every paper toward the mode of its training distribution. Asking for a parametric test when the non-parametric one was correct, the distribution is not normal and the sample is small because the materials are expensive, is not a one-off slip. It is a tool preferring the more common pattern to the right one, and at scale that is a conformity pump.

The example is exactly right and we cannot fix the general case. The system does not know why your n is twelve, and nothing in it reasons about test selection or the cost of your materials. We have told the experiment reader by hand not to mark you down for missing error bars or a missing sensitivity analysis, but that is a list of known nuisances rather than domain competence.

This is the objection where the recipient matters most. A reviewer demanding that test costs you a month or the paper. The same suggestion on your own screen costs you the second it takes to skip the paragraph.

04

It cannot tell real novelty from an unmotivated collage

Novelty a search engine can measure is a negative property: nothing matched. The novelty that matters is a positive one, a claim that moves what the field can do. Retrieval cannot close that gap, and the failure is exploitable in the worst direction: the more arbitrary the combination, the less likely anything matches, so the incoherent paper scores best.

That mechanism is real and it applies to us literally. The novelty reader has three tools and all three are searches. What it establishes is that your combination has not been published, which is not the same claim as your combination being worth making, and on an arbitrary enough pairing it will return a comfortable answer.

So read that score as unprecedented rather than as important, and read soundness and significance for the part that is supposed to hurt. This is also the direction of error that should worry you more generally: a warning you disagree with costs you nothing, and false reassurance has no corrective at all, because nothing prompts you to go and check.

05

Nobody has shown it agrees with experts

Predicting an accept or reject label is a different task from reviewing, and the label is contaminated: acceptance correlates with author prestige, affiliation, topic fashion and competent formatting, all of which are visible in the PDF. A system can score well by learning the sociology of a venue without touching the science once.

We have never compared what phdflow writes with what reviewers wrote, and we hold no measurement of that kind. What we compared is our call against what committees decided, on 697 papers from two machine learning conferences, and agreeing with an outcome is not evidence of reasoning like a reviewer. On genuinely new problems the expert is the only ground truth available, and we have no access to it.

06

Engineering and clinical work is not scored by novelty

In applied fields a paper earns its place by being done correctly and by changing a decision. Both require knowing the practice: the standard of care, the tolerance the part has to hold, what the regulator demands. None of that is in the text. A tool that scores novelty and significance from the manuscript is measuring the wrong construct, and a wrong measure is worse than none, because people act on it.

The criticism lands on how the verdict is built. Novelty counts for the same share of the score whether your paper is theoretical or applied, and for a field that rewards care over firstness that is the wrong share. The venue's own guidelines do push back. For somewhere like IEEE Access they tell the system in as many words not to mark a paper down for modest novelty alone. What they cannot change is the readers underneath, which are calibrated on top machine learning conferences and have been measured nowhere else.

Two gaps, and only one is serious. Literature coverage in medicine is fine: Semantic Scholar indexes PubMed. Validation in medicine does not exist.

07

It has never been tried on source-based disciplines

In history and neighbouring fields the object of review is the relation between a claim and a body of primary material that is not machine readable, often not indexed, frequently not digitised, and sometimes in a script that takes training to read. There is no Semantic Scholar for parish registers. Generalising from two computer science conferences to science as a whole is a category error, and computer science is unrepresentative on the one axis that decides this: it is the field whose entire output is text-native, indexed, open and recent.

We have not measured this on a single paper outside computer science. The literature search has almost no coverage of the humanities, so for archive-based work the novelty reader is blind rather than merely uncertain. The honest word for our output there is untested, and untested is a different thing from weak.

Literature coverage and validation by field
FieldLiterature coverageValidated by us
Machine learningGoodYes, 697 papers
Medicine, engineering, chemistryGoodNo
History and other source-based fieldsAlmost noneNo
08

Screening matters in the tails, and nobody measures the tails

Almost all the value of a screen sits in the extremes, and the extremes are statistically out of reach: you cannot assemble a test set of breakthroughs, because the label only exists in hindsight, by which time the text is in every training corpus. So "it performs well on average" is not merely insufficient, it is unfalsifiable exactly where it would matter. The mechanism is not mysterious: a model fits the central tendency, a breakthrough is by definition an outlier against it, and the expected behaviour is to punish what should be rewarded.

We cannot tell you whether this would recognise an important paper, and neither can anyone else. What we can do is put the number in front of you. On the 2025 ICLR set, eleven papers we called reject were accepted by the conference. On the older 2017 set it is fourteen out of ninety-eight, which is the worst that figure gets anywhere in our benchmark.

The honest conclusion is a limit on what the tool is for. This is built for the middle of the distribution, where good papers fail for boring and fixable reasons, and that is the overwhelming majority of papers including probably yours. It is not a discovery instrument and we do not know how to build one.

Where that leaves it

A rehearsal, not a verdict

Nothing above argues that a model should be reviewing anybody's paper. We do not think it should, and the venues that banned it for reviewers were right to. What is left once you take that off the table is narrow and worth having: a rehearsal of the reading your paper is about to get, run by you, on your own work, while you can still do something about it.

The way to judge that is not to read more of our copy about it. Open a review the system already produced, on a paper whose fate is public, and see whether the objections it raises are ones you would have raised. If they are, that is the only argument for any of this that means anything. If they are not, you have lost ten minutes and you know.

If the draft is still open, the failures that end a submission before anyone reads it are in the pre-submission checklist.

phdflow reads a finished draft against the guidelines of the venue you are aiming at and returns an accept or reject call with the objections behind it. Scored against 697 real decisions: AUC 0.91 and 96% accuracy on 297 ICLR 2025 papers, one false accept in 149 rejects. The full measurement. Read a finished review before you pay for one. Your first paper gets a free preview; the full report is 60 credits, €9. No account needed. Pricing.