Skip to content

Pre-submission review, benchmarked

Know the verdict before you submit

Upload a paper, pick a venue. Five agents read it in parallel, and an orchestrator turns their scores into one accept or reject call by applying that venue's own guidelines, with the issues behind it.

AUC 0.91 on 297 ICLR 2025 papers 96% accuracy, 1 false accept in 149 rejects 697 papers scored against real decisions 5 agents, 7 venues, 1 verdict

No card. Enough for two full reviews.

Measured, not claimed

Four numbers, and where each one comes from

Every figure on this page names its dataset, its sample size and its venue, because an accuracy number without those three things is not a number.

0.91 AUC on ICLR 2025 Rank-based, n=297
96% accuracy on the same papers 285 correct of 297
99% precision when it says accept 137 of 138 predicted accepts
697 papers scored against real decisions ICLR 2017, NeurIPS 2023, ICLR 2025

See the full result table and confusion matrix.

The problem

The review you need arrives after the decision

A submission cycle ends with a verdict and a set of reviews you can no longer act on for that venue. The feedback is real, it is often correct, and it is late by construction.

What you want before you submit is a straight answer to one question: read as a reviewer would read it, does this paper get in.

The things that sink papers are rarely the things that were hard to do. A missing baseline. Variance never reported. An assumption stated in your head and not on the page. A related work section that did not find the paper which already did this.

Your co-authors cannot give you that answer, because they already believe the paper. A grammar tool cannot give you that answer, because it was never asked to. A literature search returns papers, not a judgment.

So the decision to submit gets made on instinct, and the diagnosis arrives months later, attached to an outcome you cannot change.

phdflow does not replace that process. It gives you a rehearsal of it: the same question, answered against the same venue guidelines, in about four minutes, before the deadline rather than after it.

The cycle

  1. Submit
  2. Wait
  3. Decision
  4. Reviews you can no longer use

Node 1 is the last point where anything is cheap.

Fixable Before you submit
Locked in After you submit

A missing baseline

Costs you one weekend of compute

Costs you the next cycle

Variance never reported

Costs you an afternoon of reruns

Costs you the next cycle

An assumption stated in your head, not on the page

Costs you one paragraph

Costs you the next cycle

Related work that missed the closest 2024 paper

Costs you one search

Costs you the next cycle

Judge

An accept or reject call, with the issues that produced it

Upload a PDF, pick a venue, and five verdict agents read the paper in parallel. An orchestrator merges their scores into one call, and it is the orchestrator that applies the venue's guidelines. You get a score, a verdict, and a list of issues you can act on.

  • Every issue carries four things: a severity of error, warning or info, a category, the section it sits in, and a suggested fix. Nothing is a bare complaint.
  • The verdict is binary. Accept or reject, no comfortable middle. A borderline paper comes back as reject with the shortest route to accept written out in the issue list. The per-agent scores carry the nuance.
  • The orchestrator is constrained to synthesize. It is instructed to base the review only on what the sub-agents reported, and it is forbidden from inventing criticisms no agent raised.
  • The accept threshold is not hardcoded. The orchestrator reads the guidelines of the venue you picked, so the same scores can pass at one venue and fail at another. Seven venues are loaded today.
  • Timing, measured rather than promised: a median of about four minutes on real PDF uploads, range 3.4 to 12.9 minutes over our ten most recent runs.
  • Four minutes is short enough to run on a draft you are not proud of yet. That is the point. The value is not one verdict, it is how many times you hear it before the deadline.
ICLR 2025 Illustration
0.38
Overall score
Reject
error 4. Experiments

Main claim is not supported by the ablation in Section 4

Suggested fix: State the claim as conditional, or add the ablation that isolates the contribution.

warning 4.2 Results

No variance reported across seeds

Suggested fix: Report mean and standard deviation over at least three seeds.

info 2. Related work

Related work omits two closely matched 2024 baselines

Suggested fix: Position against the two nearest 2024 methods and say why yours differs.

Soundness
41%
Significance
36%
Statistical
29%
Coherence
74%
Novelty
52%
Orchestrator running

Explore

Before you write it, find out if it is already written

Give it a topic. Explore searches live literature, builds a graph of prior art and stated future work, scores novelty, and ends with one word: GO, PIVOT or PASS.

  • The output is a graph, not a list. Nodes are papers and threads, edges are the relationships between them, and the summary tells you which corner of the space is crowded and which is open.
  • The novelty score is a number between 0 and 1. Across our runs so far it has landed anywhere from 0.20 to 0.90, so it discriminates rather than flattering every idea you feed it.
  • The verdict is mandatory in the output format. The model is required to close with GO, PIVOT or PASS, which means it cannot hedge its way out of a recommendation.
  • It reads real papers. Explore searches the live literature during your run, so every node in your graph is a paper it actually retrieved, not a hit from a frozen index.
  • A run takes roughly 3 to 7 minutes, so checking five directions before committing to one costs less than a single paper review.
  • A topic that turns out to be occupied costs an afternoon to discover in month zero and six months to discover in month six.
Illustration
62%
PIVOT

The retrieval half of this space is crowded. The evaluation half is not.

Nearest prior art: three 2024 papers already do guideline-grounded retrieval.

Open gap: nobody reports faithfulness against the source guideline itself.

Reshaped angle: benchmark the faithfulness, do not rebuild the retriever.

Novelty 0.62 Depth 3 GO / PIVOT / PASS

Gap Alerts

New

Watch a topic. We re-run the search and email you when something new lands.

The literature does not hold still between the day you pick a direction and the day you submit on it. That window is exactly when someone else publishes the thing you were about to write.

Subject: 2 new papers on "guideline-grounded clinical QA"

Faithfulness benchmarks for retrieval-augmented clinical answers

Matched your watch on evaluation of guideline grounding.

Illustration

Start a watch from any topic, and 100 credits keeps it running. You get the paper, the reason it matched your topic, and a link straight back into your graph.

How it works

Five agents, one orchestrator, one verdict

Not one prompt in a loop. Five specialist agents, each with its own calibrated prompt and its own slice of the paper, and an orchestrator that turns their scores into a single venue-aware call.

1 Your PDF
2 Structured parse
3 5 agents in parallel
4 Orchestrator plus venue guidelines
5 Verdict and issues
Sets the verdict

Soundness

Does the evidence in the paper support the claims it makes.

Sets the verdict

Significance

What kind of contribution this is, and how much it moves the field.

Sets the verdict

Statistical

Experimental rigour: baselines, seeds, variance, significance.

Sets the verdict

Coherence

Whether the argument holds together from abstract to conclusion.

Sets the verdict

Novelty

Searches the live literature for the closest prior work to your claim.

Orchestrator

Reviewer

Merges the five scores, applies the venue guidelines, writes the verdict.

Each agent returns a score and its own issue list. The orchestrator is the only component that reads the venue guidelines, which is why the same paper can pass at one venue and fail at another.

The agents do not guess at the literature. Novelty searches Semantic Scholar and arXiv while your review runs, and your PDF is parsed into real sections and references first, so an issue can point at the section it came from.

Venue criteria live in a database rather than in the code, so a venue can be added or updated without touching the system.

Evidence

697 papers, scored against the decisions the venues actually made

We ran the system over three conference datasets with known outcomes and compared its call to the real one. Here is the whole table.

VenuePapersAccuracyAUC
ICLR 201720385.2%0.84
NeurIPS 202319791.9%0.87
ICLR 202529796.0%0.91
All venues69791.7%0.87

Ground truth is the venue's own decision: OpenReview for ICLR 2025 and NeurIPS 2023, the PeerRead release for ICLR 2017.

Confusion matrix, ICLR 2025

Confusion matrix for ICLR 2025, 297 papers
137 True accepts1 False accepts
11 False rejects148 True rejects

One false accept in 149 rejected papers.

SystemDatasetAUCSource
phdflowICLR 2025, n=2970.91Recomputed from our result data
Stanford Agentic ReviewerICLR 2025, their own sample0.75Kukoyi et al., 2026, preprint
The AI ScientistICLR 2022, different set0.65Lu et al., 2024

Across all 697 papers: 91.7% accuracy, precision 91.1%, recall 92.2%, F1 91.7%, pooled AUC 0.87.

On ICLR 2025 the AUC is 0.91, against the 0.75 reported by the Stanford Agentic Reviewer preprint on the same venue and year. Ground truth in both cases is the conference's own decision, published on OpenReview.

The asymmetry is the useful part: when it says accept, it has been right 137 times out of 138. It errs toward being harsh, not toward flattering you, so a weak paper does not get waved through.

Pricing

Prepaid credits, priced per run

No subscription. Buy credits, spend them on whatever mix of Judge and Explore you need, and they do not expire on a monthly cycle. New accounts get 120 credits free with no card.

Free

€0

120 credits on signup

2 full paper reviews

  • 4 topic explorations
  • No card

No trial countdown. Credits do not expire on a monthly cycle.

Starter

€29

420 credits

7 full paper reviews

  • 14 topic explorations
  • €4.14 per review

VAT included. Invoice issued for every purchase.

Best value per review

Pro

€79

1500 credits

25 full paper reviews

  • 50 topic explorations
  • €3.16 per review

24% less per review than Starter. VAT included, invoice for every purchase.

Lab

From 4800

credits, priced on request

80 full paper reviews or more

  • 160 topic explorations
  • Shared invoicing

Email us and we will quote it, with an invoice for your institution.

ActionCreditsWhat 120 free credits buys
Judge, single pass602 runs
Judge, second pass on your revised draft1201 run
Explore, one topic304 runs
Gap Alert, one topic watch1001 watch

Buying for a lab, or need the invoice made out to your institution? Email us before you pay and we will handle it as a Lab order. A VAT number cannot be added to an invoice after the fact.

Prices include VAT. An invoice is issued for every purchase, which matters if this is going on a grant.

Operator: Finance Foresight AI Kamil Szczepanik, NIP 7252336619, Poland.

Questions

Frequently asked questions

The ones we would ask.

Will using this get my paper desk-rejected?

The conference policies people are thinking of here govern two things: reviewers running confidential submissions through an LLM, and authors passing off LLM-generated text as their own writing. phdflow is neither. You run it on your own manuscript, before submission, to get feedback, the same way you would ask a lab-mate or pay an editor. It does not write your paper and it does not submit a review anywhere. That said, disclosure rules are set by your venue and they change every cycle. Read your venue's policy and remain responsible for it. We are not going to claim blanket compliance on your behalf.

Where does my unpublished manuscript go?

Your PDF is uploaded to our infrastructure and parsed by Grobid, which runs in our own container rather than a third-party service. The extracted text is then sent to Google's Gemini API for the agent runs, and the results are stored in our database so you can reload the review later. You can delete a review and its stored content from your account. If you need contractual guarantees about retention windows or sub-processors before you upload unpublished work, email us and we will answer in writing rather than in marketing copy.

Is this just ChatGPT with a wrapper?

You can test that claim cheaply. Five agents with separately calibrated prompts, a tool layer behind a three-replica cluster, venue guidelines loaded from a database, and an orchestrator constrained to synthesize rather than to invent. The result is measurable: AUC 0.91 on 297 ICLR 2025 papers with the confusion matrix printed above. A single prompt does not get there.

Can an LLM really judge novelty? That is the only thing that matters.

Novelty is the hardest dimension for any reviewer, human or otherwise, which is why we do not leave it to one judgement. Five agents score different dimensions and the orchestrator combines them against the venue criteria: that ensemble is what reaches AUC 0.91. For the novelty question specifically, Explore is the sharper tool, because it goes and reads the literature around your idea rather than scoring your text.

How do I know AUC 0.91 is real?

Three ways to check. The dataset is 297 ICLR 2025 papers we scraped from OpenReview, which is also where the Stanford Agentic Reviewer preprint drew its ICLR 2025 sample, so the task and the ground-truth source match even though the paper lists are not identical. The confusion matrix is printed above, so accuracy, precision and recall are all recomputable from four numbers. And the strongest possible test costs you nothing: spend your free credits on a paper whose outcome you already know.

Will it just tell me what I want to hear?

The measured failure mode is the opposite. On ICLR 2025 it produced 11 false rejects and 1 false accept, so it is biased toward being harsh rather than toward flattering you. Precision when it says accept is 99%. In practice that means an accept from phdflow is a meaningful signal, and a reject from it is a prompt to look harder at what it flagged rather than a prediction you should treat as settled.

Why would I pay when Consensus and ResearchRabbit are free?

Those are retrieval tools and they are good at retrieval. They do not tell you whether your paper reads as a reject. The comparison is not against a free search box, it is against finding out at the decision notification. At about four euro per review, one caught missing baseline pays for the pack several times over.

Does it give a Revise verdict?

No. The verdict is accept or reject today. There is no middle option, deliberately, because a soft verdict is agreeable and unactionable. The nuance lives in the per-agent scores and the issue list, and the second pass on your revised draft is designed for exactly the case where the first answer is reject.

How long does a review take?

About four minutes on a real PDF. Across our ten most recent uploads the median was roughly four minutes and the range was 3.4 to 12.9 minutes. Five agents run in parallel, so paper length matters less than you would expect. An Explore run is usually 3 to 7 minutes. We would rather quote a measured range than a marketing number.

Does this replace peer review?

No, and it is not trying to. It prepares you for peer review: the same question, the same venue criteria, answered before the deadline instead of after the decision. Human reviewers know things about your subfield that no benchmark captures. What this gives you is one more competent read of your paper at a point where you can still act on it.

Run it on a paper whose outcome you already know

That is the honest test, and it is the one we would run. 120 free credits is two full reviews or four topic explorations. If the verdict matches what happened, you have your answer about whether to trust it on the next paper.

No card. No subscription. Credits do not expire on a monthly cycle. Finance Foresight AI Kamil Szczepanik, NIP 7252336619, Poland.