Architecture
A panel of specialists, not one prompt in a loop
Your paper is not handed to one model with one question. It is taken apart, read in parallel by specialists looking for different kinds of weakness, and judged against the venue you chose.
Inside a run
Why the output is not one opinion
The paper is taken apart before anything reads it
A PDF is a layout format, not a document. Yours is recovered into real structure first: sections, references, the citations in the text. Every reader works from that structure rather than from a wall of extracted text.
Different readers, different failures
The specialists share no prompt and never see each other's answers. Each is set up for one kind of weakness and reports separately, so strength in one respect cannot paper over thinness in another.
The bar belongs to the venue
What counts as good enough is not written into the system. It comes from the guidelines of the venue you picked, so the same paper can clear one and fall short at another.
Prior work is checked live
The literature is searched while your run is in progress, not recalled from a model's memory, so a claim that your work sits close to something existing was checked against the record rather than remembered.
The verdict is assembled, not improvised
It is composed from independent assessments and the evidence behind them, and nothing can enter it that was not found in your paper. If part of the panel fails, the run still finishes and reports what held.
Evidence
Measured against real decisions
| Venue | Papers | Accuracy | AUC |
|---|---|---|---|
| ICLR 2017 | 203 | 85.2% | 0.84 |
| NeurIPS 2023 | 197 | 91.9% | 0.87 |
| ICLR 2025 | 297 | 96.0% | 0.91 |
| All venues | 697 | 91.7% | 0.87 |
Across all 697: precision 91.1%, recall 92.2%, F1 91.7%. Ground truth is each venue's own decision. The two conferences in this table do not draw the line in the same place, and we measured the gap: ICLR vs NeurIPS.
Confusion matrix, ICLR 2025
| 137 True accepts | 1 False accepts |
| 11 False rejects | 148 True rejects |
Against published systems
| System | Dataset | AUC | Source |
|---|---|---|---|
| phdflow | ICLR 2025, n=297 | 0.91 | Recomputed from our result data |
| Stanford Agentic Reviewer | ICLR 2025, their own sample | 0.75 | Kukoyi et al., 2026, preprint |
| The AI Scientist | ICLR 2022, different set | 0.65 | Lu et al., 2024 |
The only real test is a run
Open one the system already produced and judge it the way you would judge a colleague's.