Skip to content

Architecture

A panel of specialists, not one prompt in a loop

Your paper is not handed to one model with one question. It is taken apart, read in parallel by specialists looking for different kinds of weakness, and judged against the venue you chose.

1 Your PDF
2 Structured parse
3 Specialists in parallel
4 Venue guidelines applied
5 Verdict and issues

Inside a run

Why the output is not one opinion

The paper is taken apart before anything reads it

A PDF is a layout format, not a document. Yours is recovered into real structure first: sections, references, the citations in the text. Every reader works from that structure rather than from a wall of extracted text.

Different readers, different failures

The specialists share no prompt and never see each other's answers. Each is set up for one kind of weakness and reports separately, so strength in one respect cannot paper over thinness in another.

The bar belongs to the venue

What counts as good enough is not written into the system. It comes from the guidelines of the venue you picked, so the same paper can clear one and fall short at another.

Prior work is checked live

The literature is searched while your run is in progress, not recalled from a model's memory, so a claim that your work sits close to something existing was checked against the record rather than remembered.

The verdict is assembled, not improvised

It is composed from independent assessments and the evidence behind them, and nothing can enter it that was not found in your paper. If part of the panel fails, the run still finishes and reports what held.

And what a panel like this cannot do

Evidence

Measured against real decisions

0.91 AUC on ICLR 2025 Rank-based, n=297
96% accuracy on the same papers 285 correct of 297
99% precision when it says accept 137 of 138 predicted accepts
697 papers scored against real decisions ICLR 2017, NeurIPS 2023, ICLR 2025
VenuePapersAccuracyAUC
ICLR 201720385.2%0.84
NeurIPS 202319791.9%0.87
ICLR 202529796.0%0.91
All venues69791.7%0.87

Across all 697: precision 91.1%, recall 92.2%, F1 91.7%. Ground truth is each venue's own decision. The two conferences in this table do not draw the line in the same place, and we measured the gap: ICLR vs NeurIPS.

Confusion matrix, ICLR 2025

Confusion matrix for ICLR 2025, 297 papers
137 True accepts1 False accepts
11 False rejects148 True rejects

Against published systems

SystemDatasetAUCSource
phdflowICLR 2025, n=2970.91Recomputed from our result data
Stanford Agentic ReviewerICLR 2025, their own sample0.75Kukoyi et al., 2026, preprint
The AI ScientistICLR 2022, different set0.65Lu et al., 2024

The only real test is a run

Open one the system already produced and judge it the way you would judge a colleague's.

See a real exploration · See a real match