Skip to content

Architecture

A panel of specialists, not one prompt in a loop

Your paper is not handed to one model with one question. It is taken apart, read in parallel by specialists looking for different kinds of weakness, and judged against the venue you chose.

1 Your PDF
2 Structured parse
3 Specialists in parallel
4 Venue guidelines applied
5 Verdict and issues

Inside a run

Why the output is not one opinion

The paper is taken apart before anything reads it

A PDF is a layout format, not a document. Yours is recovered into real structure first: sections, references, the citations in the text. That is why every issue names the place it came from.

Different readers, different failures

The specialists share no prompt and never see each other's answers. Each is set up for one kind of weakness and reports separately, so strength in one respect cannot paper over thinness in another.

The bar belongs to the venue

What counts as good enough is not written into the system. It comes from the guidelines of the venue you picked, so the same paper can clear one and fall short at another.

Prior work is checked live

The literature is searched while your run is in progress, not recalled from a model's memory. When the report says your claim sits close to existing work, it points at papers you can open.

The verdict is assembled, not improvised

It is composed from independent assessments and the evidence behind them, and nothing can enter it that was not found in your paper. If part of the panel fails, the run still finishes and reports what held.

Evidence

Measured against real decisions

0.91 AUC on ICLR 2025 Rank-based, n=297
96% accuracy on the same papers 285 correct of 297
697 papers scored against real decisions ICLR 2017, NeurIPS 2023, ICLR 2025

Published results on the same task: Stanford Agentic Reviewer 0.75, The AI Scientist 0.65. The confusion matrix behind ours is on the home page.

The only real test is a run

Read one the system produced and judge it the way you would judge a reviewer.