Architecture
A panel of specialists, not one prompt in a loop
Your paper is not handed to one model with one question. It is taken apart, read in parallel by specialists looking for different kinds of weakness, and judged against the venue you chose.
Inside a run
Why the output is not one opinion
The paper is taken apart before anything reads it
A PDF is a layout format, not a document. Yours is recovered into real structure first: sections, references, the citations in the text. That is why every issue names the place it came from.
Different readers, different failures
The specialists share no prompt and never see each other's answers. Each is set up for one kind of weakness and reports separately, so strength in one respect cannot paper over thinness in another.
The bar belongs to the venue
What counts as good enough is not written into the system. It comes from the guidelines of the venue you picked, so the same paper can clear one and fall short at another.
Prior work is checked live
The literature is searched while your run is in progress, not recalled from a model's memory. When the report says your claim sits close to existing work, it points at papers you can open.
The verdict is assembled, not improvised
It is composed from independent assessments and the evidence behind them, and nothing can enter it that was not found in your paper. If part of the panel fails, the run still finishes and reports what held.
Evidence
Measured against real decisions
Published results on the same task: Stanford Agentic Reviewer 0.75, The AI Scientist 0.65. The confusion matrix behind ours is on the home page.
The only real test is a run
Read one the system produced and judge it the way you would judge a reviewer.