Skip to content
MetademicMETADEMICRESEARCH LAB
ResearchAug 01, 2026 · 11 min readPeer-reviewed where applicable

Evaluating LLMs on peer review: a benchmark and its limits

We test model assistance in review and show where human judgment still dominates.

Template article — no fake benchmarks or citations are claimed here.

Research pages use the same template as openai.com — title, deck, hero art card, then long-form prose with hairline dividers. This copy is editorial placeholder in a serious-lab voice.

Method

Describe the question, the corpus, and the evaluation. Link to metademic.com where the scholarly record is the source of truth.

Results

Report what was found without inventing numbers. If a claim is not yet verified, say so plainly.

Limitations

Every paper has them. Name them early.