ResearchAug 01, 2026 · 11 min readPeer-reviewed where applicable
Evaluating LLMs on peer review: a benchmark and its limits
We test model assistance in review and show where human judgment still dominates.
Template article — no fake benchmarks or citations are claimed here.
Research pages use the same template as openai.com — title, deck, hero art card, then long-form prose with hairline dividers. This copy is editorial placeholder in a serious-lab voice.
Method
Describe the question, the corpus, and the evaluation. Link to metademic.com where the scholarly record is the source of truth.
Results
Report what was found without inventing numbers. If a claim is not yet verified, say so plainly.
Limitations
Every paper has them. Name them early.