Skip to main content
Watch: I Asked Claude to Build Me a Business

Evaluation / Agents / Training

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Sy-Tuyen Ho, Minghui Liu, Furong Huang

arXiv:2609.209425 upvotes

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Authors: Sy-Tuyen Ho, Minghui Liu, Furong Huang

arXiv ID: 2609.20942

Problem: LLMs increasingly act as automated reviewers and assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review becomes recursive: later reviewers learn from judgments produced by earlier models. The effects of this single feedback step had not been studied in a controlled setting.

Key Methodology:

  • Start from Llama 3.1 8B fine-tuned as a reviewer on official ICLR reviews from 2018-2023
  • Train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews
  • Measure rating distribution compression and same-paper plus corpus-level semantic diversity
  • TrustReviewer mitigation: single-stage training on a curated corpus designed to reduce low-quality and semantically degenerate supervision (training-time), plus paired activation steering for test-time correction with no further training or expert annotation

Key Results:

  • Introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity - the authors name this pattern scientific-judgment collapse
  • TrustReviewer intervenes at both stages: curated-corpus training prevents propagation of low-quality supervision, and paired activation steering mitigates residual collapsed-judgment tendencies

What it means for developers: Model collapse has a judgment twin: any pipeline that feeds generated reviews back into training corpora flattens the judgment distribution of later reviewers, which quietly erodes rating spread and semantic diversity in evaluation systems. The cheap reverses are corpus curation and activation steering rather than more annotation. For anyone building judge or eval pipelines, rating-distribution spread becomes a health metric to monitor, and synthetic-review mixtures are a contamination class that must be tracked alongside benchmark leakage.

Paper: arXiv:2609.20942