GEM 2026 Outstanding Paper · ACL 2026

Are LLM Benchmarks Already Contaminated? A Systematic Review of Contamination Detection Methods

Erfan Nourbakhsh1, Mohammad Sadegh Sirjani1, Amir Mousavi1, Khoa Nguyen1, John Quarles1, Mimi Xie1, Rocky Slavin1

¹University of Texas at San Antonio (UTSA)

Fifth Workshop on Generation, Evaluation and Metrics (GEM 2026)
colocated with ACL 2026 · San Diego, California, USA

Paper

Abstract

Large Language Models (LLMs) are trained on web-scale corpora, increasing the risk that benchmark test data appears in training sets and inflates reported performance. We present a systematic literature review of 55 studies on LLM benchmark contamination through late 2025. Our contributions are: (1) a four-tier contamination taxonomy (Exact, Syntactic, Semantic, Task-Level; T1-T4); (2) a comparative analysis of five detection families (string-matching, likelihood-based, membership inference, LLM-prompted detection, and benchmark auditing), including access assumptions and failure modes; (3) a synthesis of contamination evidence on MMLU, GSM8K, HumanEval, and HellaSwag by measurement construct; (4) a comparative evaluation of mitigation strategies across lifecycle points, access assumptions, and evidence maturity; and (5) a Contamination Transparency Card (CTC) framework for future releases. Across studies, no detection method is consistently reliable across contamination tiers, model-access settings, and training stages. We identify instruction tuning as a persistent blind spot, note that RL/post-training contamination auditing is only beginning to mature, and report inflation estimates spanning roughly 6%-40% under benchmark- and setting-dependent assumptions.

Framework

Four-Tier Contamination Taxonomy

Mechanism-first tiers map overlap type to expected detectability and mitigation leverage, unifying fragmented prior terminology.

Tier Type Mechanism Representative Evidence Severity Detectability
T1 Exact Verbatim or near-verbatim test instances in the training corpus MMLU 57% option guessing; HumanEval 10-gram collisions Critical High
T2 Syntactic Test data present after surface transformation (paraphrase, shuffle, retokenization) GPT-4-level performance from paraphrased fine-tuning; MMLU-CF lexical leakage High Medium
T3 Semantic Semantically equivalent content without lexical overlap (translations, reformulations) Cross-lingual inflation Moderate Low
T4 Task-level Exposure to task format, reasoning pattern, or domain knowledge without specific test instances GSM1K accuracy drops; zero-shot task contamination Variable Very Low

Table 1. Taxonomy of benchmark contamination in LLM evaluation (T1-T4).

Methods

Contamination Detection Families

Five detection families along the white-box to black-box access spectrum. Probing methods need corpus or log-probability access; prompting methods use outputs alone.

Taxonomy of contamination detection methods from white-box probing to black-box prompting.

Figure 1. Taxonomy of contamination detection methods organized along the white-box to black-box access spectrum. Probing-based methods (left) require corpus or log-probability access; prompting-based methods (right) operate on model outputs alone.

Family Access Best for Main failure modes
String-matching White-box T1 exact overlap audits Weak under paraphrase; misses T3-T4
Likelihood-based Gray-box Open-weight memorization signals Sensitive to format and probability access
Membership inference Gray-box Document / collection scale Often near-random at instance level
LLM-prompted probes Black-box Proprietary API settings Heuristic, gameable, weakly calibrated
Benchmark-level auditing Black / gray Leaderboards and closed models (e.g., ConStat) Needs careful controls and effect-size reporting

Table 2. Family-level synthesis of detection methods surveyed across 55 studies.

Evidence

What Benchmarks Show

Evidence stratified by benchmark and measurement family. Hit rates, likelihood scores, and benchmark-level deltas measure different constructs and are not collapsed into one metric.

Benchmark Selected signals
MMLU GPT-4 TS-Guessing 57% (baseline 25%); ITD drops up to 19.0%; top 7B Open LLM Leaderboard models with ConStat δ̂ > 10%; weaker match on MMLU-CF
GSM8K Up to 8% drop on GSM1K (ρ = 0.60); ConStat δ̂ ≈ 27%-40% for InternLM-2-Math-7B
HumanEval Confirmed 10-gram collisions; post-cutoff drops on LiveCodeBench
HellaSwag / PIQA Flagged by multiple techniques; ConStat effects about 6%-11% on top-ranked 7B models

Table 3. Benchmark-specific contamination signals synthesized in the review. Effect sizes are not directly comparable across methods.

Defense

Mitigation Landscape

Static, inference-time, and dynamic strategies trade coverage, standardization, and cost. No single fix covers all tiers.

Strategy What it does Caveat
Static decontamination n-gram / semantic filters; contamination-free redesigns (e.g., MMLU-CF, Clean-Eval) Scalable baselines miss deeper semantic/task leakage
Inference-time decontamination (ITD) Detect contaminated items at eval time, rewrite, re-score without retraining Model-specific rewritten exams hurt leaderboard comparability
Dynamic benchmarks Rolling refresh / post-cutoff novelty (LiveBench, LiveCodeBench, LatestEval-style) Strongest preventive path for T3/T4; needs continuous generation cost

Table 4. Comparative mitigation strategies across lifecycle points.

Disclosure

Contamination Transparency Card

Minimal five-dimension disclosure framework modelled on Model Cards and Datasheets. Absence of contamination evidence is not evidence of absence.

Dimension Required disclosures
Training Data Pretraining corpus names/versions; cutoff dates; deduplication; total token count
Decontamination Methods and thresholds; datasets checked; known failure modes
Evidence Provided Post-hoc overlap, likelihood, and/or calibrated performance audits (e.g., ConStat, TED); known contaminated benchmarks
IFT Stage Fine-tuning composition; whether benchmark examples included; answer augmentation strategy
Reproducibility Prompt templates; few-shot ordering/count; scoring; multi-run variance; generation params

Table 5. Proposed Contamination Transparency Card (CTC). All dimensions required for benchmark releases; Training Data, Decontamination, and IFT Stage additionally recommended for model technical reports.

Conclusion

Main Takeaways

BibTeX

@inproceedings{nourbakhsh-etal-2026-llm,
    title = "Are {LLM} Benchmarks Already Contaminated? A Systematic Review of Contamination Detection Methods",
    author = "Nourbakhsh, Erfan  and
      Sirjani, Mohammad Sadegh  and
      Mousavi, Amir  and
      Nguyen, Khoa  and
      Quarles, John  and
      Xie, Mimi  and
      Slavin, Rocky",
    booktitle = "Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics ({GEM})",
    month = jul,
    year = "2026",
    address = "San Diego, California, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.gem-main.50/",
    doi = "10.18653/v1/2026.gem-main.50",
    pages = "518--539",
}