Paper
Large Language Models (LLMs) are trained on web-scale corpora, increasing the risk that benchmark test data appears in training sets and inflates reported performance. We present a systematic literature review of 55 studies on LLM benchmark contamination through late 2025. Our contributions are: (1) a four-tier contamination taxonomy (Exact, Syntactic, Semantic, Task-Level; T1-T4); (2) a comparative analysis of five detection families (string-matching, likelihood-based, membership inference, LLM-prompted detection, and benchmark auditing), including access assumptions and failure modes; (3) a synthesis of contamination evidence on MMLU, GSM8K, HumanEval, and HellaSwag by measurement construct; (4) a comparative evaluation of mitigation strategies across lifecycle points, access assumptions, and evidence maturity; and (5) a Contamination Transparency Card (CTC) framework for future releases. Across studies, no detection method is consistently reliable across contamination tiers, model-access settings, and training stages. We identify instruction tuning as a persistent blind spot, note that RL/post-training contamination auditing is only beginning to mature, and report inflation estimates spanning roughly 6%-40% under benchmark- and setting-dependent assumptions.
Framework
Mechanism-first tiers map overlap type to expected detectability and mitigation leverage, unifying fragmented prior terminology.
| Tier | Type | Mechanism | Representative Evidence | Severity | Detectability |
|---|---|---|---|---|---|
| T1 | Exact | Verbatim or near-verbatim test instances in the training corpus | MMLU 57% option guessing; HumanEval 10-gram collisions | Critical | High |
| T2 | Syntactic | Test data present after surface transformation (paraphrase, shuffle, retokenization) | GPT-4-level performance from paraphrased fine-tuning; MMLU-CF lexical leakage | High | Medium |
| T3 | Semantic | Semantically equivalent content without lexical overlap (translations, reformulations) | Cross-lingual inflation | Moderate | Low |
| T4 | Task-level | Exposure to task format, reasoning pattern, or domain knowledge without specific test instances | GSM1K accuracy drops; zero-shot task contamination | Variable | Very Low |
Table 1. Taxonomy of benchmark contamination in LLM evaluation (T1-T4).
Methods
Five detection families along the white-box to black-box access spectrum. Probing methods need corpus or log-probability access; prompting methods use outputs alone.
Figure 1. Taxonomy of contamination detection methods organized along the white-box to black-box access spectrum. Probing-based methods (left) require corpus or log-probability access; prompting-based methods (right) operate on model outputs alone.
| Family | Access | Best for | Main failure modes |
|---|---|---|---|
| String-matching | White-box | T1 exact overlap audits | Weak under paraphrase; misses T3-T4 |
| Likelihood-based | Gray-box | Open-weight memorization signals | Sensitive to format and probability access |
| Membership inference | Gray-box | Document / collection scale | Often near-random at instance level |
| LLM-prompted probes | Black-box | Proprietary API settings | Heuristic, gameable, weakly calibrated |
| Benchmark-level auditing | Black / gray | Leaderboards and closed models (e.g., ConStat) | Needs careful controls and effect-size reporting |
Table 2. Family-level synthesis of detection methods surveyed across 55 studies.
Evidence
Evidence stratified by benchmark and measurement family. Hit rates, likelihood scores, and benchmark-level deltas measure different constructs and are not collapsed into one metric.
| Benchmark | Selected signals |
|---|---|
| MMLU | GPT-4 TS-Guessing 57% (baseline 25%); ITD drops up to 19.0%; top 7B Open LLM Leaderboard models with ConStat δ̂ > 10%; weaker match on MMLU-CF |
| GSM8K | Up to 8% drop on GSM1K (ρ = 0.60); ConStat δ̂ ≈ 27%-40% for InternLM-2-Math-7B |
| HumanEval | Confirmed 10-gram collisions; post-cutoff drops on LiveCodeBench |
| HellaSwag / PIQA | Flagged by multiple techniques; ConStat effects about 6%-11% on top-ranked 7B models |
Table 3. Benchmark-specific contamination signals synthesized in the review. Effect sizes are not directly comparable across methods.
Defense
Static, inference-time, and dynamic strategies trade coverage, standardization, and cost. No single fix covers all tiers.
| Strategy | What it does | Caveat |
|---|---|---|
| Static decontamination | n-gram / semantic filters; contamination-free redesigns (e.g., MMLU-CF, Clean-Eval) | Scalable baselines miss deeper semantic/task leakage |
| Inference-time decontamination (ITD) | Detect contaminated items at eval time, rewrite, re-score without retraining | Model-specific rewritten exams hurt leaderboard comparability |
| Dynamic benchmarks | Rolling refresh / post-cutoff novelty (LiveBench, LiveCodeBench, LatestEval-style) | Strongest preventive path for T3/T4; needs continuous generation cost |
Table 4. Comparative mitigation strategies across lifecycle points.
Disclosure
Minimal five-dimension disclosure framework modelled on Model Cards and Datasheets. Absence of contamination evidence is not evidence of absence.
| Dimension | Required disclosures |
|---|---|
| Training Data | Pretraining corpus names/versions; cutoff dates; deduplication; total token count |
| Decontamination | Methods and thresholds; datasets checked; known failure modes |
| Evidence Provided | Post-hoc overlap, likelihood, and/or calibrated performance audits (e.g., ConStat, TED); known contaminated benchmarks |
| IFT Stage | Fine-tuning composition; whether benchmark examples included; answer augmentation strategy |
| Reproducibility | Prompt templates; few-shot ordering/count; scoring; multi-run variance; generation params |
Table 5. Proposed Contamination Transparency Card (CTC). All dimensions required for benchmark releases; Training Data, Decontamination, and IFT Stage additionally recommended for model technical reports.
Conclusion
@inproceedings{nourbakhsh-etal-2026-llm,
title = "Are {LLM} Benchmarks Already Contaminated? A Systematic Review of Contamination Detection Methods",
author = "Nourbakhsh, Erfan and
Sirjani, Mohammad Sadegh and
Mousavi, Amir and
Nguyen, Khoa and
Quarles, John and
Xie, Mimi and
Slavin, Rocky",
booktitle = "Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics ({GEM})",
month = jul,
year = "2026",
address = "San Diego, California, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.gem-main.50/",
doi = "10.18653/v1/2026.gem-main.50",
pages = "518--539",
}