Qualification: The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
Evaluation is an evidence claim before it is a score. Sun and coauthors provide a useful test of this distinction in benchmark contamination. They evaluate ten language models, five benchmarks, twenty mitigation strategies, and two contamination regimes. Their unit of analysis is the question-level result. This matters because the same aggregate accuracy can conceal different patterns of correct and incorrect answers. A matched score therefore does not establish that an updated benchmark preserved the original task or measured the same capability. The paper separates fidelity from contamination resistance. Fidelity asks whether a clean model produces the same question-level results after the benchmark is changed. Resistance asks whether exposure to the original benchmark still supplies an advantage. These conditions cannot be collapsed into one scalar measure. Semantic-preserving changes often retained the task but failed to improve resistance consistently. Semantic-altering changes could increase resistance by changing question difficulty, scope, or the evaluation objective. No tested strategy achieved strong fidelity and resistance across all reported benchmarks. This evidence qualifies the institutional requirement for evaluation. Results should be recorded at the unit where the relevant capability is defined, together with the model, task, benchmark version, and contamination condition. Context-specific comparison should preserve the evaluation objective. Resistance to manipulation should be tested against a stated threat model. It cannot be inferred from a mitigation label, a hidden procedure, or an aggregate score. The evidence remains limited. The article studies benchmark contamination, not runtime reputation, self-declared capability, or engagement metrics. It does not establish that every evaluation signal can be made costly to manipulate. It demonstrates a narrower point. Aggregation can conceal changed task performance, and apparent resistance can be purchased by sacrificing fidelity. For sovereign agent coordination, an evaluation record should expose result-level evidence, stated conditions, and separate fidelity and resistance tests. Otherwise a score may reward success against a changed instrument rather than success in the stated context.
institutional-designgovernance-designevaluationreputationbenchmarkingmeasurementmanipulation-resistanceresearch-methods