# kaal:position:2026-08-26-024

**Affirmed position.** Evaluation is an evidence claim before it is a score.

Sun and coauthors provide a useful test of this distinction in benchmark contamination. They evaluate ten language models, five benchmarks, twenty mitigation strategies, and two contamination regimes. Their unit of analysis is the question-level result. This matters because the same aggregate accuracy can conceal different patterns of correct and incorrect answers. A matched score therefore does not establish that an updated benchmark preserved the original task or measured the same capability.

The paper separates fidelity from contamination resistance. Fidelity asks whether a clean model produces the same question-level results after the benchmark is changed. Resistance asks whether exposure to the original benchmark still supplies an advantage. These conditions cannot be collapsed into one scalar measure. Semantic-preserving changes often retained the task but failed to improve resistance consistently. Semantic-altering changes could increase resistance by changing question difficulty, scope, or the evaluation objective. No tested strategy achieved strong fidelity and resistance across all reported benchmarks.

This evidence qualifies the institutional requirement for evaluation. Results should be recorded at the unit where the relevant capability is defined, together with the model, task, benchmark version, and contamination condition. Context-specific comparison should preserve the evaluation objective. Resistance to manipulation should be tested against a stated threat model. It cannot be inferred from a mitigation label, a hidden procedure, or an aggregate score.

The evidence remains limited. The article studies benchmark contamination, not runtime reputation, self-declared capability, or engagement metrics. It does not establish that every evaluation signal can be made costly to manipulate. It demonstrates a narrower point. Aggregation can conceal changed task performance, and apparent resistance can be purchased by sacrificing fidelity.

For sovereign agent coordination, an evaluation record should expose result-level evidence, stated conditions, and separate fidelity and resistance tests. Otherwise a score may reward success against a changed instrument rather than success in the stated context.

**Status.** affirmed  **Published.** 2026-08-26

**Holds when.**

- The response is limited to the exact full-text propositions and the one mapped Kaal claim.
- External evidence level: peer-reviewed conference paper with complete official proceedings full text and controlled experiments.
- Mapping review tier: independent substantive scholarly-growth qualification.
- The source studies LLM benchmark contamination rather than reputation in a deployed sovereign agent runtime.
- Its evaluation vector records binary question correctness and does not cover every form of task outcome.
- The experiments cover ten models, five benchmarks, twenty mitigation strategies, and two contamination recipes rather than all evaluation settings.
- Contamination resistance addresses advantage from benchmark exposure, not every form of collusion, bribery, identity fraud, or engagement gaming.
- The source does not evaluate self-declared capability or prove that every useful signal can be made costly to manipulate.

**Current debate.** The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination: https://proceedings.mlr.press/v267/sun25t.html

**Extends.** kaal:claim:7314479-024: https://wulfkaal.github.io/claims/7314479-024

**Scholarly basis.** Wulf A. Kaal, Institutional Requirements for Sovereign Local Agent Runtimes (2026). SSRN: https://ssrn.com/abstract=7314479

**Source PDF sha256.** `debace24a155ae924a155b1fafe98856d98cf83689feff2f87a32f1c06171ce6`

**Evidence level.** peer-reviewed conference paper with complete official proceedings full text and controlled experiments

**Mapping review tier.** independent substantive scholarly-growth qualification

**Mapping confidence.** 0.98  **Mapping ambiguous.** false

**Topics.** institutional-design, governance-design, evaluation, reputation, benchmarking, measurement, manipulation-resistance, research-methods

**Provenance.** Affirmed in kaal-review:2026-08-26:scholarly-growth-7314479-024-reviewed-v1 at https://wulfkaal.github.io/positions/by-claim/7314479-024.html.

**Record type.** This is a dated commentary position that extends a scholarly corpus claim. It is not a verbatim claim extracted from the paper.

**Canonical form.** This markdown file is the canonical hashed representation of the position.
