{
 "@context": "https://schema.org",
 "@type": "Claim",
 "@id": "https://wulfkaal.github.io/positions/2026-08-26-024",
 "identifier": "kaal:position:2026-08-26-024",
 "additionalType": "https://wulfkaal.github.io/positions/schema.json#AffirmedPositionClaim",
 "name": "Evaluation Needs Result Level Fidelity And Resistance",
 "text": "Evaluation is an evidence claim before it is a score.\n\nSun and coauthors provide a useful test of this distinction in benchmark contamination. They evaluate ten language models, five benchmarks, twenty mitigation strategies, and two contamination regimes. Their unit of analysis is the question-level result. This matters because the same aggregate accuracy can conceal different patterns of correct and incorrect answers. A matched score therefore does not establish that an updated benchmark preserved the original task or measured the same capability.\n\nThe paper separates fidelity from contamination resistance. Fidelity asks whether a clean model produces the same question-level results after the benchmark is changed. Resistance asks whether exposure to the original benchmark still supplies an advantage. These conditions cannot be collapsed into one scalar measure. Semantic-preserving changes often retained the task but failed to improve resistance consistently. Semantic-altering changes could increase resistance by changing question difficulty, scope, or the evaluation objective. No tested strategy achieved strong fidelity and resistance across all reported benchmarks.\n\nThis evidence qualifies the institutional requirement for evaluation. Results should be recorded at the unit where the relevant capability is defined, together with the model, task, benchmark version, and contamination condition. Context-specific comparison should preserve the evaluation objective. Resistance to manipulation should be tested against a stated threat model. It cannot be inferred from a mitigation label, a hidden procedure, or an aggregate score.\n\nThe evidence remains limited. The article studies benchmark contamination, not runtime reputation, self-declared capability, or engagement metrics. It does not establish that every evaluation signal can be made costly to manipulate. It demonstrates a narrower point. Aggregation can conceal changed task performance, and apparent resistance can be purchased by sacrificing fidelity.\n\nFor sovereign agent coordination, an evaluation record should expose result-level evidence, stated conditions, and separate fidelity and resistance tests. Otherwise a score may reward success against a changed instrument rather than success in the stated context.",
 "author": {
  "@type": "Person",
  "name": "Wulf A. Kaal",
  "identifier": "https://orcid.org/0009-0008-7840-1847"
 },
 "datePublished": "2026-08-26",
 "dateModified": "2026-08-26",
 "creativeWorkStatus": "Affirmed",
 "responseType": "qualification",
 "keywords": [
  "institutional-design",
  "governance-design",
  "evaluation",
  "reputation",
  "benchmarking",
  "measurement",
  "manipulation-resistance",
  "research-methods"
 ],
 "scope_conditions": [
  "The response is limited to the exact full-text propositions and the one mapped Kaal claim.",
  "External evidence level: peer-reviewed conference paper with complete official proceedings full text and controlled experiments.",
  "Mapping review tier: independent substantive scholarly-growth qualification.",
  "The source studies LLM benchmark contamination rather than reputation in a deployed sovereign agent runtime.",
  "Its evaluation vector records binary question correctness and does not cover every form of task outcome.",
  "The experiments cover ten models, five benchmarks, twenty mitigation strategies, and two contamination recipes rather than all evaluation settings.",
  "Contamination resistance addresses advantage from benchmark exposure, not every form of collusion, bribery, identity fraud, or engagement gaming.",
  "The source does not evaluate self-declared capability or prove that every useful signal can be made costly to manipulate."
 ],
 "currentDebate": {
  "name": "The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination",
  "url": "https://proceedings.mlr.press/v267/sun25t.html"
 },
 "extends": {
  "identifier": "kaal:claim:7314479-024",
  "url": "https://wulfkaal.github.io/claims/7314479-024",
  "citation": "Wulf A. Kaal, Institutional Requirements for Sovereign Local Agent Runtimes (2026). SSRN: https://ssrn.com/abstract=7314479",
  "paper": "Wulf A. Kaal, Institutional Requirements for Sovereign Local Agent Runtimes",
  "authors": [
   "Wulf A. Kaal"
  ],
  "year": "2026",
  "ssrn": "https://ssrn.com/abstract=7314479",
  "source_pdf_sha256": "debace24a155ae924a155b1fafe98856d98cf83689feff2f87a32f1c06171ce6"
 },
 "isBasedOn": [
  {
   "@id": "https://wulfkaal.github.io/claims/7314479-024"
  },
  {
   "@type": "CreativeWork",
   "name": "The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination",
   "url": "https://proceedings.mlr.press/v267/sun25t.html"
  }
 ],
 "batch_id": "kaal-review:2026-08-26:scholarly-growth-7314479-024-reviewed-v1",
 "review_provenance": "https://wulfkaal.github.io/positions/by-claim/7314479-024.html",
 "publicationStatus": "public",
 "recordTypeNote": "Dated commentary position extending a scholarly corpus claim. Not a verbatim claim extracted from the paper.",
 "isPartOf": {
  "@id": "https://wulfkaal.github.io/positions/index.json"
 },
 "version": "1.0",
 "canonical_url": "https://wulfkaal.github.io/positions/2026-08-26-024",
 "canonicalForm": "https://wulfkaal.github.io/positions/2026-08-26-024.md",
 "candidateId": "kaal:response-candidate:2026-08-26:scholarly-growth-7314479-024-evaluation-needs-result-level-fidelity-and-resistance-01",
 "evidenceLevel": "peer-reviewed conference paper with complete official proceedings full text and controlled experiments",
 "reviewTier": "independent substantive scholarly-growth qualification",
 "mappingConfidence": 0.98,
 "mappingAmbiguous": false,
 "mappingMethod": "independent substantive scholarly-growth one-to-one qualification review",
 "mappingWhyRelevant": "The source independently tests the exact institutional distinction between a summary score and evidence tied to stated conditions. Question-level result vectors reveal when matched aggregate accuracy conceals changed capability measurement. Separate fidelity and contamination-resistance tests make the context and manipulation threat explicit. The mapping is a qualification because the study is limited to benchmark contamination and does not test self-declared capability or engagement metrics.",
 "sourceProvenance": {
  "source": "ICML 2025 peer-reviewed proceedings paper with complete official full text",
  "sourceRecordId": "pmlr:v267:sun25t",
  "canonicalUrl": "https://proceedings.mlr.press/v267/sun25t.html",
  "publicFullTextUrl": "https://raw.githubusercontent.com/mlresearch/v267/main/assets/sun25t/sun25t.pdf",
  "retrievedAt": "2026-08-27T13:11:43.530Z",
  "fullTextPdfSha256": "3059bc2ee141055029ac3a9257ed6ec762d73ae9602df36ee216d3f86404146a",
  "extractedTextSha256": "edb4001025a6a4650680363c73f13bb9a7f5e874fc601c804b29480a6d710151",
  "officialProceedingsRecordSha256": "00281a3c9c862b2421f5436cc81d40a976a1ede83826e50aad58b246b1abba24",
  "primaryEvidenceReceiptSha256": "ef870e0d6c83918760d5d0120ea45510c74a310de3cd0289f95ef9d25e69d752",
  "sourceProposition": "Sun and coauthors show that matched aggregate accuracy can conceal different question-level outcomes and test benchmark updates by separate fidelity and contamination-resistance measures under declared model, benchmark, and contamination conditions.",
  "sourcePropositionSha256": "75af5f97ec7b1d1797cd3fde682266aae48b5aa15f00a1c82930cdb637e0bc26",
  "sourceEvidenceSetSha256": "e30866b9f00bf4ee38074e8f4f766ac5eae21b7848be4fb7214f20cacd143cba",
  "sourceEvidencePassages": [
   {
    "text": "even when the accuracies match, the question-level evaluation results differ significantly",
    "locator": {
     "publication": "Proceedings of the 42nd International Conference on Machine Learning",
     "pdfPage": 2,
     "section": "1 Introduction"
    },
    "sha256": "cde5d68159ad5aa0dc1254583a02dca67f893dfdc3f5166c8544545846b3b9a7"
   },
   {
    "text": "Focusing solely on matching the scalar accuracy is not sufficient and can even be misleading.",
    "locator": {
     "publication": "Proceedings of the 42nd International Conference on Machine Learning",
     "pdfPage": 2,
     "section": "1 Introduction"
    },
    "sha256": "a7767fa3099818a47135972efe562a9c0ad9531fa2df59f3e2872e3472b635cf"
   },
   {
    "text": "This evaluation vector is a critical component of our framework, as it captures the model’s performance on the benchmark at a question-by-question level.",
    "locator": {
     "publication": "Proceedings of the 42nd International Conference on Machine Learning",
     "pdfPage": 4,
     "section": "3 Methodology"
    },
    "sha256": "7cb1425aac2fe050445ca8991c0f1fe8f8fd1280859a09225a391d1ae3bd918d"
   },
   {
    "text": "A contamination-resistant strategy ensures that an LLM does not gain any advantage on the updated benchmark from being exposed to the original benchmark.",
    "locator": {
     "publication": "Proceedings of the 42nd International Conference on Machine Learning",
     "pdfPage": 4,
     "section": "3 Methodology"
    },
    "sha256": "384335a713ef414ac298d63d7126e8b584c66f40167cbeb0f4a1f50a6301b023"
   },
   {
    "text": "However, no existing strategy effectively achieves this balance.",
    "locator": {
     "publication": "Proceedings of the 42nd International Conference on Machine Learning",
     "pdfPage": 8,
     "section": "5 Results"
    },
    "sha256": "636c342095afc5ee4d0654364038831a137d0b87ad67a83aa839f2cba358fae9"
   }
  ],
  "workId": "work:pmlr:v267:sun25t",
  "workAuthors": [
   "Yifan Sun",
   "Han Wang",
   "Dongbai Li",
   "Gang Wang",
   "Huan Zhang"
  ],
  "workPublishedAt": "2025-07",
  "identityKeys": [
   "pmlr:v267:sun25t",
   "pdf:3059bc2ee141055029ac3a9257ed6ec762d73ae9602df36ee216d3f86404146a",
   "proposition:75af5f97ec7b1d1797cd3fde682266aae48b5aa15f00a1c82930cdb637e0bc26"
  ],
  "claimMappings": [
   {
    "claimId": "kaal:claim:7314479-024",
    "claimUrl": "https://wulfkaal.github.io/claims/7314479-024",
    "rank": 1,
    "confidence": 0.98,
    "method": "independent substantive scholarly-growth one-to-one qualification review",
    "whyRelevant": "The source independently tests the exact institutional distinction between a summary score and evidence tied to stated conditions. Question-level result vectors reveal when matched aggregate accuracy conceals changed capability measurement. Separate fidelity and contamination-resistance tests make the context and manipulation threat explicit. The mapping is a qualification because the study is limited to benchmark contamination and does not test self-declared capability or engagement metrics.",
    "ambiguous": false
   }
  ],
  "substantiveReview": {
   "reviewedAt": "2026-08-27T13:11:43.530Z",
   "sourceIdentityVerified": true,
   "authorIndependenceVerified": true,
   "kaalReferenceFoundInSource": false,
   "temporalIndependence": "The paper was published in 2025, before Kaal's 2026 paper.",
   "canonicalPublicStatusVerified": true,
   "peerReviewedStatusVerified": true,
   "evidenceClassification": "peer-reviewed controlled benchmark-contamination study",
   "retractionOrSupersessionFound": false,
   "propositionFidelityVerified": true,
   "mechanismCorrespondence": "question-level observed result vectors are compared under explicit clean, contaminated, original, and updated benchmark conditions using separate fidelity and contamination-resistance measures",
   "compatibleScope": "evaluation integrity under benchmark contamination, limited because the paper does not study production runtime reputation or all strategic manipulation channels",
   "responseWordingDefensible": true,
   "oneToOneExtendsMapping": true,
   "exactSupportingQuotesVerified": true,
   "nonOverlap": {
    "candidateIdMatches": false,
    "canonicalUrlMatches": false,
    "propositionHashMatches": false,
    "priorPositionForClaim": false
   },
   "limitations": [
    "The source studies LLM benchmark contamination rather than reputation in a deployed sovereign agent runtime.",
    "Its evaluation vector records binary question correctness and does not cover every form of task outcome.",
    "The experiments cover ten models, five benchmarks, twenty mitigation strategies, and two contamination recipes rather than all evaluation settings.",
    "Contamination resistance addresses advantage from benchmark exposure, not every form of collusion, bribery, identity fraud, or engagement gaming.",
    "The source does not evaluate self-declared capability or prove that every useful signal can be made costly to manipulate."
   ],
   "rejectionReasonsRecorded": true
  },
  "contentMap": {
   "proposition": "Evaluation requires result-level evidence under stated conditions rather than an unqualified aggregate score.",
   "evidenceLayer": "peer-reviewed conference paper with complete official proceedings full text and controlled experiments",
   "strongestLimitation": "The source tests benchmark contamination, not runtime reputation or every strategic manipulation channel.",
   "consequence": "Matched aggregate performance can conceal a changed evaluation objective, while apparent resistance can be purchased by sacrificing fidelity.",
   "requestedAction": "Publish result-level outcomes, conditions, and separate fidelity and resistance tests before using a score for agent coordination."
  },
  "stylePack": {
   "profile": "M1 early sole-author baseline v1.2.0",
   "verifiedProfileWorks": [
    "1428387",
    "1998455",
    "2150377",
    "2267560"
   ],
   "passageCount": 4,
   "rhetoricalFunctions": [
    "classification before inference",
    "empirical interpretation",
    "limitation",
    "institutional consequence"
   ],
   "sameRegisterPassagePackAvailable": true,
   "limitation": "The short public position permits only bounded stylometric comparison."
  },
  "m1Validation": {
   "status": "M1-PASS-WITH-LIMITS",
   "deterministicGate": "pass",
   "hardFailures": 0,
   "warnings": 0,
   "words": 314,
   "reason": "The publication-bound position passed strict and public deterministic controls against a task-local four-work style pack. Its short length limits stylometric comparison."
  }
 },
 "userAffirmation": "Authorized under public authority SHA-256 87aad20196a753015a36d970f742c885eb763efdbada4869949bfffe3298130c and event supersession SHA-256 7d47ef36085c4dce590f287c986e4106f3bf35a7da5a25322d6fc3d4abf456d4. Publication remains receipt-bound to successful workflows and exact live-byte verification.",
 "sha256": "478cbac82fbaf60bb94682228d675c4682405bce2a1b047221ec213f16cc7401"
}
