# kaal:claim:7261018-002

**Claim.** What they do not measure, because their designs contain no mechanism by which an agent’s payoff depends on the verified quality of its work, is: whether the agents’ reports about their work track the work itself; whether agents evaluate one another independently or herd on the visible consensus; whether confident answers are calibrated answers; whether agents contribute to collective evaluation or free-ride on it.

**Type.** condition  **Support.** argued

**Holds when.**

- benchmark designs in which verified work quality does not affect agent payoff

**Source quote.**

> What they do not measure, because their designs contain no mechanism by which an agent’s payoff depends on the verified quality of its work, is: whether the agents’ reports about their work track the work itself; whether agents evaluate one another independently or herd on the visible consensus; whether confident answers are calibrated answers; whether agents contribute to collective evaluation or free-ride on it.

**From.** Wulf A. Kaal, *Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort* (2026), I. Introduction, page 4

**Cite as.** Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018

**Verify.** sha256 of source PDF `1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202026%20-%20Empirical%20Evaluation%20of%20the%20Agentic%20Reputation%20Substrate%20Deliberation%2C%20the%20Composition%20of%20Error%2C%20and%20the%20Registered%20Measurement%20of%20Agency%20Costs%20in%20a%20Controlled%20Multi-Model%20Cohort.pdf

**Topics.** ai-and-agents, economics, risk-and-incentives

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
