kaal:claim:7261018-002

What they do not measure, because their designs contain no mechanism by which an agent’s payoff depends on the verified quality of its work, is: whether the agents’ reports about their work track the work itself; whether agents evaluate one another independently or herd on the visible consensus; whether confident answers are calibrated answers; whether agents contribute to collective evaluation or free-ride on it.

Source quote, verbatim
What they do not measure, because their designs contain no mechanism by which an agent’s payoff depends on the verified quality of its work, is: whether the agents’ reports about their work track the work itself; whether agents evaluate one another independently or herd on the visible consensus; whether confident answers are calibrated answers; whether agents contribute to collective evaluation or free-ride on it.
From

Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026), I. Introduction, p. 4
https://ssrn.com/abstract=7261018 · source PDF

Cite as

Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018

Holds when
Classification

conditionsupport: arguedai-and-agentseconomicsrisk-and-incentives

Verify

The quote above is an exact passage from the source PDF, whose sha256 is 1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8.
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/7261018-002.md | sha256sum

Positions extending this scholarly claim