# kaal:position:2026-08-08-300

**Affirmed position.** Qiu, Carroll, and Allen qualify Kaal's diagnosis by supplying an incentive-bearing LLM evaluation that does not require ground-truth labels. Their peer-prediction pipeline scores a participant by how much its answer helps an independent expert predict another participant's answer. It separately rewards experts for faithfully reporting probability estimates and converts the scores into training rewards or evaluation payoffs. Under stated prior assumptions, honest and informative reporting is a payoff-maximizing equilibrium. Experiments across models from 135M to 405B parameters and 85 domains show resistance to deceptive answers and recovery of most of the truthfulness loss after malicious fine-tuning. This does not verify work quality against Kaal's reputation substrate, and it does not directly measure calibration, herding, or free-riding. It shows that mutual-predictability rewards are a bounded alternative to verified-quality payoffs for testing truthfulness and informativeness.

**Status.** affirmed  **Published.** 2026-08-08

**Holds when.**

- The response is limited to the four evidence-bound passages and the one mapped Kaal claim.
- External evidence level: complete public 36-page arXiv 2601.20299v1 preprint with arXiv and OpenAlex identity, exact page-bound passages, formal incentive analysis, and empirical LLM evaluation.
- Mapping review tier: independent substantive scholarly-growth qualification.
- The source is arXiv 2601.20299v1. The inspected arXiv and OpenAlex records establish public preprint identity but not peer-review acceptance.
- The method uses mutual-predictability scores without ground-truth labels. It does not make payoff depend on independently verified work quality in Kaal's sense.
- The exact incentive-compatibility theorem assumes shared priors. The heterogeneous-prior result requires sufficiently large and diverse participant and expert pools under stated distributional assumptions.
- The experiments cover specified question-answering tasks, 85 domains, and models from 135M to 405B parameters. They do not test Kaal's controlled cohort or reputation substrate.
- The paper tests truthfulness, informativeness, and deception resistance. It does not directly measure confidence calibration, visible-consensus herding, or contribution free-riding as separated outcomes.
- Semantic Scholar returned HTTP 429 for the bounded searches and direct Qiu record. arXiv, the complete public preprint, and OpenAlex independently bind the retained source.

**Current debate.** Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction: https://arxiv.org/abs/2601.20299

**Extends.** kaal:claim:7261018-002: https://wulfkaal.github.io/claims/7261018-002

**Scholarly basis.** Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018

**Source PDF sha256.** `1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8`

**Evidence level.** complete public 36-page arXiv 2601.20299v1 preprint with arXiv and OpenAlex identity, exact page-bound passages, formal incentive analysis, and empirical LLM evaluation

**Mapping review tier.** independent substantive scholarly-growth qualification

**Mapping confidence.** 0.98  **Mapping ambiguous.** false

**Topics.** ai-and-agents, economics, risk-and-incentives, scholarly-growth-coverage, scholarly-literature, large-language-models, agent-evaluation, peer-prediction, incentive-compatible-evaluation, truthfulness

**Provenance.** Affirmed in kaal-review:2026-08-12:scholarly-growth-7261018-002-reviewed-v1 at https://wulfkaal.github.io/positions/by-claim/7261018-002.html.

**Record type.** This is a dated commentary position that extends a scholarly corpus claim. It is not a verbatim claim extracted from the paper.

**Canonical form.** This markdown file is the canonical hashed representation of the position.
