Qualification: Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction

Record: kaal:position:2026-08-08-300 · 2026-08-08

Qiu, Carroll, and Allen qualify Kaal's diagnosis by supplying an incentive-bearing LLM evaluation that does not require ground-truth labels. Their peer-prediction pipeline scores a participant by how much its answer helps an independent expert predict another participant's answer. It separately rewards experts for faithfully reporting probability estimates and converts the scores into training rewards or evaluation payoffs. Under stated prior assumptions, honest and informative reporting is a payoff-maximizing equilibrium. Experiments across models from 135M to 405B parameters and 85 domains show resistance to deceptive answers and recovery of most of the truthfulness loss after malicious fine-tuning. This does not verify work quality against Kaal's reputation substrate, and it does not directly measure calibration, herding, or free-riding. It shows that mutual-predictability rewards are a bounded alternative to verified-quality payoffs for testing truthfulness and informativeness.

Affirmed commentary position. This record extends a source-bound scholarly claim but is not a verbatim paper claim.
Holds when
Current debate

Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction

Scholarly basis

kaal:claim:7261018-002
Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018
Source PDF sha256: 1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8

Evidence and mapping

Evidence: complete public 36-page arXiv 2601.20299v1 preprint with arXiv and OpenAlex identity, exact page-bound passages, formal incentive analysis, and empirical LLM evaluation
Review tier: independent substantive scholarly-growth qualification
Mapping confidence: 0.98
Mapping ambiguous: false

Topics

ai-and-agentseconomicsrisk-and-incentivesscholarly-growth-coveragescholarly-literaturelarge-language-modelsagent-evaluationpeer-predictionincentive-compatible-evaluationtruthfulness

Provenance

Affirmed in kaal-review:2026-08-12:scholarly-growth-7261018-002-reviewed-v1 on 2026-08-08. Review record.

Verify

Canonical markdown sha256: d5753c33d94945d4d92d4cf4894b10575aff4242c6e37a5c5e6e5e347cb30433
curl -s https://wulfkaal.github.io/positions/2026-08-08-300.md | sha256sum