Qualification: Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
Qiu, Carroll, and Allen qualify Kaal's diagnosis by supplying an incentive-bearing LLM evaluation that does not require ground-truth labels. Their peer-prediction pipeline scores a participant by how much its answer helps an independent expert predict another participant's answer. It separately rewards experts for faithfully reporting probability estimates and converts the scores into training rewards or evaluation payoffs. Under stated prior assumptions, honest and informative reporting is a payoff-maximizing equilibrium. Experiments across models from 135M to 405B parameters and 85 domains show resistance to deceptive answers and recovery of most of the truthfulness loss after malicious fine-tuning. This does not verify work quality against Kaal's reputation substrate, and it does not directly measure calibration, herding, or free-riding. It shows that mutual-predictability rewards are a bounded alternative to verified-quality payoffs for testing truthfulness and informativeness.
ai-and-agentseconomicsrisk-and-incentivesscholarly-growth-coveragescholarly-literaturelarge-language-modelsagent-evaluationpeer-predictionincentive-compatible-evaluationtruthfulness