kaal:claim:4855607-015

Reward modeling learned through interaction with users carries two structural pathologies: majority views disproportionately influence the learned reward function, and the agent may engage in reward hacking.

Source quote, verbatim
but it comes with potential issues such as the prevalence of majority views disproportionately influencing the learned reward function and the risk of reward hacking.
From

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024), Model Overview: Reinforcement Learning (RL), p. 27
https://ssrn.com/abstract=4855607 · source PDF

Cite as

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

Holds when
Classification

failuresupport: evidencedfailure: Majority capture of the reward modelfamily: ai-oversight-and-alignment-gapai-and-agents

Verify

The quote above is an exact substring of the source PDF, whose sha256 is eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/61a867481c493ba4...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/4855607-015.md | sha256sum