# kaal:claim:4855607-015

**Claim.** Reward modeling learned through interaction with users carries two structural pathologies: majority views disproportionately influence the learned reward function, and the agent may engage in reward hacking.

**Type.** failure  **Support.** evidenced

**Holds when.**

- reward functions learned from aggregated user interaction

**Source quote.**

> but it comes with potential issues such as the prevalence of majority views disproportionately influencing the learned reward function and the risk of reward hacking.

**From.** Wulf A. Kaal, *How AI Models are Optimized Through Web3 Governance* (2024), Model Overview: Reinforcement Learning (RL), page 27

**Cite as.** Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

**Verify.** sha256 of source PDF `eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202024%20-%20How%20AI%20Models%20are%20Optimized%20Through%20Web3%20Governance.pdf

**Failure mode.** Majority capture of the reward model  (family: ai-oversight-and-alignment-gap)

**Topics.** ai-and-agents

**Keywords.** reward-modeling, majority-bias, reward-hacking, generative-ai, preference-aggregation

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
