kaal:claim:4855607-015
Reward modeling learned through interaction with users carries two structural pathologies: majority views disproportionately influence the learned reward function, and the agent may engage in reward hacking.
Source quote, verbatim
but it comes with potential issues such as the prevalence of majority views disproportionately influencing the learned reward function and the risk of reward hacking.
From
Cite as
Holds when
Classification
failuresupport: evidencedfailure: Majority capture of the reward modelfamily: ai-oversight-and-alignment-gapai-and-agents
Verify
Positions extending this scholarly claim