kaal:claim:4855607-018

RLHF fails on several fronts at once: humans can pursue harmful goals either innocently or maliciously, human feedback degrades when examples are hard to evaluate and especially when RLHF is applied to superhuman models, and reward models diverge from humans through misspecification and misgeneralization.

Source quote, verbatim
Moreover, humans can pursue harmful goals, either innocently or maliciously, and can provide poor feedback when examples are hard to evaluate, especially when applying RLHF to superhuman models. Reward models can differ from humans due to misspecification and misgeneralization
From

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024), Model Overview: Reinforcement Learning Through Human Feedback (RLHF), p. 30
https://ssrn.com/abstract=4855607 · source PDF

Cite as

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

Holds when
Classification

failuresupport: evidencedfailure: Compound RLHF failurefamily: ai-oversight-and-alignment-gapai-and-agents

Verify

The quote above is an exact substring of the source PDF, whose sha256 is eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/acd93c15e8213d7d...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/4855607-018.md | sha256sum