kaal:claim:4855607-017
Balancing helpfulness against harmlessness is an inherent tension in Safe RLHF rather than a tuning problem that can be resolved once.
Source quote, verbatim
Balancing the dual objectives of helpfulness and harmlessness remains an inherent tension in Safe RLHF.
From
Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024), Model Overview: Reinforcement Learning Through Human Feedback (RLHF), p. 30
https://ssrn.com/abstract=4855607 · source PDF
Cite as
Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607
Holds when
Classification
failuresupport: evidencedfailure: Helpfulness harmlessness tensionfamily: ai-oversight-and-alignment-gapai-and-agents
Related claims
Verify
The quote above is an exact substring of the source PDF, whose sha256 is eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/cde141bf7ae6f5c9...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/4855607-017.md | sha256sum