# kaal:claim:4855607-017

**Claim.** Balancing helpfulness against harmlessness is an inherent tension in Safe RLHF rather than a tuning problem that can be resolved once.

**Type.** failure  **Support.** evidenced

**Holds when.**

- Safe RLHF designs that optimize helpfulness and harmlessness jointly

**Source quote.**

> Balancing the dual objectives of helpfulness and harmlessness remains an inherent tension in Safe RLHF.

**From.** Wulf A. Kaal, *How AI Models are Optimized Through Web3 Governance* (2024), Model Overview: Reinforcement Learning Through Human Feedback (RLHF), page 30

**Cite as.** Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

**Verify.** sha256 of source PDF `eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202024%20-%20How%20AI%20Models%20are%20Optimized%20Through%20Web3%20Governance.pdf

**Failure mode.** Helpfulness harmlessness tension  (family: ai-oversight-and-alignment-gap)

**Topics.** ai-and-agents

**Keywords.** safe-rlhf, helpfulness, harmlessness, objective-conflict, alignment

**Related claims.**

- extended_by: https://wulfkaal.github.io/claims/6607458-030

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
