# kaal:claim:4855607-016

**Claim.** There is a trade off in RLHF between the agent imitating human advice and learning autonomously, and human guidance that is too specific will prevent the agent from discovering novel optimal strategies.

**Type.** failure  **Support.** evidenced

**Holds when.**

- human in the loop RL where the extent of human involvement is a design choice

**Source quote.**

> there is a trade-off between the extent to which the agent should imitate human advice versus learning autonomously. Overspecific human guidance can hinder the agent's ability to discover novel optimal strategies.

**From.** Wulf A. Kaal, *How AI Models are Optimized Through Web3 Governance* (2024), Model Overview: Reinforcement Learning Through Human Feedback (RLHF), page 30

**Cite as.** Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

**Verify.** sha256 of source PDF `eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202024%20-%20How%20AI%20Models%20are%20Optimized%20Through%20Web3%20Governance.pdf

**Failure mode.** Overspecific human guidance  (family: ai-oversight-and-alignment-gap)

**Topics.** ai-and-agents

**Keywords.** rlhf, human-in-the-loop, autonomy-tradeoff, strategy-discovery, guidance-overfitting

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
