kaal:claim:4855607-016

There is a trade off in RLHF between the agent imitating human advice and learning autonomously, and human guidance that is too specific will prevent the agent from discovering novel optimal strategies.

Source quote, verbatim
there is a trade-off between the extent to which the agent should imitate human advice versus learning autonomously. Overspecific human guidance can hinder the agent's ability to discover novel optimal strategies.
From

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024), Model Overview: Reinforcement Learning Through Human Feedback (RLHF), p. 30
https://ssrn.com/abstract=4855607 · source PDF

Cite as

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

Holds when
Classification

failuresupport: evidencedfailure: Overspecific human guidancefamily: ai-oversight-and-alignment-gapai-and-agents

Verify

The quote above is an exact substring of the source PDF, whose sha256 is eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/3701bee8d234b34c...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/4855607-016.md | sha256sum