kaal:claim:4855607-014

Deep reinforcement learning demands large amounts of training data, which suggests its algorithms differ fundamentally from human learning, and learning without supervision becomes particularly hard when rewards are sparse, as they typically are in sequence generation tasks.

Source quote, verbatim
Deep RL methods often demand large amounts of training data, suggesting that the algorithms may differ fundamentally from those underlying human learning. Moreover, learning without supervision is particularly hard when the reward is sparse, which is likely to happen for sequence generation tasks.
From

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024), Model Overview: Reinforcement Learning (RL), p. 26
https://ssrn.com/abstract=4855607 · source PDF

Cite as

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

Holds when
Classification

failuresupport: evidencedfailure: Sparse reward learning failurefamily: ai-model-and-training-failureeconomicsempirical-evidence

Verify

The quote above is an exact substring of the source PDF, whose sha256 is eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/9eaf31289e4a67e6...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/4855607-014.md | sha256sum