# Data exhaustion

`kaal:entity:data-exhaustion`

**Status.** derived

This node is assembled mechanically from the 4 claims that carry the concept tag `data-exhaustion`. It is a roster of what the corpus says under this term. It is **not** an adjudicated definition: no single statement here has been ruled canonical, and no first-appearance call has been made. Read the claims and judge for yourself.

## Every claim under this term

4 claims across 1 works, 2025 to 2025.

**2025**

- [5095633-001](https://wulfkaal.github.io/claims/5095633-001) [predictive/evidenced] -- The accessible reserves of publicly available human-created text usable for training large language models could be exhausted by 2028 at current usage trajectories.
  > the accessible reserves of publicly available human-created text could be exhausted by 2028, given current usage trajectories.
  Wulf A. Kaal, Artificial Intelligence The Final Frontier (2025). SSRN: https://ssrn.com/abstract=5095633
- [5095633-002](https://wulfkaal.github.io/claims/5095633-002) [mechanism/argued] -- Data exhaustion is caused primarily by the exponential growth in the size of datasets needed to build increasingly sophisticated AI models, not by any sudden loss of existing text.
  > The phenomenon—often referred to as "data exhaustion"—arises largely from the exponential increase in the size of datasets required to develop increasingly sophisticated AI models.
  Wulf A. Kaal, Artificial Intelligence The Final Frontier (2025). SSRN: https://ssrn.com/abstract=5095633
- [5095633-003](https://wulfkaal.github.io/claims/5095633-003) [empirical/evidenced] -- The total effective stock of human-generated text is estimated at roughly 300 trillion tokens, with a plausible range from 100 trillion to 1 quadrillion tokens.
  > Researchers estimate that the total effective stock of human-generated text may currently stand at approximately 300 trillion tokens, with the range varying from 100 trillion to 1 quadrillion tokens.
  Wulf A. Kaal, Artificial Intelligence The Final Frontier (2025). SSRN: https://ssrn.com/abstract=5095633
- [5095633-004](https://wulfkaal.github.io/claims/5095633-004) [mechanism/argued] *(failure mode)* -- The apparent abundance of internet text overstates the usable supply, because much of it fails quality thresholds for model training due to redundancy, noise, or irrelevance.
  > while the internet contains a vast corpus of textual material, not all content meets quality thresholds suitable for model training, given issues such as redundancy, noise, or irrelevance.
  Wulf A. Kaal, Artificial Intelligence The Final Frontier (2025). SSRN: https://ssrn.com/abstract=5095633

## Verify

Every claim above resolves to a record carrying a verbatim source quote, the sha256 of the source PDF, and a preformatted citation. Nothing here asks to be taken on trust.

    curl -s https://wulfkaal.github.io/entities/data-exhaustion.md | sha256sum

**Canonical form.** This markdown file is the canonical hashed representation of this entity node. Its sha256 is the content hash.
