entity · derived
Data exhaustion
Derived node: assembled mechanically from the claims carrying data-exhaustion. A roster, not an adjudicated definition.
Every claim under this term
- 5095633-001 : The accessible reserves of publicly available human-created text usable for training large language models could be exhausted by 2028 at current usage trajectories.
- 5095633-002 : Data exhaustion is caused primarily by the exponential growth in the size of datasets needed to build increasingly sophisticated AI models, not by any sudden loss of existing text.
- 5095633-003 : The total effective stock of human-generated text is estimated at roughly 300 trillion tokens, with a plausible range from 100 trillion to 1 quadrillion tokens.
- 5095633-004 : The apparent abundance of internet text overstates the usable supply, because much of it fails quality thresholds for model training due to redundancy, noise, or irrelevance.