# Overfitting

`kaal:entity:overfitting`

**Status.** derived

This node is assembled mechanically from the 4 claims that carry the concept tag `overfitting`. It is a roster of what the corpus says under this term. It is **not** an adjudicated definition: no single statement here has been ruled canonical, and no first-appearance call has been made. Read the claims and judge for yourself.

## Every claim under this term

4 claims across 3 works, 2017 to 2024.

**2017**

- [2998033-021](https://wulfkaal.github.io/claims/2998033-021) [failure/argued] *(failure mode)* -- Repeated use of the same dataset by data scientists creates an overfitting risk: the training model fits the test set so closely that its performance on a different dataset degrades.
  > When data scientists use the same data set repetitively a risk exists that the training model will overfit the test set of data which can limit the performance of the applied model on a different dataset.
  Wulf A. Kaal, Blockchain Innovation for Private Investment Funds (2017). SSRN: https://ssrn.com/abstract=2998033

**2019**

- [3409548-009](https://wulfkaal.github.io/claims/3409548-009) [failure/argued] *(failure mode)* -- Repeated use of the same data set by data scientists creates an overfitting risk: the training model overfits the test set, which limits the performance of the applied model on a different dataset.
  > When data scientists use the same data set repetitively a risk exists that the training model will overfit the test set of data, which can limit the performance of the applied model on a different dataset.
  Kaal, Financial Technology and Hedge Funds (2019). SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3409548
- [3409548-010](https://wulfkaal.github.io/claims/3409548-010) [design/argued] -- Requiring data scientists to stake a cryptocurrency on their own predictions is a workable remedy for overfitting, because the stake expresses confidence in live performance and lets the fund select the optimal model.
  > The staking process, in turn, enables Numerai to choose the optimal model and in the process improve the performance of its hedge fund.
  Kaal, Financial Technology and Hedge Funds (2019). SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3409548

**2024**

- [4755632-003](https://wulfkaal.github.io/claims/4755632-003) [failure/argued] *(failure mode)* -- The move by AI developers toward smaller training datasets raises the risk of overfitting, especially with complex models, which forces LLM developers to rely on regularization to counteract overfitting of the model to the training data.
  > However, with small datasets in LLMs, the risk of overfitting also rises, especially with complex models. Therefore, LLM developers have to turn to regularization in an effort to address overfitting of the model with the training data.
  Wulf A. Kaal, AI Learning - Decentralized Governance to Optimize Human Output Datasets for AI Learning (2024). SSRN: https://ssrn.com/abstract=4755632

## Verify

Every claim above resolves to a record carrying a verbatim source quote, the sha256 of the source PDF, and a preformatted citation. Nothing here asks to be taken on trust.

    curl -s https://wulfkaal.github.io/entities/overfitting.md | sha256sum

**Canonical form.** This markdown file is the canonical hashed representation of this entity node. Its sha256 is the content hash.
