# kaal:claim:2998033-021

**Claim.** Repeated use of the same dataset by data scientists creates an overfitting risk: the training model fits the test set so closely that its performance on a different dataset degrades.

**Type.** failure  **Support.** argued

**Holds when.**

- applies to adaptive data analysis where the same data set is used repetitively

**Source quote.**

> When data scientists use the same data set repetitively a risk exists that the training model will overfit the test set of data which can limit the performance of the applied model on a different dataset.

**From.** Wulf A. Kaal, *Blockchain Innovation for Private Investment Funds* (2017), III.2. Combining AI, Big Data, and Blockchain, page 21

**Cite as.** Wulf A. Kaal, Blockchain Innovation for Private Investment Funds (2017). SSRN: https://ssrn.com/abstract=2998033

**Verify.** sha256 of source PDF `aafb1be3c25cd33da477d759df9ca2f856f0a8fe133d6396da2e75d0af573dbd` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202017%20-%20Blockchain%20Innovation%20for%20Private%20Investment%20Funds.pdf

**Failure mode.** overfitting-in-repeated-dataset-use  (family: ai-model-and-training-failure)

**Topics.** ai-and-agents, risk-and-incentives

**Keywords.** overfitting, adaptive-data-analysis, machine-learning, model-risk

**Related claims.**

- restated_by: https://wulfkaal.github.io/claims/3409548-009

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
