# kaal:claim:3409548-009

**Claim.** Repeated use of the same data set by data scientists creates an overfitting risk: the training model overfits the test set, which limits the performance of the applied model on a different dataset.

**Type.** failure  **Support.** argued

**Holds when.**

- data scientists reusing a single data set repetitively

**Source quote.**

> When data scientists use the same data set repetitively a risk exists that the training model will overfit the test set of data, which can limit the performance of the applied model on a different dataset.

**From.** Kaal, *Financial Technology and Hedge Funds* (2019), II.2 Artificial Intelligence and Big Data, page 12

**Cite as.** Kaal, Financial Technology and Hedge Funds (2019). SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3409548

**Verify.** sha256 of source PDF `74227ab2656b06bfe4a29c942fc2ba26df9f476917e0a9f85ec038b8c3402c40` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202019%20-%20Financial%20Technology%20and%20Hedge%20Funds.pdf

**Failure mode.** Overfitting from repeated reuse of the same data set  (family: ai-model-and-training-failure)

**Topics.** ai-and-agents, risk-and-incentives

**Keywords.** overfitting, adaptive-data-analysis, machine-learning, model-risk

**Related claims.**

- restates: https://wulfkaal.github.io/claims/2998033-021

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
