kaal:claim:3409548-009

Repeated use of the same data set by data scientists creates an overfitting risk: the training model overfits the test set, which limits the performance of the applied model on a different dataset.

Source quote, verbatim
When data scientists use the same data set repetitively a risk exists that the training model will overfit the test set of data, which can limit the performance of the applied model on a different dataset.
From

Kaal, Financial Technology and Hedge Funds (2019), II.2 Artificial Intelligence and Big Data, p. 12
https://ssrn.com/abstract=3409548 · source PDF

Cite as

Kaal, Financial Technology and Hedge Funds (2019). SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3409548

Holds when
Classification

failuresupport: arguedfailure: Overfitting from repeated reuse of the same data setfamily: ai-model-and-training-failureai-and-agentsrisk-and-incentives

Related claims
Verify

The quote above is an exact substring of the source PDF, whose sha256 is 74227ab2656b06bfe4a29c942fc2ba26df9f476917e0a9f85ec038b8c3402c40. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/04292da2c0073a86...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/3409548-009.md | sha256sum