kaal:claim:2998033-021

Repeated use of the same dataset by data scientists creates an overfitting risk: the training model fits the test set so closely that its performance on a different dataset degrades.

Source quote, verbatim
When data scientists use the same data set repetitively a risk exists that the training model will overfit the test set of data which can limit the performance of the applied model on a different dataset.
From

Wulf A. Kaal, Blockchain Innovation for Private Investment Funds (2017), III.2. Combining AI, Big Data, and Blockchain, p. 21
https://ssrn.com/abstract=2998033 · source PDF

Cite as

Wulf A. Kaal, Blockchain Innovation for Private Investment Funds (2017). SSRN: https://ssrn.com/abstract=2998033

Holds when
Classification

failuresupport: arguedfailure: overfitting-in-repeated-dataset-usefamily: ai-model-and-training-failureai-and-agentsrisk-and-incentives

Related claims
Verify

The quote above is an exact substring of the source PDF, whose sha256 is aafb1be3c25cd33da477d759df9ca2f856f0a8fe133d6396da2e75d0af573dbd. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/0954a2c47a31da71...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/2998033-021.md | sha256sum