# kaal:claim:5541658-032

**Claim.** Benchmark results for legal LLMs may overstate capability because of data contamination: if a model saw a benchmark's ground truth answers during training, its measured performance reflects memorization rather than genuine generalization.

**Type.** failure  **Support.** evidenced

**Holds when.**

- closed source LLMs evaluated on public benchmarks

**Source quote.**

> If a model has already seen a benchmark's ground-truth answers during training, its performance may reflect memorization rather than genuine generalization, making it hard to assess its ability on truly unseen tasks.

**From.** Wulf A. Kaal, Morgan A. Gray, *The Evolving Role of Artificial Intelligence in Law* (2025), A. Improved Algorithms for Complex Legal Datasets, page 35

**Cite as.** Wulf A. Kaal, Morgan A. Gray, The Evolving Role of Artificial Intelligence in Law (2025). SSRN: https://ssrn.com/abstract=5541658

**Verify.** sha256 of source PDF `e543a2d698fcd522d4d02e034cc9ee1344d0015d2c824b40b9e05ab7c0728c60` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20and%20Gray%20-%202025%20-%20The%20Evolving%20Role%20of%20Artificial%20Intelligence%20in%20Law.pdf

**Failure mode.** benchmark data contamination  (family: measurement-and-metric-failure)

**Topics.** economics

**Keywords.** data-contamination, benchmarks, evaluation-validity, large-language-models

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
