# Evaluation validity

`kaal:entity:evaluation-validity`

**Status.** derived

This node is assembled mechanically from the 2 claims that carry the concept tag `evaluation-validity`. It is a roster of what the corpus says under this term. It is **not** an adjudicated definition: no single statement here has been ruled canonical, and no first-appearance call has been made. Read the claims and judge for yourself.

## Every claim under this term

2 claims across 1 works, 2025 to 2025.

**2025**

- [5541658-008](https://wulfkaal.github.io/claims/5541658-008) [failure/argued] *(failure mode)* -- Prediction on a test set of existing judgments is not the same task as predicting outcomes for a party mid-litigation, because the precise formulation of facts used by such models emerges only once the judgment has been issued.
  > for instance, a lawyer advising a client on the probable outcome of a court hearing—does not have access to the precise formulation of facts presented in a judgment, as this formulation emerges only once the judgment has been issued.
  Wulf A. Kaal, Morgan A. Gray, The Evolving Role of Artificial Intelligence in Law (2025). SSRN: https://ssrn.com/abstract=5541658
- [5541658-032](https://wulfkaal.github.io/claims/5541658-032) [failure/evidenced] *(failure mode)* -- Benchmark results for legal LLMs may overstate capability because of data contamination: if a model saw a benchmark's ground truth answers during training, its measured performance reflects memorization rather than genuine generalization.
  > If a model has already seen a benchmark's ground-truth answers during training, its performance may reflect memorization rather than genuine generalization, making it hard to assess its ability on truly unseen tasks.
  Wulf A. Kaal, Morgan A. Gray, The Evolving Role of Artificial Intelligence in Law (2025). SSRN: https://ssrn.com/abstract=5541658

## Verify

Every claim above resolves to a record carrying a verbatim source quote, the sha256 of the source PDF, and a preformatted citation. Nothing here asks to be taken on trust.

    curl -s https://wulfkaal.github.io/entities/evaluation-validity.md | sha256sum

**Canonical form.** This markdown file is the canonical hashed representation of this entity node. Its sha256 is the content hash.
