# kaal:claim:5245185-028

**Claim.** Anomaly detection and behavioral analysis models trained on agent data may replicate the biases in that data and fail to detect novel deviations absent from the training set.

**Type.** failure  **Support.** evidenced

**Holds when.**

- machine learning monitors trained on historical agent behavior
- novel agent behaviors outside training distribution

**Source quote.**

> rely on machine learning models trained on agent data, which may replicate biases or fail to detect novel deviations not present in training sets.

**From.** Wulf A. Kaal, *How can we Best Monitor AI Agents* (2025), 6.3.1. Circular Dependency: AI Monitoring AI Creates Inherent Bias, page 13

**Cite as.** Wulf A. Kaal, How can we Best Monitor AI Agents (2025). SSRN: https://ssrn.com/abstract=5245185

**Verify.** sha256 of source PDF `4d7adba83ec722480e97bde6528cbe9ce98c709e45cb18794f157a64b8fe7da2` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202025%20-%20How%20can%20we%20Best%20Monitor%20AI%20Agents.pdf

**Failure mode.** training-set-blindness  (family: consensus-and-protocol-attack)

**Topics.** ai-and-agents, education-and-practice

**Keywords.** anomaly-detection, training-data-bias, novel-deviation, machine-learning

**Related claims.**

- specializes: https://wulfkaal.github.io/claims/4855607-003

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
