# kaal:claim:4855607-007

**Claim.** The self attention mechanism in transformers scales quadratically with input sequence length, which makes transformers expensive and slow to train and use on long sequences and disqualifies them where real time processing or limited compute is required.

**Type.** failure  **Support.** evidenced

**Holds when.**

- long input sequences
- real time or compute constrained deployments

**Source quote.**

> The self-attention mechanism in Transformers has a quadratic computational complexity with respect to the input sequence length, making them computationally expensive and time-consuming to train and use, especially for long sequences.

**From.** Wulf A. Kaal, *How AI Models are Optimized Through Web3 Governance* (2024), Model Overview: Transformer AI, page 20

**Cite as.** Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

**Verify.** sha256 of source PDF `eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113` at https://raw.githubusercontent.com/wulfkaal/Academic-Papers/main/papers/pdf/Kaal%20-%202024%20-%20How%20AI%20Models%20are%20Optimized%20Through%20Web3%20Governance.pdf

**Failure mode.** Quadratic attention cost  (family: scalability-and-throughput-limit)

**Topics.** ai-and-agents

**Keywords.** transformer-ai, self-attention, quadratic-complexity, compute-constraints, long-sequences

**Canonical form.** This markdown file is the canonical hashed representation of the claim. Its sha256 is the content hash used for attestation.
