kaal:claim:4855607-007

The self attention mechanism in transformers scales quadratically with input sequence length, which makes transformers expensive and slow to train and use on long sequences and disqualifies them where real time processing or limited compute is required.

Source quote, verbatim
The self-attention mechanism in Transformers has a quadratic computational complexity with respect to the input sequence length, making them computationally expensive and time-consuming to train and use, especially for long sequences.
From

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024), Model Overview: Transformer AI, p. 20
https://ssrn.com/abstract=4855607 · source PDF

Cite as

Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

Holds when
Classification

failuresupport: evidencedfailure: Quadratic attention costfamily: scalability-and-throughput-limitai-and-agents

Verify

The quote above is an exact substring of the source PDF, whose sha256 is eb0b3e62374b45a8fa888c6bde9725e606bcb46cf4b5e74a6e851d9f25099113. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/115428c99f1779d4...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/4855607-007.md | sha256sum