kaal:claim:4855607-007
The self attention mechanism in transformers scales quadratically with input sequence length, which makes transformers expensive and slow to train and use on long sequences and disqualifies them where real time processing or limited compute is required.
Source quote, verbatim
The self-attention mechanism in Transformers has a quadratic computational complexity with respect to the input sequence length, making them computationally expensive and time-consuming to train and use, especially for long sequences.
From
Cite as
Holds when
Classification
failuresupport: evidencedfailure: Quadratic attention costfamily: scalability-and-throughput-limitai-and-agents
Verify