The O(N^2) Complexity Curse of Self-Attention
The standard Transformer architecture has dominated deep learning since 2017. However, its core mechanism—self-attention—suffers from quadratic computational complexity with respect to sequence length, making extreme context windows computationally prohibitive.
Mamba-2 and Structured State Space Duality
Researchers introduced Mamba and selective State Space Models (SSMs), which operate with linear O(N) complexity in time and constant memory footprint during inference. Leading labs are now adopting hybrid models (such as AI21's Jamba), alternating between attention and Mamba layers to achieve the accuracy of transformers at linear processing speeds.