Today for AI
Explore indexBack to Terms

Multi-head Latent Attention (MLA)

An attention mechanism that compresses KV cache via low-rank projection to reduce memory usage and computation.

No public content is connected to this entity yet.