Tag: Transformer
All the articles with the tag "Transformer".
-
重新理解 MLA:缓存 X 而非 KV
思考:如果不缓存 KV,而是缓存输入 X,可行吗?
-
FlashAttention 详解(V1 & V2)
Fast and Memory-Efficient Exact Attention with IO-Awareness.
-
为什么KV缓存没有Q
从感性认知和数学公式两个角度,解释为什么自回归推理只缓存 KV 而不缓存 Q,以及 GQA 背后的设计哲学。