Tag: Transformer
All the articles with the tag "Transformer".
-
重新理解 MLA:缓存 X 而非 KV
思考:如果不缓存 KV,而是缓存输入 X,可行吗?
-
FlashAttention 详解(V1 & V2)
Fast and Memory-Efficient Exact Attention with IO-Awareness.
-
Roofline 分析:瓶颈的判定与局限
用 Roofline 判断性能瓶颈,并分析它在混部资源预估中为什么会失准。
-
PagedAttention 详解:缓存的存储结构
GPU 时代的系统设计,从 CPU 时代继承了什么。