Flash Attention from Scratch: Appendix B - Block Size Configuration
This appendix dives into how block size configurations affect instruction patterns and performance in Flash Attention.
Last 10 notes on Sonny's Blog
This appendix dives into how block size configurations affect instruction patterns and performance in Flash Attention.
Intro In Part 6, we improved our kernel to slightly outperform the reference kernel on the RTX 3090, but found that on the A100, it only reached 80.3% of the reference.
Ampere Microarchitecture This part dives into the SM architecture to reveal how execution units compete for resources.
In Part 6, we exceeded reference performance on the RTX 3090, hitting 101.5% through FP instruction fusion and auto-tuning.
In the previous part, we implemented three major optimizations from the CUTLASS GEMM library: eager block loading, sub-tiling with fragment interleaving, and double buffering.
Intro In the previous part, we implemented swizzling and achieved a dramatic 2x performance improvement by eliminating bank conflicts.
Intro In the last part, we used the instructions covered in Part 2 to construct our first kernel and reached nearly half the performance of the official implementation on the RTX 3090.
Intro In Part 2, we explored the fundamental CUDA building blocks - tensor core operations (mma) and efficient memory transfers (cp.async & ldmatrix).
Intro In this part, we ll explore the CUDA operations that form the foundation of our Flash Attention kernel.
Intro In this 10-part series, we re going to implement Flash Attention 2 from scratch on Ampere GPUs.