RSSAmplifier

Blog

Sonny's Blog

Last 10 notes on Sonny's Blog

lubits.chRSS feed ↗10 posts

Latest posts

Flash Attention from Scratch: Appendix B - Block Size Configuration

This appendix dives into how block size configurations affect instruction patterns and performance in Flash Attention.

Flash Attention from Scratch Part 8: Instruction Reduction

Intro In Part 6, we improved our kernel to slightly outperform the reference kernel on the RTX 3090, but found that on the A100, it only reached 80.3% of the reference.

Appendix A - Ampere Microarchitecture

Ampere Microarchitecture This part dives into the SM architecture to reveal how execution units compete for resources.

Flash Attention from Scratch Part 7: A100 Profiling

In Part 6, we exceeded reference performance on the RTX 3090, hitting 101.5% through FP instruction fusion and auto-tuning.

Flash Attention from Scratch Part 6: FP Instruction Fusion and Auto-Tuning

In the previous part, we implemented three major optimizations from the CUTLASS GEMM library: eager block loading, sub-tiling with fragment interleaving, and double buffering.

Flash Attention from Scratch Part 5: Cutlass GEMM Optimizations

Intro In the previous part, we implemented swizzling and achieved a dramatic 2x performance improvement by eliminating bank conflicts.

Flash Attention from Scratch Part 4: Bank Conflicts & Swizzling

Intro In the last part, we used the instructions covered in Part 2 to construct our first kernel and reached nearly half the performance of the official implementation on the RTX 3090.

Flash Attention from Scratch Part 3: Kernel 1

Intro In Part 2, we explored the fundamental CUDA building blocks - tensor core operations (mma) and efficient memory transfers (cp.async & ldmatrix).

Flash Attention from Scratch Part 2: Building Blocks

Intro In this part, we ll explore the CUDA operations that form the foundation of our Flash Attention kernel.

Flash Attention from Scratch Part 1: Intro

Intro In this 10-part series, we re going to implement Flash Attention 2 from scratch on Ampere GPUs.