Learn CUTLASS the hard way - part 2!
Exploring Hopper Architecture on H100 and try to match PyTorch/cuBLAS performance for GEMMs using CUTLASS
This is the website/technical blog of Kapil Sharma. I work at Meta on the Pytorch team.
Exploring Hopper Architecture on H100 and try to match PyTorch/cuBLAS performance for GEMMs using CUTLASS
Walkthrough of optimization techniques for GEMMs from a naive fp32 kernel to CUTLASS bf16 kernel
Worklog: Performance debugging Triton Kernel
Fused Softmax Triton kernel exploration
RMS Normalization Triton kernel implementation for LLMs