How I Beat NVIDIA NCCL by 2.4x
I built a 2-GPU NVLink AllReduce library that outperforms NVIDIA NCCL by 1.2x-2.4x with 50x+ more stable tail latency.
µs, ns, 80% speed-of-light...
I built a 2-GPU NVLink AllReduce library that outperforms NVIDIA NCCL by 1.2x-2.4x with 50x+ more stable tail latency.
A digestible high-level overview of what happens in The Die - understanding CPUs, GPUs, Von Neumann Architecture, and why GPUs excel at parallel processing.