Understanding GPU Straggler Ranks and NCCL Basics - p1
introduction to DDP ml training and NCCL network stragglers
k10s is two things: kitty: a Daemonset that lives on your Kubernetes cluster that collects node-level GPU + Network telemetry k10s: (kittens) a tui/cli tha...
introduction to DDP ml training and NCCL network stragglers
Here is the raw diagram at the end of the stream: I'll write a follow-up blog-post with a slightly more well-organized notes.
I vibe-coded a GPU-aware Kubernetes TUI for 7 months, archived it, and started over. Here's what AI gets wrong when projects grow complex.
How picking the wrong design atoms cost me a rewrite, and what the right atoms for a GPU-aware Kubernetes TUI actually are.
I'm building k10s because I want to learn how to create value, and k10s is the thing I'm going to learn it on. k10s is a TUI for operating Kubernetes clusters, much like k9s , but simpler and focused on GPU workloads. I'm the first user. I built the first version because I wanted it to exist for myself. That's already a real reason to build something. The bigger question, the one that gets me out…