RSSAmplifier

Blog

PyTorch

pytorch.orgRSS feed ↗10 posts

Latest posts

FP8 Training on AMD GPUs with TorchTitan and TorchAO: Upstreaming Performance Improvements

At the PyTorch Conference 2025, we demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters using Primus-Turbo, an AMD optimization library for training frameworks such as TorchTitan. We have since upstreamed those AMD optimizations so TorchTitan supports AMD Instinct(™) GPUs directly, with competitive FP8 performance out of the box. All contributions mentioned have been merged into…

Fast, On Device Agentic AI with Muse Glimmer on ExecuTorch

Today, Meta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse Spark for on-device agentic workflows. Alongside, ExecuTorch is adding end-to-end support for running Muse Glimmer on NVIDIA...

PyTorch Conference North America Announces 2026 Keynotes

PyTorch Conference North America will be held in San Jose, California, on October 20–21, 2026. Featured PyTorchCon NA 2026 keynote speakers include: Mark Collier, Executive Director, PyTorch Foundation Mazin Gilbert,...

PyTorch by the Sea: The inaugural Santa Cruz PyTorch Meetup

TL;DR The inaugural Santa Cruz PyTorch Meetup brought together 45 local engineers, students, and leaders for GPU/CUDA talks and lightning presentations on chemistry, plant health, and autonomous driving demonstrating...

FBTriton Infra: Upstream Ingestion, Hierarchical Validation, Ideals vs Realities

TL:DR Learn how Meta’s FBTriton infrastructure powers custom GPU compiler innovations like TLX and autoWS while staying synced with upstream Triton using agentic ingestion and a stratified L1/L2/L3 validation framework....

PyTorch Foundation Flare Pin Community Design Contest

We invite you to design the 2026 PyTorch Foundation flare pin for PyTorch Conference North America. The winning entrant will receive one complimentary ticket to PyTorch Conference North America in...

Helion on TPU: Towards Hardware Heterogeneous Kernel Authoring

TL;DR Helion is PyTorch s high-level DSL for writing performance-portable ML kernels. Partnering with Google, we have built a TPU backend that compiles Helion kernels to Pallas, providing a PyTorch-friendly way...

Driving the Future of Open Source AI: An Update from PyTorch Foundation Projects

TL;DR In April 2025, the PyTorch Foundation evolved into a multi-project Foundation, with the objective to support deeper collaboration across domains and help scale innovation throughout the AI lifecycle. Today,...

PyTorch Conference North America Schedule Is Live

PyTorch Conference North America will bring developers, researchers, and practitioners to San Jose on October 20–21 for sessions spanning training and inference, compiler innovations, responsible AI, applications, and the PyTorch...

Triton Plugin Extensions: Enabling TLX and Custom Compiler Passes Out of the Box

TLDR The PyTorch-Triton 3.7 release introduces the Triton Plugin Extensions system, a framework for dynamically loading custom compiler passes, dialects (including their ops), and DSL extensions into upstream Triton at...