Hey there! It’s been a while since our last newsletter and we have a lot of updates. We’ve added AMD training support and completely reimagined Unsloth Studio, introducing a new Hub, expanded customization, more language support, many new features, models, and much more.
We collaborated with AMD to enable you to train and run LLMs locally on AMD GPUs across 500+ models. Works on AMD Radeon, Instinct, Ryzen and data center GPUs across Windows, WSL, and Linux. You can train models 2× faster with 70% less VRAM and no accuracy loss on just 3GB VRAM. We also have optimized ROCm builds for GGUF and Safetensors inference. GitHub repo
We’re introducing a new Model Hub which makes it easier to discover, download and run models, with a trending feed, search, README previews and a resumable download manager. We also rebuilt model context auto-fitting, delivering up to ~3× longer context in several scenarios.
We’ve released new Dynamic NVFP4 quants for NVIDIA Blackwell GPUs, covering Qwen3.6 and Gemma 4. They enable faster and more accurate 4-bit inference. Unsloth can now also create merged FP8 and NVFP4 exports after training, plus imatrix-assisted GGUF quantization.
The full Gemma 4, QAT family now runs and trains in Unsloth, including the new 12B model alongside E2B, E4B, 26B-A4B and 31B. We also uploaded new Dynamic NVFP4 quants. Recently, Google also added better tool calling and MTP for faster and improved local inference.
Unsloth Studio’s appearance is now fully customizable, palettes, custom accent/background colors, fonts and accessibility settings, all synced across devices. We also now support 12 languages including Japanese, Chinese, Portuguese, Hindu, Arabic and others. Local transcription is also coming this week!
We now have much more customization for model loading. Loading now exposes Flash Attention and tensor-parallel options. A new GPU Memory panel lets you choose which GPUs to use, set GPU layers and control MoE expert offload, or leave everything to Unsloth’s automatic fit mode.
DeepSeek-V4-Flash, GLM-5.2, Qwen3.6 - run the new agentic models
Inkling, DiffusionGemma, Qwen-AgentWorld, Ornith, Kimi K2.7 Code and MiniMax-M3 - are all now available and supported in Unsloth.
Hope you have a lovely rest of July! We’ve got even more hardware, Unsloth and model updates on the way. 🦥
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.