Sebastian Raschka, PhD · Dec 3, 2025
From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates
0Sign in to vote or save
This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
Similar to DeepSeek V3, the team released their new flagship model over a major US holiday weekend. Given DeepSeek V3.2's really good performance (on GPT-5...
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.