Prism ML has released impressive ternary and 1-bit variants of Qwen3.6 27B. The 1-bit version is only 3.9 GB, making it dramatically smaller than the original, which is roughly 55 GB.
Bonsai 27B Models (Hugging Face)
These Bonsai 27B models behave quite differently from the original Qwen3.6 model, particularly in terms of accuracy and token efficiency.
Can you get Qwen3.6 27B-level accuracy from a 3.9 GB model?
With reasoning enabled, Bonsai 27B can nearly match the accuracy of Qwen3.6 27B running without reasoning. Moreover, that dramatic reduction in memory comes at a substantial computational cost: Bonsai needs to generate far more tokens to achieve the same result, up to 14 times more on some coding tasks.
When reasoning is enabled for both models, Qwen3.6 27B remains clearly ahead. It uses fewer tokens while achieving significantly higher accuracy. Still, what Prism ML has achieved is remarkable.
In this article, I take a deep dive into Bonsai 27B’s recipe, benchmark performance (my own numbers), and token efficiency. We will examine where the model struggles most, why its raw accuracy is lower, and why the results are nevertheless highly promising.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.