64 GGUF quants from 6 uploaders, measured by KL divergence against the BF16 reference using ~250,000 tokens of coding, chat, tool calling, science, non-Latin scripts, and long documents. Full methodology.
unsloth/Qwen3.6-35B-A3B-GGUF (22 quants)
bartowski/Qwen_Qwen3.6-35B-A3B-GGUF (27 quants)
lmstudio-community/Qwen3.6-35B-A3B-GGUF (3 quants)
ggml-org/Qwen3.6-35B-A3B-GGUF (1 quant)
mudler/Qwen3.6-35B-A3B-APEX-GGUF (7 quants)
AesSedai/Qwen3.6-35B-A3B-GGUF (4 quants)
The best quant at each size. If it’s not in this table, a smaller file with lower KL exists.
Qwen 3.6 35B A3B quantizes significantly better than its predecessor Qwen 3.5 35B A3B. Q8_0 has KL 0.069, compared to 0.121 for Qwen 3.5 35B A3B.
UD-Q8_K_XL (38.4 GB, KL 0.073) is not on the frontier. ggml-org’s Q8_0 (36.9 GB, KL 0.069) is smaller and has lower KL.
unsloth dominates: 14 of 26 frontier positions. AesSedai earns 4 by filling size gaps between unsloth’s UD quants.
lmstudio-community and mudler (APEX) never appear on the frontier.
Tool calling is the worst category at Q8_0 (KL 0.177), worse than long documents (0.121). This is unusual compared to other Qwen models.
Even at UD-IQ1_M (10.0 GB), the model retains 86.9% top-1.
Tool calling dominates the quality loss at Q8_0 (KL 0.177), followed by long documents (0.121). Coding and science stay at or below 0.010. This pattern is different from Qwen 3.5 models where long documents were always the worst category.
Inference: TextGen + patched llama.cpp (logprob extraction from prompt)
Reference: BF16 GGUF by unsloth
Dataset: ~250,000 tokens across 6 categories (coding, general chat, tool calling, science, non-Latin scripts, long documents)
Input format: full OpenAI-compatible messages rendered through the model’s Jinja2 chat template
Metric: KL divergence, computed token-by-token between reference and quantized top-40 log-probability distributions
Please do not share the plots and tables in this post publicly. Instead, share the post URL. These measurements are expensive to produce and more subscribers means more models benchmarked.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.