RSS Amplifier

localbench · Apr 20, 2026

Qwen 3.6 35B A3B GGUF Quality Benchmark: unsloth, bartowski, lmstudio-community, ggml-org, mudler, AesSedai compared

0
Sign in to vote or save

oobabooga · localbench

64 GGUF quants from 6 uploaders, measured by KL divergence against the BF16 reference using ~250,000 tokens of coding, chat, tool calling, science, non-Latin scripts, and long documents. Full methodology.

The best quant at each size. If it’s not in this table, a smaller file with lower KL exists.

  • Qwen 3.6 35B A3B quantizes significantly better than its predecessor Qwen 3.5 35B A3B. Q8_0 has KL 0.069, compared to 0.121 for Qwen 3.5 35B A3B.

  • UD-Q8_K_XL (38.4 GB, KL 0.073) is not on the frontier. ggml-org’s Q8_0 (36.9 GB, KL 0.069) is smaller and has lower KL.

  • unsloth dominates: 14 of 26 frontier positions. AesSedai earns 4 by filling size gaps between unsloth’s UD quants.

  • lmstudio-community and mudler (APEX) never appear on the frontier.

  • Tool calling is the worst category at Q8_0 (KL 0.177), worse than long documents (0.121). This is unusual compared to other Qwen models.

  • Even at UD-IQ1_M (10.0 GB), the model retains 86.9% top-1.

Tool calling dominates the quality loss at Q8_0 (KL 0.177), followed by long documents (0.121). Coding and science stay at or below 0.010. This pattern is different from Qwen 3.5 models where long documents were always the worst category.

  • Inference: TextGen + patched llama.cpp (logprob extraction from prompt)

  • Reference: BF16 GGUF by unsloth

  • Dataset: ~250,000 tokens across 6 categories (coding, general chat, tool calling, science, non-Latin scripts, long documents)

  • Input format: full OpenAI-compatible messages rendered through the model’s Jinja2 chat template

  • Metric: KL divergence, computed token-by-token between reference and quantized top-40 log-probability distributions

Full methodology

Please do not share the plots and tables in this post publicly. Instead, share the post URL. These measurements are expensive to produce and more subscribers means more models benchmarked.

Read the original on localbench.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.