Qwen 3.8 27B Is Excellent, But Slow

from blog Tao of Mac, | ↗ original
There is certainly quite a bit more to optimise in local inference–Simon got around 72% more throughput from MTP speculative decoding–even if most readily available (and not hugely overpriced) consumer hardware still can’t quite get to the point where memory bandwidth makes dense models usable interactively. I can’t wait for an A3B or adaptive...