adlrocha · May 24, 2026
72 tok/s on my Strix Halo. Qwen + MTP enabled is AMAZING. I'll test its perf...
0Sign in to vote or save
This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.
72 tok/s on my Strix Halo. Qwen + MTP enabled is AMAZING. I'll test its performance in coding and other tasks and report back, but is the closest that I've got to a usable LLM model. Next step: Gemma4 with a drafter model. If you have no idea what I am talking about give these two a read: * https://adlrocha.substack.com/p/adlrocha-towards-local-plug-and-play *…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.