RSS Amplifier

adlrocha · May 24, 2026

72 tok/s on my Strix Halo. Qwen + MTP enabled is AMAZING. I'll test its perf...

0
Sign in to vote or save

This site does not allow itself to be embedded. You can still read it on the original site — the toolbar below keeps your place in the directory.

72 tok/s on my Strix Halo. Qwen + MTP enabled is AMAZING. I'll test its performance in coding and other tasks and report back, but is the closest that I've got to a usable LLM model. Next step: Gemma4 with a drafter model. If you have no idea what I am talking about give these two a read: * https://adlrocha.substack.com/p/adlrocha-towards-local-plug-and-play *…

72 tok/s on my Strix Halo. Qwen + MTP enabled is AMAZING. I'll test its performance in coding and other tasks and report back, but is the closest that I've got to a usable LLM model. Next step: Gemma4 with a drafter model. If you have no idea what I am talking about give these two a read: * https://adlrocha.substack.com/p/adlrocha-towards-local-plug-and-play * https://adlrocha.substack.com/p/adlrocha-in-a-quest-to-becoming-ai

Read on x.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.