RSSAmplifier

Home on Blake Hamm · Feb 17, 2026

Self-hosted AI Model Upgrade

0
Sign in to vote or save

Blake Hamm · Blake Hamm

In this blog, I test some new-to-me AI models that fit on my homelab . I record the latency and provide a ‘Vibe Score’ to see if they replace existing models for my use cases.

TL;DR

Homelab Review

First, let’s dig into how I run a homelab that supports AI workloads; you can find more details here . Essentially, I have two AMD AI Max+ 395 (strix halo) and one AMD AI R9700.

The AI Max+ provides 128GB of unified memory (allocatable as VRAM) and I can run two decent sized models (~30b-120b models). I have an open ticket to enable llama.cpp rpc which will support larger models, but it’s still a WIP… The R9700 is another solid system with 32GB of VRAM and I use it for smaller, embedding models. So, as of now, I have capacity to run three models at a time.

The models I test are dictated by my hardware. I’ve made an effort to avoid Nvidia because I believe in the underdog (and suffering apparently). Given these constraints, what am I actually trying to accomplish with local AI?

Dual Framework Mainboards (AMD AI Max+ 395 - Strix Halo) in a 2u server rack case. 256GB of unified memory and a rats nest of cables; LFG!

Dual Framework Mainboards (AMD AI Max+ 395 - Strix Halo) in a 2u server rack case. 256GB of unified memory and a rats nest of cables; LFG!

Use Cases

In my day-to-day as an AI/ML Engineer, I use the latest models from OpenAI, Anthropic, AWS and Google to build agentic applications. These proprietary models are state of the art (SOTA), no cap. They can service almost any use case, given the right prompt and context.

In contrast, even with decent hardware and incredible research published, the use cases for running local models have their limitations… At the moment, I have tested local models in Open WebUI , Roo Code and OpenCode . I’ve had success with Open WebUI and would recommend it for local models. Roo Code and OpenCode are trickier and depending on the codebase and task, local models will struggle. When context exceeds 20k, these models run at a snail’s pace on my AMD hardware. I believe this can be solved with better context management and prompting, but I didn’t make much progress with my first JavaScript project and VSCode plugin - camber

I used to dismiss “this is the worst AI will ever be” as marketing fluff for AGI. But with local models, that cliché has never felt more accurate.

Why the upgrade?

So why keep testing new models when the current ones have all these limitations? Because open source research is moving fast. Right before Chinese New Year, labs like DeepSeek , MiniMax , Z.AI and Moonshot AI dropped some serious heat.

I stay up to date on these developments through communities like LocalLLama and bycloud on Discord and Reddit. While I’m no AI researcher, as an AI/ML engineer and self-hoster, I need to understand model architectures at a high level - specifically their impact on quality, latency, and context length. This matters both for my homelab and production AI applications.

Given these limitations and my daily dependence on cloud APIs, you might wonder why I bother with self-hosted AI at all.

Ultimate goal

The goal is simple: find an open source model I can run locally and actually use with Roo Code and OpenCode. I started using Open Router for my personal projects and found that Kimi K2.5 fits that bill; I would consider this the best open source model for my use case and almost on par with SOTA closed-source models (while being a fraction of the cost). Unfortunately, I can’t fit this model on 128GB of VRAM…

In addition to replacing paid models, I want to see the benefits of the latest model architecture research in latency. Through self-hosting AI models, I learned how MoE architectures, attention mechanisms and quantization techniques impact performance. GPT-OSS 120b drove this home; its speed and capability made these architectural trade-offs tangible. Testing new models helps me understand how different LLM families and architectures affect resource usage and latency.

In production AI applications, closed-source models become deprecated and you are required to upgrade. One of the great things with local AI is you don’t have to worry about models being removed. Regardless, it’s best practice to have a framework to validate AI quality and latency in case you find a faster, cheaper model with better outcomes.

I’m hopeful that this new round of open source models will be more intelligent and faster; only one way to find out… Before testing the new batch, let’s review what I’m currently running.

Baseline Vibe

Unfortunately, I don’t have a formal eval process comparing models so the review is more of a ‘vibe’. Also, I have monitoring and tracing configured on my AI gateway ( kube-ai-stack ), so I will be collecting data for a post-deployment evaluation.

First off, let’s review the fully arbitrary ‘Vibe Score’ which is my personal feeling towards the model. Basically, I will send the same prompt in Roo Code and OpenWebUI and record the latency along with a 1-5 Vibe Score on quality. I’ll base it off of how well it solves the problem, any failures it might encounter and overall how I like the response. To put it simply, a 5 means the model has either incredibly high quality results OR it runs quick and has sufficient results, but may need some direction and hand holding.

Also, I collected a one-time ’latency’ metric. This is not a scientific average or anything of that nature. It is simply a record of the latency for the first response in my test query in Roo Code. This includes the coldstart time when scaling from zero which is highly correlated with model size. Here is the prompt provided in Roo Code for my homelab project :

docker/ceph-cleanup/README.md:1-3

Ceph cleanup

This python script will cleanup ceph orphaned data in the kubernetes pool. AKA: it deletes data that is not in the prod kubernetes cluster as pv or volume snapshots

Write unit tests for this project. Use pytest.

I chose this prompt because I found LLMs do a great job writing unit tests. This ‘project’ is one simple python script and I intentionally left it vague. Also, this file is part of a monorepo and because of how Roo sends prompts, the context starts out at 10k+, stressing the models a bit.

So, let’s dive right into my review of the current models I have available:

ModelFamilyPrimary use caseFeaturesNotesLatencyVibe Score
Qwen/Qwen3-Embedding-8B-GGUF:F16QwenEmbeddings (Roo Code and LiteLLM cache with Qdrant )Embedding Only, TextAt one point I tried a few embedding models like Nomic Embed Code and Mistral Instruct Embed , but I decided Qwen embed can do it allNA5
ggml-org/Qwen3-32B-GGUF:F16QwenRarely used in practice; in theory, a good base model for fine tuning which was my original purposeDense, Base ModelBeing a dense model, this was a little too slow on my hardware4m 29s3
unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF:Q6_K_XLQwenQuick chat queries; likely a good model for custom agent, but didn’t testMoE, Hybrid Attention, InstructRecently replaced qwen3-vl-30b-a3b ; there wasn’t a noticeable difference, but it appeared higher on benchmarks38.3s5
unsloth/Qwen3-Next-80B-A3B-Thinking-GGUF:Q6_K_XLQwenMore complex chat and search queriesMoE, Hybrid Attention, ReasoningSimilar to above, it replaced the 30b thinking version; however, I noticed it thinks a bit too excessively2m 46s3
ggml-org/Qwen3-Coder-30B-A3B-Instruct-Q8_0-GGUFQwenGreat with Roo Code for coding-specific tasks, less reliable for tool calls/agentsMoE, Coding, InstructWorked better with simpler tools/prompts in OpenCode, but struggled with the complexity of Roo Code49.8s4
unsloth/Seed-OSS-36B-Instruct-GGUF:Q8_K_XLByteDanceBest for its size with Roo Code; good balance of speed and qualityDense, Reasoning, InstructVery impressed with this model; one of my top picks; recent versions of llama.cpp have degraded performance - I am still unsure which container to use2m 51s5
gghfez/gpt-oss-120b-Derestricted.MXFP4_MOE-ggufGPTExcellent for chat and search toolsMoE, Reasoning, Instruct, DerestrictedThis is the jailbroke version of gpt oss; TBH I didn’t notice much of a difference1m5
unsloth/gemma-3-27b-it-GGUF:BF16GoogleI’ve tested it with images, seemed okay; I really don’t have a good use caseVision, Reasoning, Instruct, MultimodalThis is a bit faster than the latency suggests when using for multimodal use cases; the model struggles in agentic coding6m 42s3
bartowski/cerebras_GLM-4.5-Air-REAP-82B-A12B-GGUF:Q4_K_LGLMExcellent for agentic tasks, coding and more difficult problems; top tier with Roo CodeReasoning, MoE, InstructCan be a bit slow with significant context3m 7s5
unsloth/GLM-4.5-Air-GGUF:Q4_K_XLGLMVery similar to above model, but slightly slowerReasoning, MoE, InstructPretty slow3m 44s4
ggml-org/Llama-4-Scout-17B-16E-Instruct-GGUF:Q4_K_MLlamaIt works with images as well; don’t have much of a use caseVision, MoE, Instruct, MultimodalAt one point in time, I had quite a few llama 3 models; after adding qwen and glm models, the llama 3 models showed their age with poor latency and quality2m 5s2

Looking back at this, there are only a few five star models for my use cases. These include Qwen embed, Qwen instruct, Seed OSS, GPT-OSS and GLM-Air-REAP; I could probably cover all my use cases with these models. Many of the other models are complementary and serve different purposes. Gemma and Llama are multimodal so they stand out slightly in that sense. Qwen 32b is a dense, base model which would be ideal for fine-tuning, but I haven’t gotten around to that.

That’s the joy of self hosting! No worries if I have model parameters sitting around on my hard drive. Better to have them accessible and in my possession than not at all! With that baseline established, here’s what I’m testing next and the new vibe check.

New model Vibe

ModelFamilyUse CaseFeaturesLatencyVibe Score
unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:BF16MistralAgentic coding (Roo Code)Vision, Reasoning, Code1m 27s3
unsloth/Devstral-2-123B-Instruct-2512-GGUF:Q5_K_XLMistralAgentic coding (heavily quantized)Vision, Reasoning, Code10m 15s2
unsloth/GLM-4.7-Flash-REAP-23B-A3B-GGUF:BF16GLMGeneral purpose, quick queriesMoE, Reasoning, Lightweight40s5
unsloth/GLM-4.7-Flash-GGUF:BF16GLMGeneral purposeMoE, Fast Inference1m 41s4
unsloth/GLM-4.7-REAP-218B-A32B-GGUF:Q2_K_XLGLMComplex reasoning tasksMoE, Reasoning, Large Scale3m 58s2
bartowski/stepfun-ai_Step-3.5-Flash-GGUF:Q3_K_XLStepfunLong context tasksMoE, Reasoning, Long Context2m 47s2
unsloth/Qwen3-Coder-Next-GGUF:Q6_K_XLQwenCoding tasks, tool useMoE, Coding, Tool Use1m 31s5
unsloth/Nemotron-3-Nano-30B-A3B-GGUF:BF16NVIDIAFunction calling, general purposeMoE, Efficient, Function Calling1m 22s4
unsloth/Kimi-Dev-72B-GGUF:Q8_K_XLMoonshotCoding, long contextCoding, Long Context11m 55s2
bartowski/moonshotai_Kimi-Linear-48B-A3B-Instruct-GGUF:Q8_0MoonshotGeneral purpose, daily driverMoE, Linear Attention, Efficient1m 25s5
unsloth/MiniMax-M2.5-GGUF:Q2_K_XLMiniMaxLong context, research tasksMoE, Long Context3m 24s3

A few models emerged as standouts. Kimi Linear proved to be a fantastic all-purpose model; it was fast, capable, and consistent across different tasks. Qwen Coder Next was incredible which is echoed by communities online; its coding capabilities and speed are exceptional and it has immediately become my go-to for AI-assisted development. GLM 4.7 Flash REAP also impressed with its speed and solid general performance, vibing similar to GPT OSS 120b, but faster. Nemotron also caught my attention with its impressive speed; I’m curious to see how it performs in more targeted agentic workflows. On the other hand, MiniMax and GLM 4.7 REAP were high quality all-rounders but their latency made them tough for realtime workflows, but I could imagine them excelling in batch processes. Unfortunately, Devstral has some known issues with its chat template and proved unusable.

Conclusion

This round of vibing transformed my model lineup, retiring my previous go-tos: Qwen Coder Instruct, Seed OSS and GLM Air in favor of these newer drops. Kimi Linear has become my daily driver for general tasks and Qwen Coder Next is a game-changer for coding. It’s wild how good these open source models are getting; they’re approaching the quality of the big proprietary ones.

The biggest surprise? Heavy quantization actually works. I used to avoid Q2_K and Q3_K quants, thinking they’d be garbage, but MiniMax and GLM 4.7 REAP at Q2_K_XL proved me wrong. They’re very high quality, but noticeably slow. I can imagine leveraging them for background, research-focused tasks which is something I have planned .

Remember that goal of finding a model that actually works locally? This round delivered. Kimi Linear and Qwen Coder Next proved the dream is getting closer. But what’s exciting isn’t just my homelab win - it’s what this means for everyone.

The pace of open source AI is insane right now. You have small research teams with limited hardware dropping models that compete with (and sometimes beat) what billion-dollar companies are putting out. Meanwhile the industry is frothing at the mouth about AGI and valuations, but the real story is that a handful of dedicated researchers are just… giving away powerful AI for free!

While Dario keeps promising SWEs will be obsolete in 6 months and my dad reads the WSJ convinced there’s no bubble, these labs are actually democratizing access. Self-hosting used to be about privacy or avoiding vendor lock-in, but now? It’s becoming a real alternative to feeding the cloud monopoly. My dream is that soon, anyone can afford a little box under their desk running models that rival the APIs. No subscriptions, no rate limits, no data harvesting. Just you and your AI.

What’s next

Regardless of my philosophy on local AI, my process for evaluating models is a mess. I’m manually downloading models, testing one or two quants if I’m feeling ambitious and patient enough, juggling different llama.cpp containers to compare Vulkan vs ROCm backends, checking Arize Phoenix for latency and giving an arbitrary ‘vibe check’. It’s tedious and gets in the way of actually using these models.

I’d love to test all the different versions of kimi linear , but it would take me another 3 days and would still be just a vibe (not the good kind).

I need to automate this properly. I want to systematically test different quantization levels, container images, and llama.cpp CLI args to figure out the sweet spot of quality vs speed for each model. Sort of like what I do in my day-to-day for production AI applications. Real benchmarks, not just vibes. Checkout my next project beyond-vibes here to follow along.

Read the original on site.bhamm-lab.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.