When I tested an older version of Grok this spring, my only comment was blunt: “grok-4.20 was useless.” So when Grok 4.5 was released a month ago and started receiving very positive reviews, I was quite surprised. Today, Grok 4.6 was released and I’ve finally had a chance to test it against HTB challenges, and the team behind it deserves an apology. I’m not sure what the folks at xAI did, but the…
DeepSeek V4 Pro was the best (and most expensive) LLM I tested with Strix a few months ago. After the heartbreaking results of its smaller sibling, DeepSeek V4 Flash 0731, in my HTB-Challenger tests, I was curious to see how the Pro version would handle the new challenges.
I tested the previous version of DeepSeek V4 Flash, now called DeepSeek V4 Flash 0423, with Strix back in April, and I absolutely fell in love with it. It delivered great results at a very low price and became my go-to model for most tasks that didn’t require the capabilities of frontier models. That’s why I was looking forward to testing the latest, improved version on my HTB-Challenger…
Qwen3.8 Max was released recently, but it didn’t create much buzz. I was wondering if it was simply overshadowed by Kimi K3, which was released just a few days earlier, or if its performance wasn’t strong enough to raise any eyebrows.
Kimi K3 landed with a big splash just three weeks ago, and the initial reviewers seemed to agree on one thing: it’s very good, but it’s also quite expensive. So here I am, the late reviewer, with my own numbers. Let’s see if they line up with the prevailing opinion.
What the heck is GPT-5.6 Luna Pro, I hear you asking. That’s a very good question! To answer it, let me quote its description on OpenRouter.ai: “GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.” Okay, so how much better, and how much more expensive, is it compared with GPT-5.6 Luna, I hear you asking…
My first thought was: let’s kick off the HTB-Challenger Benchmark with some state-of-the-art models. Let’s see what these Fables, Opuses, Terras and Sols can do with offensive security challenges. As it turned out, not much - sooner or later, all of them refused to continue with the HTB challenge-solving workflow, with messages like “This content was flagged for possible cybersecurity risk.” So I…
Choosing an LLM model for offensive security work is difficult. Public benchmarks can help compare model capabilities, but they rarely show how much it costs to complete a task from start to finish. The price per million tokens does not tell the whole story: different models may require vastly different numbers of tokens and tool calls before they solve a problem - or fail to solve it. On top of…
It has been a while since my last blog post, but I definitely haven’t stopped playing with AI-assisted security tools. Summer is here, and my schedule is finally a little less crowded, so I’m planning to write a few blog posts about what I’ve learned lately. This post should be short and to the point. After a few months in its v0.x phase, Strix recently reached adulthood with the release of…