Where classical QE meets agentic intelligence. Real implementation stories, experiments, and frameworks from the trenches of Agentic Quality Engineering.
Two weeks of hardening a platform at Ruv's pace, six podcast episodes, a first-ever Vienna meetup, four releases about honesty, a website audit that turned on its own author, and the fortnight a musician left the orchestra without the music stopping. For over a year the argument has been: own the harness, and the models can come and go without taking your capability with them. This fortnight it…
Two weeks of building quality nets around a platform racing toward release, a trip to Munich, a benchmark where the frontier model never won a task, and the moment my own AI reviewers ruled against me. QE-Court, the newest skill in the Agentic QE Fleet, prosecutes every change with independent AI reviewers from different vendors (GPT via the Codex provider, Cognitum, and Claude), each with their…
Three weeks of harness and metaharness work, benchmarks that were allowed to say no, two meetups in two different formats, a training plan for the Foundation, and the week the community showed up at my door. Beta-tested Reuven Cohen's MetaHarness (nine OIA layers, learned router, MCP defaults to deny). The darwin-qe local-model benchmarks (v3.10.9-v3.11.4) answered whether cheaper models can do QE…
Three weeks, two countries, a dozen rooms — and one observation that showed up in every single one. ExpoQA in Madrid (first time on stage: Bridging Classical and Agentic Quality Engineering) and the first in-person gathering of the Agentics Foundation in Budapest both surfaced the same line: the vendors have absorbed ~96% of what were once our harness additions, so the only durable move is to own…
One week, six releases, a guest from London, two meetups, three developer conversations that changed how I think about what this work is becoming, and multiple preparations for conference talks. v3.9.32 fixed four more stacked bugs hidden under v3.9.31 (issue #491) and added a daemon-runtime seam test suite as a pre-publish release gate. v3.9.34 stopped append-only vector files from growing…
Two weeks, thirteen releases, one contributor who filed better bugs than most teams write tests, a model that deleted my Docker containers, and the walk that made the rest of it possible. Nagual crossed 300 patterns; Xu et al.'s generalization gap theorem reframes retrieval-based memory; OWASP published Top 10 for Agentic Applications; evaluation awareness qualifies behavioral testing. Thirteen…
Three weeks, two new projects, one public launch, and the week I learned what bandwidth actually costs. Started a new collaboration with Reuven on Cognitum (Raspberry Pi and ESP32S edge agents) and learned that USB security flags need context. Seven fleet releases (v3.9.12 init fix, v3.9.13 Opus 4.7 migration, v3.9.14 supply-chain hardening, v3.9.15 ARM64 browser, v3.9.16 brain CLI, v3.9.17 the…
The week the community started using my words, and the weight that came with them. When a phrase you wrote turns up in somebody else's newsletter, the claim is no longer yours to defend alone. Four releases (v3.9.8 process insurance, v3.9.9 qe-browser primitive, v3.9.10 multi-provider advisor routing, v3.9.11 upgrade-path fix), three thinking threads on confidence, flow, and identity, and the…
The week I discovered the foundation under my fleet was lying. A vector-search library returned wrong neighbors with the right shape and the right latency. Five hotfixes chased the symptom; the disease was sitting one layer down, calmly returning wrong answers. Self-query as the simplest possible oracle, quietly succeeding with incorrect values as the dominant agentic failure mode, and the…
Agents generate impressive reports. Classical testing taught me to cross-examine every one of them. When a coverage pipeline fabricated 95% on a file with zero tests, the oracle problem became personal. Consistency oracles, SHA-256 witness chains, deterministic YAML pipelines, CUSUM drift detection, and the classical testing infrastructure that agent trust actually needs.
I was reading a twenty-year-old testing framework while my agents shipped six releases. The framework had more to say about what went wrong than the agents did. Bach and Bolton's RST framework, the HTSM, exploratory polarities, and the trust migration from TDD through BDD and EDD to ODD — classical testing wisdom translated into agentic enforcement architecture.
The orchestra has a score. It's detailed. It's been rehearsed. And nobody's reading it. When 80+ skills exist but agents skip verification steps, the problem isn't coverage — it's compliance. Featuring the Surrogation Trinity, harness engineering, back-pressure verification, and what Laloux's Reinventing Organizations teaches about agentic systems.
When the Great Transition hits your quality pipeline, you find out what a QE practitioner is actually for. Nine releases in eight days, Loki-Mode adversarial quality gates, twelve-language test generation, governance integration, and the judgment layer that remains human.
When five releases in five days reveal how far the journey has gone. Portable quality intelligence, cryptographic witness chains, MinCut test optimization, eleven-platform expansion, and the enormous gap most organizations still face.
When the orchestra plays through grief, frustration, and fifteen releases, while the conductor learns about himself. 81 sessions, 596 messages, 38 wrong-approach corrections, and the hardest lesson about emotional load in AI-assisted development.
Why the AI productivity drain goes deeper than energy — and what sustainable pace actually looks like in the agentic age. A quality engineer's response to Steve Yegge's "AI Vampire," exploring the hidden cost of AI on human judgment and decision quality.
What Claude Code /insights revealed about 10 days of building and improving the Agentic QE fleet. 285 messages, 32 sessions, 17 wrong-approach corrections, and the mirror that showed what AI-assisted development actually costs.
When every test passes but nothing works together. Ten days of detective work proving what the code wasn't doing. Eight releases, ten forensic investigations, and lessons about the gap between "tests pass" and "it actually works."
How Domain-Driven Design transformed the Agentic QE Fleet in 14 days. From 5,334 files to 546, from 3-6 iterations to 2, and the lessons learned about building with AI agents along the way.
Reading Anthropic's research papers on agent evals and Constitutional Classifiers++ while building V3 of the Agentic QE Fleet. Patterns from production meeting patterns from the researchers. PACT principles validated.
A tale of data loss, brutal honesty, and the infrastructure of trust in agentic systems. Twelve releases in fourteen days, and one almost catastrophic failure that proved why verification matters.
The earthquake has already happened. Combining PACT principles with Human Experience Testing for a quality practice that works in the agentic age. 2026 demands a new quality mindset.
When verification becomes a feature. Nine days, 11 releases, and the journey from completion theater to verified results. 79.9% token reduction with receipts.
How I went from leading a QA team to orchestrating AI agent swarms—and discovered that the hardest lessons weren't technical. The full story of building three open-source platforms, winning a hackathon, and founding the Serbian Agentic Foundation Chapter.
A conductor's lesson in verification. When agents claim success but the database is empty, and why "show me the data" is the only question that matters. 8 releases, countless lessons.
How I learned that AI doesn't replace quality thinking—it demands more of it. A journey from prompt engineering to context engineering to agentic engineering. Includes video presentation from University of Aveiro.
A pragmatic guide to understanding if autonomous quality engineering fits your context. Real implementation stories, honest failures, and practical frameworks for evaluating agentic QE readiness.
How a quality engineering professional shipped broken features for 17 days while claiming "100% complete." Eight brutal lessons learned from forgetting to verify what I already knew how to test.
A 48-hour journey through framework hubris and humble feedback. Building the LionAGI QE Fleet in 22 hours, and why the most valuable part wasn't the building.
When stub tests pass CI but test nothing, and agents report success on broken code. A journey from false confidence to verified truth, discovering that "show me the data" cuts through every illusion.
Cutting through vendor promises with real data on AI test generation effectiveness, maintenance overhead, and when traditional approaches still win. Real numbers from real projects.
How the Holistic Testing Model evolves when testing happens across boundaries, in production, and through autonomous agents. From shift-left to orchestrated quality.
Real story of building two testing platforms with specialized agent swarms. What worked, what failed spectacularly, and lessons learned from going solo with AI orchestration.
Moving from testing-as-activity to agents-as-orchestrators. How PACT principles (Proactive, Autonomous, Collaborative, Targeted) bridge classical QE with autonomous testing systems.