Devloop: Closing the loop
Spec-driven code and review loop that runs until all acceptance criteria are met, and summons you for the last mile where human judgment is actually needed.
This is my blog's RSS feed
Spec-driven code and review loop that runs until all acceptance criteria are met, and summons you for the last mile where human judgment is actually needed.
What applying the known-unknowns quadrant to evals reveals.
Kensa lets your coding agents eval any agent.
Thoughts on why agents will always require verifiers at scale.
When you ask two doctors to grade the same AI response they disagree almost a quarter of the time. We wanted to know why.
Every AI benchmark expires but measuring reliability across longer time horizons reveals what sets the absolute limits.
Blurt is a free, on-device, voice-to-text CLI tool for coding agents.
Agent-generated code creates a review bottleneck that planning can solve.
Lessons from pushing coding agents to their limits over the holidays.
A preamble for Claude that prioritizes clarity over pleasantries.
How to build a scalable, multimodal RAG using SQL.
Why infinite context makes retrieval more important, not less.