Most AI benchmarks tell you if a model can solve a LeetCode puzzle. They don’t tell you if it can ship a product.
I got tired of guessing which model to use for actual work. So I built gmickel-bench.
It is a living scoreboard based on the real-world engineering tasks I actually ship:
Full-Stack: Building OAuth 2.1 servers on Convex.
Frontend: Designing comp…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.