Run Claude Code with Your ChatGPT Subscription
Use your existing ChatGPT Plus/Pro/Max quota to power Claude Code via LiteLLM - no separate API key needed. Now with background launchd service and silent startup.
AI engineer, consultant, and founder. Writing about LLMs, distributed systems, and building things that work.
Use your existing ChatGPT Plus/Pro/Max quota to power Claude Code via LiteLLM - no separate API key needed. Now with background launchd service and silent startup.
Google's TurboQuant compresses the KV cache by ~5x with minimal quality loss - what this means for open-source LLM inference.
An open-source Python library that turns text prompts into procedural music in real time, on CPU, sub-second.
An RPG-style simulation of X's recommendation algorithm - pick a persona, write a tweet, and watch the pipeline decide if it goes viral. Real ML running in the browser.
Making LLMs produce reliable JSON - Part 3 of 3. Temperature tuning, graceful degradation, observability, token caching, jailbreak detection.
Making LLMs produce reliable JSON - Part 2 of 3. Multi-model architectures, writer-critic loops, best-of-N sampling, and the freeform thinker + JSON formatter pattern.
Making LLMs produce reliable JSON - Part 1 of 3. Simple schemas, detailed descriptions, constrained decoding, and why enums beat strings.
The post-mortem: 5 things that killed a bootstrapped B2B startup, what's salvageable, and what I'd do differently.
Selling MonitorIntent to startup founders, the Clay defection, the vitamin problem, and watching our moat disappear in real time.
Building MonitorIntent's MVP - 2 months to a Replit dashboard, 5x LLM cost reduction, and the realization that LinkedIn likes are not buying intent.
The first product at MonitorIntent was clever - maybe too clever. AI messages that sounded human but made judgment errors nobody forgave.