RSSAmplifier

Blog

Prabal

AI engineer, consultant, and founder. Writing about LLMs, distributed systems, and building things that work.

prabal.caRSS feed ↗11 posts

Latest posts

Run Claude Code with Your ChatGPT Subscription

Use your existing ChatGPT Plus/Pro/Max quota to power Claude Code via LiteLLM - no separate API key needed. Now with background launchd service and silent startup.

Google Made Long Context 5x Cheaper

Google's TurboQuant compresses the KV cache by ~5x with minimal quality loss - what this means for open-source LLM inference.

LatentScore: Text to Real-time Music on CPU

An open-source Python library that turns text prompts into procedural music in real time, on CPU, sub-second.

The X Algorithm, Visualized

An RPG-style simulation of X's recommendation algorithm - pick a persona, write a tweet, and watch the pipeline decide if it goes viral. Real ML running in the browser.

Shipping in Production

Making LLMs produce reliable JSON - Part 3 of 3. Temperature tuning, graceful degradation, observability, token caching, jailbreak detection.

When One LLM Isn't Enough

Making LLMs produce reliable JSON - Part 2 of 3. Multi-model architectures, writer-critic loops, best-of-N sampling, and the freeform thinker + JSON formatter pattern.

Schema Design

Making LLMs produce reliable JSON - Part 1 of 3. Simple schemas, detailed descriptions, constrained decoding, and why enums beat strings.

The Post-Mortem

The post-mortem: 5 things that killed a bootstrapped B2B startup, what's salvageable, and what I'd do differently.

They Wouldn't Stay

Selling MonitorIntent to startup founders, the Clay defection, the vitamin problem, and watching our moat disappear in real time.

Scraping LinkedIn

Building MonitorIntent's MVP - 2 months to a Replit dashboard, 5x LLM cost reduction, and the realization that LinkedIn likes are not buying intent.

The AI Trust Deficit

The first product at MonitorIntent was clever - maybe too clever. AI messages that sounded human but made judgment errors nobody forgave.