Designing Verifiable RL Environments for DevOps Agents
What separates a rigorous, trainable agentic task from a vibe-coded one: reproducible sandboxes, hidden verifiers, difficulty calibration, and graders that resist reward hacking.
This is my portfolio RSS feed
What separates a rigorous, trainable agentic task from a vibe-coded one: reproducible sandboxes, hidden verifiers, difficulty calibration, and graders that resist reward hacking.
A deep dive into building a vector database from scratch in Python, understanding embeddings, distance metrics, and semantic search.
How I built HybridRAG - an enterprise search engine that balances quality and cost through intelligent model routing.