We’re open-sourcing our fork of Harbor , an RL environment sandboxing platform, that adds support for scalably doing RL training with Modal and Tinker. Context Reinforcement learning is all the rage today for pushing the frontier of our LLMs. Teams build environments that simulate some task, and teach the models how to navigate and work with those environments. We’ve been experimenting…
(This is a companion piece to my substack post ) A good chunk of my professional career has been spent trying to spend more money . Pre-LLMs, this was trying to scale fraud pipelines at close to Google search scale. With LLMs, I’ve spent probably O($1M) to programatically label and distill data. Being able to spend that money somewhat intelligently requires real infrastructure and planning…
Since the first time anyone said “coding agents,” I’ve basically only wanted one thing: an intern I can assign tasks to, whenever they pop into my mind, and expect results eventually. Claude Code and OpenAI Codex are getting there, in terms of the pure capability. But they’re also chew through tokens quickly, add huge markups (4-5x) on sandbox usage, and are not that easily…
For a pile of sand, people really seem to think LLM’s are good at understanding human emotion and psychology. I certainly agree that they’re helpful. But as an AI researcher, I know that RLHF training pushes them towards responses that people like - not necessarily what they need to hear. Sycophancy is problematic in many areas, but especially when trying to resolve conflicts: if both parties talk…