RSSAmplifier

Blog

Ibrahim Ahmed

personal website and blog

ibrahimahmed.caRSS feed ↗7 posts

Latest posts

Fast RL using off-policy sampling

First open-source implementation of Soft Policy Optimization, an off-policy RL algorithm that works with LMs. This makes many RL experiments faster and cheaper.

Real Work

We are developing the first open-source LLM RL environment framework for real work.

How to Vibe Code Effectively

My intuitive and counterintuitive learnings

Proposal: Self-Refined RL (SRRL)

Policy gradient RL algorithms like GRPO have been used to improve LLMs' performance on verifiable tasks like math and coding problems.

How scaling pretraining affects RL sample efficiency

Insights from a small-scale transformers experiment.

Bugs in LLM Benchmark Grading

I used Claude to audit the grading code of 8 major LLM benchmarks and found issues throughout all of them.

The path from Fable to superintelligence

How real-world feedback loops could turn capable AI agents into recursively improving systems.