RSSAmplifier

Blog

blank

jashvira.comRSS feed ↗10 posts

Latest posts

Initials Atlas

Notes on Variational Inference and Jensen’s Inequality

Spatial Competence Benchmark

List of open source contributions to Inspect AI

Preview: Visual Geometry Bench

What do we measure?

Hill climb on MBPP using verifiers

From scratch: SFT and GRPO on Qwen 2.5

Intuiting Policy Gradient methods

Recently, I found it imperative to grok Policy Gradient (PG) methods. As much as I enjoy entering rabbit holes of adjacent techniques, which are abundant in RL, I have refrained. The motivation is to think effectively about PG research in LLMs/foundational models.

LLM Agent ~ Reddit Consensus

Neural Network precision pitfalls in the wild

In my work, I draw on concepts from computational geometry and graphics, often demanding high-precision guarantees. A typical case is sampling an object’s implicit function to generate a parameterised form for downstream use.