
Noisy Data Breaks RLVR
Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs.
My personal Substack
Live Last read · last published · next check

Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs.

Text-to-SQL is increasingly important in data analytics and database-driven applications.

Reinforcement learning with verifiable rewards (RLVR) has become a critical tool for modern large language model (LLM) development.

Semantic joins over unstructured data have become essential to modern data analytics.

In research workflows using public data, answering a single analytical question requires collecting and organizing data from many different web sources.

Last year, we introduced CVE-Bench, a rigorous benchmark with real-world web vulnerabilities to evaluate the cyberoffensive capabilities of AI agents.

In our ACL 2025 paper, we introduced REPRO-Bench (GitHub), a benchmark designed to evaluate whether AI agents can accurately assess the reproducibility of social science research papers, and showed that existing AI agents struggled significantly when powered by GPT-4o.

LLMs are rapidly expanding their built-in knowledge from training.

Household humanoid robots promise to assist everyone in daily life, with several exciting demos released recently (NEO, Figure 03, Tesla Optimus).

Data science workflows generally include two major phases: data retrieval and data analysis.