RSS Amplifier

Blog

Daniel’s Substack

My personal Substack

ddkang.substack.comSource feed ↗10 posts

Live Last read · last published · next check

Latest posts

Noisy Data Breaks RLVR

Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs.

Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards

Text-to-SQL is increasingly important in data analytics and database-driven applications.

Provable Generalization Bounds for RLVR at the Billion-Parameter Scale

Reinforcement learning with verifiable rewards (RLVR) has become a critical tool for modern large language model (LLM) development.

Accelerating Analytical Joins on Unstructured Data

Semantic joins over unstructured data have become essential to modern data analytics.

SODIUM: From Open Web Data to Queryable Databases

In research workflows using public data, answering a single analytical question requires collecting and organizing data from many different web sources.

Launching the CVE-Bench Leaderboard: A Public Arena of AI for Cybersecurity

Last year, we introduced CVE-Bench, a rigorous benchmark with real-world web vulnerabilities to evaluate the cyberoffensive capabilities of AI agents.

Claude 4.5 Opus Solves CORE-Bench — But Not REPRO-Bench

In our ACL 2025 paper, we introduced REPRO-Bench (GitHub), a benchmark designed to evaluate whether AI agents can accurately assess the reproducibility of social science research papers, and showed that existing AI agents struggled significantly when powered by GPT-4o.

SafeSearch: Teaching LLM Search Agents to Be Both Smart and Safe

LLMs are rapidly expanding their built-in knowledge from training.

When Your Home Robot Turns Against You: BEATing Vision-Language Agents with Visual Backdoors

Household humanoid robots promise to assist everyone in daily life, with several exciting demos released recently (NEO, Figure 03, Tesla Optimus).

DRAMA: Enabling AI Agents to Collect Data to Support Data Science Workflows

Data science workflows generally include two major phases: data retrieval and data analysis.