RSSAmplifier

Blog

june.kim

Blog posts by June Kim

june.kimRSS feed ↗483 posts

Latest posts

test post

Interface-Driven Code Scavenging

Nested Agentic Iteration

The Bible Is Terminal

Replayable Claims: Epistemic Status Without Trusting the Sender

AI Safety is Reinventing the Law

Auditing SlopCodeBench

τ²-bench Doesn't Check the Rules

The Complementarity Test

The Virtues of Typed Reasoning

Borrowed Vocabulary for the Hypothesis Edge

Why There Are No Ads in Your Chatbot

Intent Extraction in the Wild

The agent is the inbox

Auditing Frontier-Bench

Assurance at the Boundary: The Level Below AAL-1

Game of Benches

An Epistemic Ablation

Auditing FrontierCode

Auditing MirrorCode

Vulnerability Disclosure Queue Overload

Compensating Controls

PPC Science

Adtech Is an Accounting Problem

Hiring Is Evals

797,444 Impressions, 7 Clicks

How to Audit a Benchmark

Trust Is a Cache

Terminal-Bench Is Blind to Destruction

Auditing DeepSWE v1.1

Nodewise E-Values Under Graph Interference: Causal Abuse Filtering for Agent Platforms

The Power Diagram Auction: A Formally Verified VCG Mechanism for LLM Advertising

Regeneration Re-Prices Contamination

Who Authored the Spec? Evidence for the Line-Drawing Question Thaler Left Open

Union-Find Compaction: Provenance-Preserving Context Compression for LLM Agents

The Kilogram Was Losing Weight

ProgramBench Measures Recall

Generalize or Specialize? Retaining Reusable Skills for World-Model Agents

Detecting and Refining Ambiguous Specifications, Automatically (DRAFT)

A Determinacy Audit of SWE-bench Pro

Go Check

Verifiable Knowledge

You Cannot Ring a Semiring

What Cannot Be False Cannot Be True

Wrong Again

Wrong Questions

Compress and Unfold

How Not to Run SWE-bench Pro

Precisely Wrong

Truly Untrue?