RSSAmplifier

Blog

Kyle Wild - dorkitude.com

Essays and bookmarks on technology, AI, and software engineering

dorkitude.comRSS feed ↗20 posts

Latest posts

[Bookmark] IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications by Conor McCauley

Bookmarked on August 04, 2026 A benchmark for whether LLM applications honor instruction priority under realistic conflicts.

[Bookmark] Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations by Mohammadreza Rashidi

Bookmarked on August 04, 2026 A careful look at a dangerous gap between what an MCP approval surface shows and what reaches the model.

[Bookmark] MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation by Kuan Wang

Bookmarked on August 04, 2026 A call for controlled memory benchmarks that do not confuse method gains with pipeline changes.

[Bookmark] Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge by Neeraj Yadav

Bookmarked on August 04, 2026 A focused treatment of stale facts: retrieval memory needs a model of time, not just similarity.

[Bookmark] The Token Tax of Epistemic Accuracy: Comparing RAG and Long-Context Architectures for Document-Grounded Generative AI Applications by Austin Hamilton

Bookmarked on August 04, 2026 A timely comparison of the cost and grounding trade-offs between retrieval and long-context systems.

[Bookmark] BioHarness: Substrate-Aware Evidence Assembly for Biomedical Question Answering across Literature, Knowledge Bases, and Biological Atlases by Meng Xiao

Bookmarked on August 04, 2026 A domain-specific example of assembling evidence across text, knowledge bases, and structured data.

[Bookmark] SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval by Jiandong Ding

Bookmarked on August 04, 2026 A benchmark for the subtle but costly problem of choosing the wrong skill among near-equivalent options.

[Bookmark] Towards Retrieving Interaction Spaces for Agentic Search by Shengyao Zhuang

Bookmarked on August 04, 2026 An argument for retrieving the possible interactions with a corpus, not only its documents.

[Bookmark] Entity-Collision: A Stratified Protocol for Attributing Retrieval Lift in Agent Memory by Youwang Deng

Bookmarked on August 04, 2026 A sharper way to separate real retrieval gains from entity overlap and benchmark leakage.

[Bookmark] Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay by Xiaohua Wang

Bookmarked on August 04, 2026 A provocative replay-first approach to making repetitive agent work cheaper and more predictable.

[Bookmark] Is Grep All You Need? How Agent Harnesses Reshape Agentic Search by Sahil Sen

Bookmarked on August 04, 2026 A study of how an agent harness changes the value of search and retrieval choices.

[Bookmark] Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval by Jeffrey Flynt

Bookmarked on August 04, 2026 A reminder that memory evaluation should measure precision, not reward dumping the whole store.

[Bookmark] Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation by Geert Trooskens

Bookmarked on August 04, 2026 A concrete case for moving recurring LLM workflow work into deterministic compiled artifacts.

[Bookmark] From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents by Meftun Akarsu

Bookmarked on August 04, 2026 A useful wide-angle comparison of retrieval choices for mixed text-and-table RAG.

[Bookmark] How We Improved Agentic Search by Evis Drenova

Bookmarked on May 07, 2026 Search is not a side operation in the agent loop; it is one of the main things the agent does.

[Essay] Codex's `/goal` Is Underrated

I think Codex’s `/goal` is underrated. Easy one-task ralph loops for everyone.

[Bookmark] Content for Content’s Sake by Armin Ronacher

Bookmarked on May 05, 2026 "The fact that it was cheap for you to produce does not make it cheap for someone else to receive."

[Bookmark] A quote from Andy Masley by Simon Willison

Bookmarked on May 05, 2026 "A farmer in Loudoun County sells a few acres of mediocre hay field to a hyperscaler for ten times its agricultural value."

[Essay] Simcluster, and Letting the Agent Play the Meta-Game

I asked Alan to look at Simcluster and think about how he would try to win it.

[Bookmark] Introducing workspace agents in ChatGPT by OpenAI

Bookmarked on April 22, 2026 OpenAI's first claw-like thing since hiring the OpenClaw founder, with UX tight enough for non-nerds