Research
benchmarks, reference tables, literature reviews.
Software Correctness Tools Compared
43-tool reference table across runtime verification, model checking, SMT solvers, theorem provers, and verified artifacts.
Benchmarking Four Approaches to Agent-Assisted Effect Code
Curated reference vs. human docs vs. raw source vs. nothing. Five tasks, regex judges, quantitative results.
effect-first research (live site)
Active research front-end: what guidance helps an agent write stronger Effect code, and what turns out to be noise.
What Survives When the Context Window Resets
A 75-call GPT-4.1 experiment comparing five context representations across five knowledge categories.
Cursing Agents
Comparison of three instruction tones across 36 OpenCode runs, scored with executable hidden tests.
What Happens Around Compaction
Academic papers, builder experiments, and a few who skip memory entirely. Literature review on context and memory.
20 Ways to Look at Agent Memory
A catalogue of 20 generated interface concepts for agent memory, kept as ideation rather than evidence.