HOT TAKE
Most teams paying for a frontier model would get more from an hour spent on their AGENTS.md.
What moves the needle more? Rules or Model
LAST WEEK’S TAKE
The consequence of asking what should guardrails minimize: a clean sweep.
FREE VIRTUAL EVENT - TODAY
Speakers from eight teams share what holds up once MCP has to handle production security, observability, governance, and multiple agents.
Expect live demos and practical lessons, including source-level findings from 100 MCP servers, agent swarm orchestration, an MCP-powered observability stack, and what the 7-28 spec changes and how to migrate.
Happening today, 15:00-19:00 UTC / 08:00-12:00 PDT. Free and virtual.
HIDDEN GEMS
Built around a portable, provider-neutral harness, Tau reads and edits files, runs shell commands, preserves sessions, and provides a compact reference implementation for understanding coding-agent architecture without a large production codebase.
Tunix high-throughput agentic RL framework
Combining asynchronous trajectory collection, barrier-free training pipelines, modular agent and environment APIs, and lightweight profiling, Tunix helps JAX/TPU teams reduce idle time and identify bottlenecks in multi-turn RL workloads.
Agent developer workspace architecture
Built around Coder-based workspaces, Monaco’s setup combines isolated VMs, cloned databases, prebuilt AMIs, private networking, and controlled credentials to support multiple coding agents working in parallel.
Random-vector reasoning experiment
Testing Qwen3 models across arithmetic, planning, legal reasoning, and debugging tasks, the study measures whether noisy embedding prefixes and plurality voting improve small-model accuracy without fine-tuning or larger hardware.
JOB OF THE WEEK
Agentic AI Platform Engineer // On // London, UK
On’s newly established AI and ML Platform team is hiring an engineer to build and operate production agentic infrastructure, including gateways, evaluation systems, observability, and governance, while supporting product teams adopting agent-based workflows across the business.
Responsibilities
Architect and operate secure, observable infrastructure for production agent systems.
Build evaluation frameworks, routing layers, gateways, and monitoring capabilities.
Guide product teams on scalable architecture for agent-based workflows.
Establish engineering standards and assess tools for shared platform use.
Requirements
Three-plus years’ software engineering experience for mid-level applicants.
Production experience building agentic AI or LLM platform components.
Experience with cloud infrastructure, preferably GCP, and agent frameworks.
Practical knowledge of MCP, vector databases, evaluation, and observability.
MLOPS COMMUNITY
LLM costs are no longer falling fast enough to offset exploding usage, and hidden premiums can turn a viable feature into an expensive mistake.
Cost calculators built into evaluation workflows can flag unscalable experiments before teams spend weeks prototyping.
Token telemetry by service and feature makes spikes, model choices, and experiment costs visible.
Optimization needs researchers, engineers, and finance working from the same latency, quality, and cost data.
Teams that price AI early can ship faster and avoid learning the economics after launch.
An agent can issue a refund in seconds, but the company may later need to show which model ran, what data it used, which policy permitted the payment, and where human intervention was available.
Risk classification follows the agent’s intended use, affected users, data access, and tools rather than model capability alone.
Teams remain responsible for the assembled system, including access controls, human approvals, version records, and logs of tool requests and results.
Material changes to its purpose, permissions, or models may require a fresh assessment.
Traceability and control need to be part of the production design before deployment.
One-shot generation can produce a polished track, but it gives creators few ways to correct one weak section without starting again.
YouAndOrchestra assigns composition and critique to specialist agents coordinated through an iterative conductor loop.
Users can intervene globally, at section level, or down to a bar, beat, or instrument.
Every note carries provenance, showing which agent role proposed it, which musical rule it followed, what feedback redirected it, and how the decision was applied.
The work is forkable and versioned, so earlier states remain available.
The project offers a useful model for agent systems where people need precise control over evolving work.
A coding agent can pass every test and still build the wrong feature. Agent Levers puts the workflow on disk so intent, verification, and stopping conditions remain visible.
Tasks move through a plan, do, check, and act cycle with explicit acceptance criteria.
Verifiers run from a fresh shell, while iteration budgets stop failed loops from expanding the diff.
Recurring mistakes become proposed rules in AGENTS.md for future sessions.
The result is less supervision without surrendering control over what gets shipped.
IN-PERSON EVENTS
New York - August 14
Toronto - September 10
San Francisco, Voice Agents Forum - September 16
VIRTUAL EVENTS
Virtual MCP event - August 6
MEME OF THE WEEK
ML CONFESSIONS
Three years ago someone senior pushed for us to move every training pipeline onto an orchestration framework they’d used at their last company. I argued against it in three separate meetings. What I said was that it would add operational burden we didn’t have the headcount for, and that we’d be rewriting DAG logic that already worked fine.
What I thought was that I’d just come off a fortnight of a different migration and could not face learning another set of abstractions.
We didn’t move. About a year later the main maintainer stepped back and the company behind it got bought. The org that had standardized on it spent most of last year moving off. I’ve still got the Slack message from one of their engineers saying I’d called it early.
Nobody has asked me to explain the reasoning. I get pulled into architecture reviews for this now, and there’s one next week about an agent framework. I’ve read maybe half the docs and I already know what I’m going to say.
Share your confession here.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.