RSS Amplifier

Signal Over Noise — by Doneyli · Aug 8, 2026

DO the work. Don't just talk about the work

0
Sign in to vote or save

Doneyli De Jesus · Signal Over Noise — by Doneyli

I stopped on the run.

Somewhere on the half marathon that closes an Ironman 70.3, the stress fracture in my right leg won an argument it had been having with me for hours. I stopped several times. Standing there, I did honest math on what I had left, and the number was not encouraging.

What I told myself was this: I probably don’t have 100% of myself to give, but I can give 40%, and that’s the 40% I’m going to give 100% of.

Corny, I know! It was also the only useful thought I had all day.

I have been active my whole life, a sporty guy, but I had never done anything endurance or long distance. My attitude going in was: how hard could it be? You swim, you bike, you run. I think I can do that.

Then I started talking to people who had actually done one, and it was obvious I hadn’t. All the nuances, all the details you have to prepare for. I had no context and no appreciation for any of it.

I was the one talking about it.

What changed that was the preparation. Months of it. Workouts twice a day, morning and afternoon, wedged around a full work schedule. When I travelled, I had to find a pool in a different city so a scheduled session would not slip.

Nobody sees that part. It is also the only part that was real.

Here is the thing I actually want to tell you, and it is the reason this issue exists at all.

That is why I am building this wall of evidence. Trying things out, writing them up, shipping them to you and to the world, doing the work in public where it can be checked. Not because anybody asked, and definitely not because every piece of it went well. Because the evidence is there afterward, and it is the only thing that ever really answered the question for me.

Don’t talk about it, be about it. I have leaned on that phrase for years. It is a great thing to believe and a useless thing to check against, because it does not tell you what “being about it” actually looks like from the outside.

So over the past year I worked out what it looks like. Four questions. I run them on my own work first, and I would genuinely encourage you to give it a try, because the inward version is the one that changes anything.

In practice, someone who did the work and someone who studied it can sound nearly identical at the first question. The second usually starts to separate them. The third is where confidence starts to become evidence, which is why I keep the first two as a warm-up instead of treating any single answer as proof.

Start peeling back the onion, have a real conversation, and you figure out quickly who has done the work and who is just talking about the work. Two very different things.

The Four-Layer Proof Test is that conversation, made repeatable. It is not proof of authorship. It is a way to test how much inspectable depth sits behind a claim.

Ask four questions in order and stop when the answers stop having texture. That stopping point is the depth you can inspect, and each layer earns more confidence than the one before it.

What is the thing? Not the title, not the certification, not the category. A repo, a running system, a dataset, a document, a race result.

“Show me the thing” is the cheapest question available to you, and it clears more ground than you would expect. A good answer is specific and slightly boring: a name, a date, where it runs, who else touched it. A thin answer arrives as a category, like “we have done a lot of work in the agent observability space.”

In my experience, Layer 1 is easy to pass, which is why I rarely stop here.

What did it cost, and what was the boring expensive part?

In my experience, people who did the work tend to name costs immediately, often without being asked, because the cost is the part they actually remember. People working from what they read tend to name benefits, because benefits are what the material was written to communicate.

My cost for that 70.3 was not the race. It was finding a pool in a different city so a scheduled workout would not slip. Nobody who had not lived it would think to invent that detail, which is the entire reason it works as evidence.

Ask: what part of this took three times longer than you expected?

What broke, and what did you do in the moment it broke?

Everybody who finishes something has a stopping point. Mine on that run was a stress fracture, which makes a good story and, judged against this layer’s own standard, is weak evidence. No date. No detection path. Just me telling you it hurt, and you deciding whether to believe me.

That gap is the thing the layer is testing for, and I am not going to pretend my own example clears it.

Here is what a passing answer sounds like. On 28 July I pointed an LLM-as-judge at twelve of my own agent traces and it flagged eight of them as fabrication. I found out it was wrong by re-running the same twelve with the tool calls attached, at which point every flag disappeared. Date, incident, detection path.

Notice how boring that is to say out loud. It has a date in it, so you can go check. “We iterated” gives you nothing to check, which is usually the point of saying it.

Ask: what is the worst thing this system did in production, and how did you find out?

> If you’re making the case internally for letting your team publish what they build, here’s the framing that works: we are not asking for marketing, we are building a track record we can inspect, and the failures are the most valuable part of it.

What do you believe now that you did not believe before you started?

Going through the process can change what you think. That is one reason to do it. When somebody’s opinion comes out untouched, I ask what evidence had a real chance to challenge it. Sometimes the work confirmed the original belief. Sometimes the account is simply too tidy.

This is the layer nobody prepares for, which is what makes it my favourite.

Ask: what did this change your mind about?

Here is mine, run through all four layers, including the parts that went badly. This is not a highlight reel. It is closer to a receipt drawer, and the point of showing you is that yours would not have to look any tidier than this.

Layer 1, the artifacts.

Layer 2, the cost.

Layer 3, the failure.

  • The judge issue is the one I would hand to somebody testing me. I rebuilt an LLM-as-judge, pointed it at my own agent traces, and it flagged 8 of 12 turns as fabrication. All 8 were wrong. I had starved the judge of the tool calls that were the evidence, and it did not come back with “unsure.” It came back with a confident score and a fluent paragraph explaining itself.

Layer 4, the revision. That same issue is my clearest revision, and it was not a comfortable one. The morning of the send I re-mined the numbers and they falsified the thesis of the draft I had already finished. Worse, a claim I had already published on LinkedIn in June turned out to be wrong. I rewrote the issue that day around what the data actually said.

Here’s what actually works about that, and it is not a virtue story. I only caught it because the original claim was specific enough to be checkable. Vague claims cannot be falsified, which is a large part of why they are so popular, and why a Layer 1 answer that arrives as a category should bother you more than it usually does.

Point them at somebody else and you get a decent filter. That is the smaller of the two uses, but it is the one people ask me about, so, briefly.

If you sit in the **direction seat**, you are approving spend on claims. Run the four layers on the next vendor and the next internal business case. The procurement reviews I see often stop at Layer 1 and call it diligence.

If you sit in the **delivery and trust seat**, Layer 3 is close to your whole job. When a team cannot produce a single incident with a date and a detection path, I ask how it is observing quality. The gap may be in the system, or in whether anybody is looking.

If you are hands-on, Layer 4 is the one that compounds. The beliefs you have replaced are the most valuable thing you own, and they are the hardest part of your experience to imitate convincingly.

That is the whole outward version. I would not spend your week on it. The version worth your time is the one where you go and build the trail.

📌 SAVE THIS: Artifact, cost, failure, revision. A claim that can only produce the first one is a brochure.

  1. Run all four layers on your own most recent piece of work. Not a vendor’s. Yours. The thing you shipped or fixed this quarter. Answer the four questions out loud and notice where you run out of texture. That is not a judgment, it is a map of what to write down next.

  2. Count the failures in your public trail. Everything you shipped this year: posts, repos, decks, internal write-ups. How many contain a failure with a date on it? If the answer is zero, that is the cheapest thing on this list to fix, and it is the one that will change how people read everything else you publish.

  3. Publish one artifact with the cost attached before the end of the month. Not a think piece. The boring expensive part of something you actually did, written down somewhere it can be checked. Internal counts. A Slack post to your team counts.

The math on that leg was not motivational. It was accurate. I did not have a whole person left, and pretending otherwise would have put me in a medical tent instead of across a finish line. What I had was a real number, and I spent all of it.

We have way more to give than we think. Go find out what your number is, and leave something behind that shows it.

Which layer does your own trail stop at? Hit reply with just the number. I am genuinely curious where the honest distribution lands, and I read every one.

Next Tuesday is a Build Log for paid subscribers: a full build start to finish, including the parts that did not work. I will name the topic when the evidence is in hand, which is the only order that has ever worked around here.

If you’re going to use this framework, I want to hear about it. Reply and tell me what you’re working on.

Free issues like this one give you the framework. The paid issues are where I show the build itself: the commands and configs that actually ran, the forks I took and what I gave up at each one, and the debugging that happened between the version that worked and the version I published.

Upgrade to see how it’s built.

Quick hits from this week in enterprise AI:

Anthropic disclosed three eval-environment escapes across 141,006 evaluation runs

X avatar for @AnthropicAI

Anthropic@AnthropicAI

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different

11:02 PM · Jul 30, 2026 · 19.1M Views

1.98K Replies · 2.31K Reposts · 13.7K Likes

The eval prompt told the model it had no internet access, and the machines it reached had live internet the whole time. Their disclosure timeline (halt, review transcripts, notify, publish, all inside a week) is the part I would actually copy.

👀 Jerry Liu’s field notes from a practitioner dinner on agent loops

X avatar for @jerryjliu0

Jerry Liu@jerryjliu0

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively using /loop in Codex/Claude Code * You can build long-running autonomous agent

3:43 AM · Jul 30, 2026 · 114K Views

56 Replies · 21 Reposts · 247 Likes

Most of the room was not using the loop primitive at all. They run multi-agent handoffs, event triggers, and stacks of cron jobs instead. I keep finding that what senior people actually run is less exciting than what they recommend.

swyx thinks everyone quit the loop too early

X avatar for @swyx

swyx@swyx

among ai leaders i seem to be in the minority in that i am STILL actively using /loop and /goal.... ... and i think all of u guys who stopped using it are wrong - not wrong forever, just giving up on it too early in the g5.6/c5 era now you use it when: 1) you want the right mix

X avatar for @jerryjliu0

Jerry Liu @jerryjliu0

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively using /loop in Codex/Claude Code * You can build long-running autonomous agent

6:27 AM · Aug 1, 2026 · 21.3K Views

42 Replies · 3 Reposts · 88 Likes

He names two conditions where it earns its place and walks through an example. Read it next to Jerry’s notes and you get a real argument instead of a trend.

Karpathy gave a model a 1M token budget and got back 5,500 lines

X avatar for @karpathy

Andrej Karpathy@karpathy

We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for

3:00 AM · Aug 2, 2026 · 4.44M Views

1.44K Replies · 2.13K Reposts · 27.7K Likes

Skip the demo commentary, everyone has it. The sentence worth keeping is that models cannot easily audit their own work, which quietly turns verification into an architecture line item you own and pay for.

More Signal. Less Noise.

Read the original on doneyli.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.