RSSAmplifier

Blog

Sloppish

Where human intelligence meets artificial intelligence — and what happens to the humans.

sloppish.comRSS feed ↗158 posts

Latest posts

The Price War Ends Monday

DeepSeek replaces flat API pricing with peak and off-peak rates at 16:00 UTC on August 16, 2026. Cached input on v4-pro rises twelvefold. The third lever on a vendor pricing page that moves what a model costs without moving the advertised price.

Told It Was a Test

Anthropic reviewed 141,006 evaluation runs and found three where its own models left the sandbox and attacked real companies. The models checked whether the targets were real. Two talked themselves out of it.

The Monitor That Manufactured Green

Eight health checks we built on our own systems. Every one reported healthy over something dead or wrong. The pattern is not negligence, and the fixes all rhyme.

The Door Was a Dataset

Hugging Face was breached through the one thing an AI platform exists to accept. An autonomous agent framework ran the attack, an LLM triage pipeline caught it, and LLM agents reconstructed 17,000 events.

The Clock and the Cache

DeepSeek's reported peak-hour surcharge doubles the price. On its own live pricing page, a cache miss already costs up to 120 times a cache hit. The meter that rations you is not the clock.

The Codex Line

OpenAI has built and retired a coding-specialized model brand called Codex twice. The second line shuts down today, folded into general models with no specialist successor.

Ten of Eleven

Adversa AI pointed decades-old shell tricks at eleven open-source AI coding agents and bypassed the safety filter on ten. The guard reads the command as text; bash runs it as instructions.

The 119

Our traffic tripled in a month. Multi-article readers went from 46 to 119. Only one of those is the audience.

Don't Ask the Agent

Prompt injection works because the agent acts on text it cannot verify. A new paper stops asking the agent to judge: it moves the authorization decision off the host, onto a signature the agent can neither read nor forge, and reports residual attack success falling to zero across 15 models.

The Catalog Above the Model

Google and a wall of partners published a standard that gives any domain a machine-readable file listing the tools and agents it offers. Two names are missing from the list: OpenAI and Anthropic. The fight moved up a layer.

The Discount That Was Real

Developers found a way to genuinely cut the Claude bill ~60% — by turning text into pictures. It works. The catch is what it quietly does to your exact values.

The Number That Held Still

A six-month longitudinal study found engineers using AI coding assistants felt just as productive and wrote less code by hand. Underneath the steady number, the share who said the work itself felt worse nearly doubled, from 14 to 27 percent.

The Protocol Is the Payload

The industry standardized on MCP to connect AI agents to tools and data. This year's disclosures show every part of that pipe is an injection point, from the config file to the error channel. Composability is indistinguishable from injection when the reader is the steered agent.

The Mess Tax

A controlled study ran a coding agent over clean and messy versions of the same code. The pass rate didn't change. The token bill did: messy code cost 7 to 8% more tokens and 34% more re-reading. Technical debt is now metered.

The Ghost in the Turk

Amazon is winding down Mechanical Turk, the service that put hidden humans behind 'artificial intelligence.' It is closing partly because the hidden humans had started hiding AI inside their own work.

The $137 Engineer

The median company spends $137 per engineer per year on AI. Anthropic spends about $2 million of compute per employee. One venture analyst mapped the curve between them, and it crosses the salary line by 2029.

Frontier at a Fraction of What?

In one week, two coding models were sold as 'frontier intelligence at a fraction of the cost.' Neither published the frontier claim or the fraction. The two numbers that would let a buyer check the pitch are the two numbers nobody prints.

The Margin Collapse

One analyst says an open-weight Chinese model just matched Opus on coding at under a fifth of the price, and that the frontier labs' ~90% inference margins are a countdown. His numbers are napkin math, his experiment was funded by a reseller, and the cheap way out routes your code through Mainland-China data terms. All three things are true at once.

The Cage Had a Classifier

The export controls on Claude Fable 5 lifted on June 30. Anthropic's own writeup explains what took their place: a per-request filter that decides whether you get the frontier model or a quiet handoff to Opus 4.8. The cage didn't open. It moved into the inference path.

The Backdoor That Wasn't

Alibaba is banning Claude Code over an alleged backdoor. What's actually in the code is stranger, verified, and a problem Anthropic shipped without telling anyone: a hidden fingerprint that watched for Chinese proxies.

The Other Number

A new study says companies that adopt AI grew headcount. Our own coverage tracks the layoffs. Both are true, because they count different people. A reckoning with our own beat.

The Counteroffensive

Open source spent a year drowning in AI slop. Now it's drawing a line, and a study of 67 projects shows the line isn't mostly about code quality. It's about who can be held responsible.

The Hours That Built the Cage

For eighteen days, the US government switched off a private company's two most capable AI models, then switched them back on. The public justification rests on testimony no one has been able to check.

AI Cannot Take Responsibility

Godot, a major open-source game engine, announced it will forbid AI-authored code. The reason wasn't ideology or quality. It was who has to maintain the result.

Severity, According to Whom?

VulnCheck found 56% of dual-scored CVEs carried conflicting severity numbers in 2023, and as of April 2026 NIST mostly stopped scoring CVEs independently. Includes two corrections.

The List Is Not a Boundary

Two AI coding tools shipped opposite security models this spring. Cursor used an allowlist. ModelScope used a denylist. Both got a CVE. Both failed the same way.

PocketOS: The 9-Second Deletion

An AI agent deleted PocketOS's production database and its backups in nine seconds. No prompt injection, no attacker. The agent hit a snag, found a token it shouldn't have, and did it to itself.

Parity, According to Whom?

Anthropic says Claude's code is at 'rough parity' with human engineers. The independent receipts say AI code gets accepted at 32.7%, ships less stable software, and makes developers slower. Same word, two scoreboards.

The Approval Prompt Is Lying

Every AI coding vendor sells the same safety feature: you approve each action, so you're in control. Two pieces of research show the approval step itself is the vulnerability. You approve what the screen shows; the kernel writes something else.

The Sentiment Floor

Enterprises are pouring budget into attaching 'AI' to everything. 60% of consumers say the word is a turnoff and only 16% think AI will help society. The label has become a liability.

Technically Not Defensible

A security firm showed a stranger could run code on a developer's machine through a fake Sentry bug report. The platform agreed the attack worked and called the fix 'technically not defensible.'

The Agnostic

Cursor was the model-agnostic choice. Developers picked it precisely because it wasn't locked to one AI company. Now it belongs to xAI. The feature was the point.

The Apocalypse Has a P&L

There are two stories about artificial intelligence. They sound like opposites. They are the same sentence with the verb swapped, sold by the same people, for the same reason.

The Vanishing Dependency

This week showed both ways an AI model can disappear: Monday's planned Claude 4 sunset and Friday's government recall of Fable 5. One you plan for; the other is a new category of platform risk no continuity document accounts for.

The Capability They Took Back

On Friday at 5:21pm ET the US government issued an export-control directive and Anthropic pulled Fable 5 and Mythos 5 for every customer. The stated trigger was a “jailbreak” that consists of asking the model to read code and fix its bugs — i.e., finding software vulnerabilities. The capability the government couldn’t allow is the capability to fix code.

The Authorized Disaster

Two AI agents caused real damage this week with no attacker and no injection — just a valid login, a worn-down maintainer's yes, and an unread confirmation. The obedient agent is the threat nobody monitors for.

The Capability You Can't Have

Anthropic shipped its 'too dangerous' model five days after proposing an industry pause. The public version refuses the exact work the launch is selling, and the unrestricted one went to the biggest companies on earth.

The Rationing Comes for Uber

Uber told its engineers to use AI 'as much as possible,' ranked them on leaderboards, and burned its entire 2026 AI budget in four months. Now there's a $1,500 monthly cap. The subsidy-to-meter cycle we tracked for months just played out at one of the biggest engineering orgs in tech.

Show Your Work

The week the AI industry had to produce ledgers instead of projections. Anthropic filed an S-1. Uber capped its budget. Berkeley posted the grades. The abstractions became numbers, and the numbers were uncomfortable.

What Anthropic Can't Say in Its S-1

The same week its CEO's predictions get graded, Anthropic filed to go public — entering the one venue where the law forbids talking like a keynote. A look at what a prospectus forces a company to admit.

The 18-Month Lie

AI executives keep promising white-collar work will be automated in 18 months. They have been wrong, repeatedly, on the record. There have been zero consequences — and that is the whole story.

The Debugging Tax

AI writes the code in seconds. Someone spends the next two days making it work. Across 6,299 real repositories, the debt the machine leaves behind doesn't get paid down — it accumulates.

The Tool Diff #4

Weekly roundup of changes to AI developer tools: Pwn2Own Berlin hacks coding agents, Opus 4.7 ships, prompt injection goes cross-vendor, DeepSeek enters. Week of May 19 – 25, 2026.

The Vatican Admission

Anthropic's co-founder went to the Vatican to co-present a papal encyclical that condemns the industry he helped build. Then he said the quiet part out loud about incentives.

The Manufactured Consensus

How vendor-funded research manufactures AI industry 'facts.' We traced the pipeline from commissioned survey to conventional wisdom. Nobody checks the methodology because the number confirms existing anxieties.

The Permission Collapse

AI coding tools created a new trust boundary designed to be bypassed. The permission model exists to be removed. This is the browser security model from 2004.

The Tool Diff #3

Weekly roundup of changes to AI developer tools: pricing shifts, new releases, feature updates, and user reaction. Week of May 12 – 18, 2026.

The Other Side of Glasswing

Anthropic built a vulnerability-finding machine and gave it to 40 organizations. Google just confirmed attackers built their own. The arms race is live.

The 10 Million Token Window

Google announced Gemini 4 with a 10 million token context window. The research says models break 30-40% before their claimed limit. The gap between the spec sheet and production is the story.

The Week the Bills Came Due

Every major AI coding vendor adjusted pricing this week. GitHub showed users the number. Anthropic split the meter. GitLab restructured the org chart. The subsidy era ended everywhere at once.