RSS Amplifier

Kaci Nguyen · May 11, 2026

tokenmaxxing is just a phase, mom

0
Sign in to vote or save

Kaci Nguyen · Kaci Nguyen

It’s been roughly 90 days since Intuit launched its 3x productivity initiative, provisioning Claude Code access to every single employee. They’re just one of many enterprises making the same bet to become AI-pilled and triple their productivity or risk obsolescence.

Across the industry, this bet has split organizations into two camps based on how they think about consumption:

On one side are the inputs-as-virtue leaders, who treat token burn as evidence of cultural transformation. They look at their workforce and conclude that human behavioral inertia is a Sev 0 threat. Setting $50,000 on fire so an employee can learn to use 10 billion tokens is a cheap tax for future-proofing the company. On the other side are the inputs-as-surveillance leaders. Performance reviews may be tied to consumption, and as a result, employees generate reams of unverified, fluffy, polished-looking AI slop.

Both camps are chasing the same 3x, but neither can verify they’re getting it. Token consumption can’t tell you whether the work is better, faster, or worth more than it costs. It can only tell you that tokens were used. Employees, sensing they’re being measured on consumption, are churning out volume to max out their allotments regardless of usability. Management watches usage climb and calls it transformation. It’s a textbook Goodhart’s Law failure: we’re incentivizing raw consumption with the goal of productivity, but optimizing for the former is corrupting the latter.

The instinct is to pick a side: either champion tokenmaxxing and accept the slop, or outright demand outcomemaxxing before employees have any foundation for what good looks like, disengage, or fall back on what they already know. But both are wrong when viewed as fixed policies. They miss the reality of how human beings learn. Tokenmaxxing and outcomemaxxing are just sequential phases of organizational maturity, and Goodhart’s Law applies differently in each.

The mistake most companies make is applying a point-in-time, static policy to a dynamic learning curve. You can’t demand ROI from someone who hasn’t figured out what the tool can do, and you shouldn’t subsidize brute-force burning from someone who has. To navigate this, management has to guide employees through three distinct phases, and recognize when to move.

One structural note: these phases apply per function rather than company-wide. I’d expect engineers running Claude Code through agentic loops to exit Phase 1 in weeks while a marketing team experimenting with drafting and synthesis to need months. For larger organizations with greater heterogeneity, there will likely be two or three phases running at once as the expected state.

In the first 90 days of rollout, the primary goal is behavioral change, not efficiency. Any workforce will sort into roughly three groups:

  1. Non-believers who tried it once, got a mediocre result, and abandoned.

  2. Eco-conscious budgeters who feel guilty of wasting electricity and money and self-cap before they discover the upside (btw, here’s how much electricity it uses)

  3. Power users – but at this stage, every power user is a Slop Cannon. Turbo Brains don’t exist yet; they’re Slop Cannons who survive Phase 2.

I’d argue Groups 1 and 2 deserve the most attention because re-activating a churned skeptic is harder than coaching an enthusiast towards discipline.

Management should give everyone an effectively unlimited budget, mandate the latest models, and actively tell employees to be wasteful on purpose. Push the models until they break. Try the absurd prompt. Burn an afternoon on a workflow that ends up not panning out. The goal is to trade short-term financial efficiency for long-term behavioral transformation.

The math checks out: If an employee’s salary is $150,000, capping them at $500/month in tokens to “protect the bottom line” is absurd if uncapped access could plausibly make them 3x more productive. The $50K you spend on tokens is buying you the option on a 6-figure productivity gain you can’t access any other way.

The Goodhart trap in Phase 1 is the inverse of what you’d think: you want people to chase the consumption metric, because the friction of adoption is the actual bottleneck. Slop is the cost of breaking the blank page problem.

You’d want to exit Phase 1 when weekly active usage stabilizes above ~70% without prompting, and senior ICs start complaining that reviewing AI output is eating their week. The second signal is the more important one because it means you’ve succeeded at adoption and created a new problem.

Phase 1 succeeds when senior ICs and directors become slop janitors – drowning in volume from juniors who’ve perhaps learned to generate but not to judge. This is the most dangerous phase, because it looks like the system is working (token usage is up! adoption is universal!) when the work has just been offloaded to a different (most expensive) group of people.

The single highest-leverage move here is to shift review from outputs to prompts. Borrowing from healthcare, directors should become AI preceptors. When a junior submits sub-par work, the feedback isn’t about the deliverable. It’s:

Show me your context window and your prompts. Let’s refine your .md files.

This single reframe does more than any policy change. It turns every review into a teaching moment, it makes prompt quality a visible artifact, and it kills the incentive to ship volume – because volume without good context just produces more work for the reviewer to dismantle.

Strategies in this phase could include tapering the universal token budget so consumption stops being a virtue signal, introducing high-taste automated linters or secondary AI review agents as a slop filter before human review, and starting to grade employees on whether they’re building systems that level up their peers.

The Goodhart trap to note in Phase 2 is that tokens-per-employee falls, and management declares victory on “efficiency” while the actual problem (judgment, taste, context) goes unmeasured.

You’d want to exit Phase 2 when prompt quality, not adoption, is the bottleneck – and when a handful of Turbo Brains have emerged who can demonstrably ship reusable workflows that other people use.

Once the behavioral shift has stuck, orgs can transition off universal unlimited budgets and onto outcome grants. If a Turbo Brain can prove a specific workflow saves their team 10 hours/week (inclusive of vertical and horizontal stakeholder alignment), their department unlocks a dedicated uncapped pool for that workflow.

The unit of reward is the reusable scaffold: a prompt library, an internal tool, a documented workflow, an agent that other people can pick up and run. You’re rewarding the thing that compounds, not the person who happened to build it. This matters because the alternative – rewarding individual productivity gains – recreates the same Slop Cannon dynamics one level up.

Outcome grants should be jointly approved by the function head (who can attest to the work) and finance (who owns the budget). Neither should own it alone since function heads may rubber-stamp and finance may starve good ideas.

What about employees who don’t graduate into Turbo Brains? They keep a reasonable baseline allotment for daily work, with their manager’s preceptorship still available.

We could counter the Goodhart trap in Phase 3 by funding attempts at scaffolding, not just successes, to avoid any drifts back toward conservative use of AI.

To get to 3x productivity, organizations will want to progress to phase 3 as quickly as possible without sacrificing the foundational integrity established in the previous phases. You’ll want to watch out for signals of missing foundational pieces that could cause an organization to standstill.

In Phase 1, focus on quality once adoption numbers take hold. If not, token usage could climb, senior ICs could become full-time slop janitors, and no one will pull the trigger on preceptorship to progress on productivity.

In Phase 2, watch out for preceptorship becoming a permanent state where no one is trusted to operate independently, every workflow still requires a senior reviewer, and Turbo Brains never get the room to build. You’ll want to look at scaling quality and enable to deploy org-level judgment once it’s developed.

Be cautious of jumping straight into Phase 3. By skipping the foundational phases and gating everything on outcomes from day one, you risk adoption for the majority. This is where eco-conscious budgeters may self-cap before they know what’s possible, where Turbo Brains work in isolation without bringing the rest of the organization with them, and where managers may not realize the competence wasn’t broadly built because exploration was never funded.

When Claude 5 (or the next leap) lands, it’s expected that parts of the org will snap back to Phase 1. Old prompts will break, new capabilities will emerge, and exploration will start over.

What’s critical for organizations is to cultivate problem-solving skills by developing a culture of play. AI has democratized the ability to build more than ever before, and the best way to understand it is to make stuff with it. This means budgeting for temporary efficiency dips as the expected cost of learning by making. The organizations that understand this don’t treat exploration as a waste. Instead, they deliberately fund exploration, protect it from short-term efficiency pressures, and build it into how teams work.

Orgs that do this don’t scramble when the next evolution of tools and capabilities land. They’ll treat it as Phase 1. That capacity to reset, explore, and build will be one of the only remaining durable advantages in a market where every competitor has access to the latest Anthropic releases.

Read the original on kacinguyen.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.