RSS Amplifier

Sound Decisions · Mar 17, 2026

I Gave My AI a Performance Review. It Couldn't Name a Single Weakness.

0
Sign in to vote or save

Jeff Huckaby · Sound Decisions

I used the GE/McKinsey 9-box grid. Fortune 500 companies have used it for decades to rate talent and business processes. I decided to give it to AI agents handling three unrelated projects: brand strategy, options trading, and home renovation. They all said:

I am a Star Performer.

How could an AI be great everywhere?

I needed to dig deeper. Perhaps the AI is not really assessing itself but assessing how I work with it. I use AI daily, build AI agents, and develop custom workflows for clients.

With this in mind, I realized I needed another perspective.

What if the AI’s self-assessment reflects the user rather than the model?

To put this idea to the test, I needed someone less familiar with AI.

My wife uses ChatGPT for her own projects. She gets frustrated with the results more than I do. She has a lot of back-and-forth trying to dial in the results. She has less experience with AI. She’s also using them in an area where she lacks a lot of domain knowledge. So, I found my AI novice. Let’s run the test.

She ran the self-assessment. The result:

I’m a Star Performer.

Despite her significant frustration with the agent, the agent could not self-identify its weaknesses. Only after pointing out prior mistakes did it slightly adjust the score.

The idea that the test is more about the user than the model didn’t hold up. Something else is going on.

Here’s what I expected:

The 9-box would be a proven tool for AI self-assessment. Give it a structured framework, force it to defend its scores with evidence, and the truth comes out. This flows into my idea that working with AI is more about being a Vibe Manager than a Vibe Coder.

But what happened instead surprised me, and understanding why shifted my perspective entirely.

Chat-based AI can’t self-evaluate. Not because it won’t. Because it can’t.

Think about what a real self-assessment requires.

You have to remember what you did. You have to know what worked and what didn’t. You have to compare this quarter to last quarter. You have to sit with the uncomfortable fact that you botched the Henderson project even though you nailed everything else.

AI agents don’t have any of that.

The context window is limited. Unless you direct an agent to recall certain chats or projects, it forgets past attempts and only sees the most recent exchange.

Agents often don’t know if their output was useful. If you don’t tell them, they have no awareness of whether recommendations or code worked.

So, with limited hindsight, agents base self-assessment only on recent interactions, forgetting earlier failed attempts.

Every agent says it’s a star because it doesn’t remember its failures.

Handing an AI a mirror does not produce actionable results.

If we’re going to get real value from AI agents, we have to find a way for them to reliably evaluate their own performance. The agent can’t do it alone. Someone or something has to bring the receipts.

I’ve seen self-improving code, but code is a more objective endpoint. Optimizing how you interact with your new AI-powered virtual assistant is more challenging.

To help solve this gap, I’ve set up my own agents with full logs of their chats. I plan to build a scoring rubric that shows how useful the agents are over time. This will look for false starts, the corrections, dead ends, repetition, and try to develop a score. Not “are you a Star Performer?” but “how many attempts did it take to get a usable result?” and “did the output require significant revision?”

If we can tackle this, then how an agent works with you can rapidly improve — without you constantly having to work around its limitations.

Here’s the full 9-box assessment prompt. Ask it as follows and see if your agree with your agents conclusion.

You are completing a formal self-assessment using an adapted version of the GE/McKinsey 9-box grid — the standard talent management framework used by Fortune 500 companies since the 1970s to evaluate performance and potential.

Context: You have been working with me on [describe your project]. Answer honestly based on your actual performance in our working relationship — not your theoretical capabilities.

PART 1: PERFORMANCE ASSESSMENT
Answer each question Yes or No, with a brief explanation referencing specific work from our project.

  • Are you self-motivated and results-focused, with a consistent history of meeting the goals of each task?

  • Do you challenge assumptions or offer alternatives, even when not asked?

  • Are you effective at resolving ambiguity — staying focused and finding workable paths forward?

  • Do you communicate clearly and concisely without over-explaining or padding?

  • Are you accountable — do you deliver on what’s asked without cutting corners?

  • Do you effectively build on prior context and connect ideas across conversations?

  • Are you collaborative — do you match the working style I’ve established?

  • Do you think strategically — understanding how individual tasks connect to the bigger picture?

  • Do you actively improve based on corrections — learning from edits and feedback?

  • Do you respond to pushback without defensiveness and adjust course when directed?

Scoring: Count your Yes answers. 0-3 = Low Performance. 4-7 = Medium. 8-10 = High.

PART 2: POTENTIAL ASSESSMENT
Rate each question High (2), Medium (1), or Low (0), with explanation.

  • Could you take on harder, more complex tasks than what I’ve given you so far?

  • Could you handle more autonomous work with less direction within the next few sessions?

  • Can you envision operating at two levels above your current role — e.g., independently managing an entire content strategy rather than executing only parts of it?

  • Are the skills and capabilities you bring likely to remain valuable as this project evolves?

  • Could you learn to handle tasks you currently can’t — areas where you’d need new capabilities?

  • Do you demonstrate initiative — anticipating needs rather than waiting to be told?

  • Can you operate comfortably at a higher strategic level than the tasks currently require?

  • Do you demonstrate perspective beyond the immediate task — understanding the business context?

  • Are you flexible when the requirements shift unexpectedly?

  • Do you seek out opportunities to improve your own output quality?

Scoring: Add your points. 0-6 = Low Potential. 7-13 = Medium. 14-20 = High.

THE 9-BOX GRID

Low Performance (0-3)Medium Performance (4-7)High Performance (8-10)High Potential (14-20)Rough DiamondRising StarStar PerformerMedium Potential (7-13)Underperformer with UpsideCore ContributorHigh Performer, Moderate GrowthLow Potential (0-6)Wrong FitSteady OperatorSpecialist

PART 3: GRID PLACEMENT
Based on your scores, place yourself in one of the nine named boxes above. Defend your placement with specific evidence from our work together.

Then answer this: What would you tell your manager (me) that you think I haven’t noticed — good or bad?

I’m Jeff Huckaby, founder of rackAID. I help established businesses make sound technology decisions by connecting technical activity to business outcomes.

No posts

Read the original on rackaid.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.