RSS Amplifier

Ann Jackson · Jun 19, 2026

Self-Service Analytics, According to Anthropic

0
Sign in to vote or save

Ann Jackson · Ann Jackson

At the beginning of June 2026, Anthropic released a blog post describing how it enables self-service data analytics using Claude, reporting 95–99% accuracy. To say I was curious is an understatement. If you have the time, I encourage you to read it here.

The headline finding: out of the box, Claude never achieved more than 21% accuracy in answering analytical questions, even after verifying Claude had read all the materials and content available. A finding made even more devastating by the follow-up that providing Claude with all pre-existing analytical content, all the SQL queries, notebooks, and dashboards didn’t move the dial on accuracy a single percentage point.

How is this possible?

If you’re an AI doomer it is easy to gloss over the whole experiment and blame it on AI, but that would also be the easy way out. The truth, however, is more subtle and more obvious. It was never about the technology, about the SQL, or the data warehouse. As Anthropic puts it, “Analytics accuracy is a context and verification problem, not a code generation issue.” It was always about the skills.

Self-service analytics has been a holy grail aspiration my entire career, mainstreaming some time around the mid-to-late 2000s (depending on an organization’s tolerance for what qualifies as self-service analytics), and is still in that same aspirational state. Let’s call it a square two decades for ease.

In that 20-year period, we’ve seen an evolution of Business Intelligence (BI) tools, from initial warehousing and query-related efforts, exportable reporting data dumps, the rise of the dashboard (and elevation out of Excel), expansion of analytical applications (often built from dashboarding tools), and now the inclusion of conversational analytics. We’ve also seen a word soup of terms applied to the practice: data analysis, business intelligence, analytics engineering, and sometimes even data science.

What we haven’t seen is mastery of the actual aspiration: enabling everyday workers to ask their own analytical questions and reliably trust the answers. Nobody has reaped the full benefits of the self-service analytics dream; everyone is still feverishly chasing it.

It’s easy to get bent out of shape at someone saying your 20-year goal hasn’t been achieved, but that kind of attitude would negate the massive benefits and improvements that have happened in those two decades. My own experience as a data professional is proof. I went from an analyst building pie charts using hand-keyed data in Excel circa 2010, to regularly connecting, modeling, and building interactive data displays from data warehouses in 2016, to consulting on the whole damn thing in 2018. A massive shift in my own workflows and a massive value-add to those I served.

As you track my own career progression it would be easy to tell a story that my skills grew in terms of working with technology, that I got better at coding, but that’s not the truth. I’ve never learned SQL enough to write complex queries from scratch. I failed a HackerRank test and still can’t reliably remember what the acronym CTE stands for. This shortcoming of mine is not because I think knowing SQL is frivolous or lacking in value. It’s because it never prevented me from being extremely effective at answering analytical questions people asked of me.

Self-service analytics has never happened because there’s been limited success in putting effort into the actual needs required to get it right. Yes, the tools are important, but the knowledge and understanding part are even more critical. And more pointedly, the people whose responsibility it was to usher in self-service analytics got caught up in the technology, kept renaming the function, and forgot about the outcomes.

The cure for 21% accuracy is plainly stated by Anthropic, they had to define skills for Claude to improve accuracy. Skills that when mapped to a human role align more toward a senior embedded data analyst than a senior analytics engineer.

Yes, they said the thing we’ve all known for a long time. More data, better code, perfect data warehouses do not equate to self-service nirvana, nor do they get the business closer to that effort. There’s a 21% accuracy ceiling.

With the rise of LLMs we’re reaching yet another iteration of what cutting edge technology for self-service analytics looks like inside organizations and it looks like having a conversation. From an outsider’s perspective it can be seen as the ultimate frontier. Ask a question and get an answer. Nothing could be more simple.

Except now we all know it was never that simple.

Here I point back to Anthropic’s research and their own description of what work is necessary to get to 95–99% self-service accuracy. When you read it, you’ll see it outlined as it would apply to their own technology, but it’s easy enough to abstract Claude away and to focus in instead on the art of what Stephen Few would likely call data-sensemaking itself.

In their research they break down four different conceptual layers: data foundations, sources of truth, skills, and validation. Layers that when stacked on top of each other produce the agentic analytics stack. Layers that when stacked together also describe how the task of answering questions grounded in data has always been.

credit to Anthropic, taken from their article

Data foundations, as described by Anthropic, while not solved everywhere, are largely legible. Serious analytics functions know this work exists. They know how to fund it, staff it, govern it, and complain about it. It is both table stakes and the proverbial iceberg under the water.

So too is the semantic layer, which is currently undergoing a glow up: being lifted out of other tools, given its own proper place, with mapping lineage and transformation along for the ride. But those remaining two living inside Anthropic’s sources of truth — query corpus and business context — those are far less native to the function.

adapted from Anthropic

A query corpus (not a SQL query, a domain-specific reference document) and business context (the lingo, concepts, and goings-on of the business), simply do not exist inside the functions we know today. These are not things that teams are spending their time on, spending their time discovering, spending their time documenting, spending their time codifying. These are skills that a senior analyst in a business unit would have. Skills which are described beautifully in Anthropic’s top two layers. Anthropic’s “skills” layer is, more plainly, the act of being a data analyst. The validation layer can be understood as error checking and answer acceptance.

adapted from Anthropic

Anthropic says it best by defining the difference between sources of truth and skills as the difference between the declarative what vs. the procedural how. As applied to a data analyst, it is the difference between knowing whatthe metric profit ratio is comprised of (profit/sales) vs. how to derive that number and from where.

Within its fully agentic approach, Anthropic strongly considers the need for adversarial agents questioning the result before task completion. A human skill that we know as skepticism. A trait when applied in an agentic workflow improves accuracy by 6% at a 32% increase in token expenditure and 72% increase in latency cost. Oof.

Someone please point me to where these formally exist within the function. Certainly there are those that have these capabilities, but they’re not formalized, hard to measure, and as witnessed in the case of skepticism, costly to implement. Nevertheless, they are of critical importance and are the mechanism by which self-service produces accuracy. An entire area we need to stop hand-waving over in tech demos.

The top layer serves as the feedback layer to everything constructed underneath it. Seeing them listed out like this reminds me of the fog of understanding that every analyst must contend with. All of which would be trivial if only they received precise questions, were proactively kept up to date, and provided the information they needed to go look something up. Something that is of the rarest occurrence in most enterprises.

With Anthropic meticulously describing all of these elements, they’ve inadvertently explained why all the work done to date hasn’t made a dent in what was promised. Through fresh eyes, they’ve written down exactly what’s necessary for self-service analytics.

They’ve also explained the other 79% that is so constantly waved away: the analyst’s judgment.

Distilled versions of the queries and questions, an arsenal of analysis techniques, the ability to understand the “why” behind the question. Those are human skills that suffer from lack of definition, lack of place, and lack of recognition in today’s self-service analytics world. A unique blend that when put together resembles a crime scene investigator more than a programmer. A person skilled enough to consider that there may be a partial fingerprint on a doorknob that can be analyzed vs. writing “if this then that” instructions to computers.

Through the lens of their solution these skills need to be written down, a photocopy version of the analyst presented as easier to read code.

For Anthropic, a five-year-old AI company with a deeply technical founder-bench, a series of markdown files and human managers approving inclusions and monitoring outputs daily can make a lot of sense. It sounds plausible. But what about the everyday enterprise? The 100-year-old health insurance company who acquired new lines of business with unique regulatory requirements. Can they too resolve all of their domain-specific knowledge to a file and maintain it in perpetuity? Daily? Is that even feasible?

No, I don’t think it is. And as the analyst in this conversation, it is time to take the attention and budget away from the tools and put it toward human judgment. The missing 79% is human judgment, meticulously verbalized by Anthropic. It is the thing we have no business pretending we can pass off to technology before understanding, funding, and formalizing it.

Author’s Note: Between March and April of 2023, I wrote a 7-post series on being a data analytics consultant (first post here). After writing this piece, I think the time has come for me to rework my 2023 material to now include how generative AI fits in — and it fits in massively.

No posts

Read the original on annujackson.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.