RSS Amplifier

AI Realist · Aug 20, 2026

How chatGPT-Taught Experts Are Crippling Agentic AI

0
Sign in to vote or save

Maria Sukhareva · AI Realist

For decades AI was an obscure field covering Natural Language Processing, Computer Vision, Bioinformatics and similar areas. People considered it a difficult research area with unclear value. When deep learning gained momentum, the field got even more complex - research papers filled with mathematical formulas, code that no one knows how to run, and constant CUDA errors when you try to train something.

The output, though, was fairly understandable - we can classify reviews on Amazon for your product into good or bad, divide your documents into invoices and contracts, find all the addresses and people’s names, translate documents. Those were obvious repetitive tasks, not hard to understand in terms of what they could do and very hard to understand in terms of how. But the capabilities were clear and they fitted neatly into existing workflows and established processes. And when the what is that clear, working out the how is a technicality.

And then something interesting happened - ChatGPT arrived.

I thought this made things worse - now it was hard to understand what it can do and hard to understand how.

Not everyone shared my opinion. A colleague, in a heated argument with me about the need for deep learning experts as such, said something like: “AI is very simple now, anyone can do it, you do not need AI experts anymore, anyone can do AI.”

What made him think this was the gap between his own lack of expertise in AI and the convincing eloquence of GPT-3.5 - the first model of its kind to reach a large audience, and the one that spread hallucination and sycophancy along with it. Very quickly that model convinced him that he had a deep and profound understanding of how AI works and what it can do. And so it happened with many others. I call them chatGPT-taught experts.

I wish I had known the word slop back then, because I was lacking a concept to describe those presentations - gorgeous in aesthetics, horrendous in content - produced by fairly expensive and well known consulting agencies that target C-level managers and decision makers.

Surprisingly to me, managers would far rather watch this slop for hours, sit through executive briefings and AI strategy workshops, again filled with slop presentations, than spend one day trying AI output themselves.

And that is how we ended up with atrocities like YourCompanyAI - a useless in-house model, built in the spirit of BloombergGPT, which was purpose-built from scratch on forty years of financial data and then beaten by general-purpose GPT-4 on almost every financial task within weeks of its release. Or a RAG bot that desperately hallucinates about the documents it found on the company intranet. Or a customer support assistant that customers avoid like the plague.

Most of the workforce was experiencing a kind of cognitive dissonance. On one side, the news and the big minds of AI were talking about the impending white collar bloodbath, and companies were boasting about AI related layoffs. On the other, all the employees actually had was Microsoft Copilot and a multitude of DIY chatbots that were of absolutely no use to them.

So how did it happen that a technology that has become genuinely good and advanced over the last four years has had so little impact on the workforce, and even worse - many of those whose productivity it was supposed to boost do not want to use it, and consider it dumb, hallucinating and dangerous?

That is where I come to the key argument of this article. You cannot understand what large language models can do, particularly in combination with their agentic functions, unless you try them. And, thus, you cannot take an informed decision about anything related to agentic AI.

Become a paid subscriber to AI Realist and unlock:

https://msukhareva.substack.com/subscribe

Historically, career paths, particularly in traditional companies, never favoured individual contributors. The ultimate growth was through becoming a manager - first a manager of 5 or 6 people, then the head of a bigger department, and finally the ultimate manager, the one with a letter C in the title.

A manager is a kind of general. He does not go to war in his armour, he defines strategy, he decides what to do and how to do it is less of his problem. But there is a difference between a general and a typical manager. The general is an expert. Once upon a time he was a soldier, he studied the art of war, and he knows what he is dealing with.

Managers do not know what they are dealing with when they decide what to do with agentic AI, or with LLMs in general. The majority of them never worked with the models, never studied them. Frequently, all their knowledge comes from the shiny slides of consulting companies, created by people who have also never worked with AI or agents, and who are merely chatGPT-taught experts.

Let us look at the following illustration:

What is this nonsense, you might say. Aren’t all cars on wheels? What is the point of a car if it does not move? What is emerging there? And every car has an engine, or a motor if it is electric, even a self-driving one. A self-driving car can be electric or not, its ability to be autonomous is not related to the powertrain at all. And what does GPS have to do with it? Cars had satellite navigation for decades before any of them could drive themselves. What exactly does this assessment framework help me to assess?

It tells nothing about the capabilities of real cars, or their cost. So is a Ferrari the same as an Opel? I mean, they are both engine-based. Or are they both moving cars? After all, they move too.

And you know why you think this is nonsense? Simply because you have seen cars and driven them. You know there are different brands, different ways to fuel them, different maximum speeds. You know that a petrol car has an engine and an electric one has a motor, and that this is not a trivial detail. And you know that a car is, at minimum, a thing on wheels.

And now imagine a world where cars never existed and suddenly appeared long after PowerPoint was invented. A group of very serious decision makers arrive at work on their horses, and a respected consultant, who also arrived by horse, tells them about this amazing automotive mobility, a term they are hearing for the first time. There is a so-called “car”. It drives faster than a horse. Unlike legs it has wheels, and it rolls by itself. It has powerful engines equal to hundreds of horses. It can even drive itself. It can tell you where to go and save your route with GPS.

You still do not believe that something this nonsensical would convince anyone?

Well, lo and behold:

This is one of the first slides I found when searching for “agentic AI gartner”, from Jim Hare’s September 2025 article Agentic AI for Vendors Is a Risk Without Oversight.

Let us imagine I am someone who never worked with an LLM. What do I learn from this? The most basic level of AI assistants is a conventional chatbot. Conventional as opposed to what? Is there an unconventional chatbot somewhere, one that rejects the norms of society and calls for anarchy? One rung up sits a conversational AI assistant, which raises the question of whether there are assistants that refuse to converse. And do not be confused: your conversational AI assistant, although it counts as agentic AI, is not an agent. Because to be an agent you need to be… LLM-based.

That is where I immediately want to ask: is this diagram also diachronic? There used to be, in 2018, task-oriented chatbots occasionally called assistants that never used an LLM, but as soon as ChatGPT appeared all the new chatbots became LLM-based.

I had to research a bit what this assessment framework is supposed to be assessing, and to me, it looks like it assesses the year a chatbot was built. Look up what "conventional" means here and you find decision trees and keyword matching, with NLU and intent detection one level up and an LLM one level above that. So, yes, if you built a conventional chatbot - congratulations, you are in year 2011.

The only assessment I can make from this slide is that it was created by someone whose entire exposure to LLMs is a corporate ChatGPT subscription.

But now imagine someone who has no idea about chatbots, assistants, AI, LLMs and so on. Someone who has never used any of them. Would this framework make sense?

Of course it would. Just google it and see how many people have unironically discussed conventional chatbots and LLM-based agents as if this classification were of any use.

And now let us get back to our example with the cars. The managers got very impressed by mobile automobiles, particularly by the self-driving ones, and started envisioning how they would fire all the coachmen, how they would cut the costs of horse maintenance, food and grooming, and how they would be able to go from Paris to Berlin in an hour while sleeping in the car as it drove itself at 500 kilometres an hour.

The sky is the limit here when noone in the room knows what a car can realistically do.

They bought some cars, put the guys that ride the horses in them, and somehow the cars were even worse than the horses and everyone hated it. One guy even got killed because he confused gas and brake. Another one thought that he did not need to steer and flew off a cliff. After all, a horse is smart enough not to jump off a cliff, and this stupid technology is dead dangerous.

They decided that from now on a car may not go faster than 10 kilometres an hour, and that the only place where one is allowed to drive it is in a circle in a designated space. Suddenly the car could not go anywhere and was barely moving, and the decision was that it was a useless and expensive technology.

If this sounds too stupid to have actually happened, it happened. In 1865 the British Parliament passed the Locomotive Act, remembered as the Red Flag Act. Road vehicles were limited to 4 miles (6.4 km) per hour in the country and 2 miles per hour in town. Every vehicle needed a crew of three, and one of those three had to walk sixty yards ahead (human in the loop so to say) of it carrying a red flag to warn the horses. The rules stood for three decades, until 1896. And they were not written by frightened bureaucrats alone: the stagecoach and railway industries lobbied for them, because they could see what the car would do to their business.

The fatality happened too. On 17 August 1896, Bridget Driscoll was struck and killed at the Crystal Palace by a car giving promotional demonstration rides, travelling at a speed the driver insisted was capped at 4 miles per hour.

They were doing everything possible, just not learning to drive.

Well, instead of relying on slides and deciding that moving cars are too dangerous, one should have just tried to drive one.

The person would quickly see for themselves that a car is not a horse. It needs good roads. It cannot turn into a narrow path in the woods, jump over a pond, or drive through deep mud. They would realise that the car will break down, get stuck, or simply not fit, unless the road meets certain criteria. They would realise that it essentially does not matter whether you have an Opel or a Ferrari: if you have no roads, you have a problem.

They would also realise that unlike a horse, a car cannot take some rest, chew on grass and keep going. Cars need fuel and places to get that fuel. Fuel costs money. And better, faster cars burn much more of it than cheap ones.

They would know that cars are not self-driving, and that even the ones that are need certain conditions in order to do it. They would know that you need to steer, because cars do not normally hold the lane on their own. They would realise that you do need rules and safety precautions, but that those rules should enable what the car is actually for: going fast, holding the lane by itself, covering distance. Crippling a car so that it can only do 6 kilometres an hour might well be safe, but it makes the car completely useless.

And after all that, they would know that the only way to learn how to drive is to practise.

The Gartner assessment framework, in my humble opinion, is utterly useless. It does not tell us anything about how to evaluate your agents.

It is very easy to understand why, if you spend one day vibe coding.

Many consultants preach safety and guardrails. They preach it to such an extent that there are no roads open for LLMs at all, and they are just doing pointless circles in the sandboxes. This is one of the most important points here. If you say to an LLM, summarise the latest news on AI, and it does not have access to the internet because you blocked it, it is useless. If you ask an LLM to edit your file but it does not have read and write permissions, it is useless. LLMs might be the engines, but if you do not give them wheels, fuel and good roads that lead to many places, they are utterly and completely useless.

The assessment that would really show how developed your agents are would not care about things like “learning”. Currently there is no true “learning” out there, in the sense that an LLM can memorise new information as it works. The weights are frozen at inference time. This is a limitation of the technology. The only learning LLMs can do is:

  • Write text files where they store a log-like record of their recent actions and findings

  • Self-correct by validating against provided criteria, which lets an LLM adjust its own prompt and the scripts it wrote, and add new rules and assets

None of this makes your agents next-level advanced, simply because all the modern harnesses do it already. They all write some kind of memory.md, user.md and a bunch of other markdown files. You do not need to implement anything here most of the time. As for validation and self-correction, this is even more trivial: all the state-of-the-art LLMs are trained on agentic loops, so they will define a plan and acceptance criteria and check their work against them.

Full automation is not more advanced than learning either. For example, I have fully automated agents, and if you spent a day building agents you would also know that full automation is not that hard to achieve. My automation agent is configured in OpenClaw. Each time I get a new paid subscriber here, this agent adds the subscriber to the airealist.org workspace and sends them an invite to see the courses there. If someone purchases a course on the website, my automated agent processes their payment, adds them to the list of course participants, sends an instruction email and reduces the number of available spots.

Is that advanced? No. In fact the only part the LLM played here was building the agent. The agent itself is a set of deterministic Python scripts. And this is one of the best use cases for an LLM: try not to build agents on the basis of LLMs, but rather build agents with LLMs, by asking them to create scripts and skills for themselves.

And none of the above is a secret knowledge - you get to realise it very fast .

That is where managers and agentic AI clash. A manager is trained to be above implementation, above hands-on. Managers learn from high-level slides, they do not code. And people who talk to managers also do not code.

In the past two decades we would constantly come up with some role that would focus on translating between individual contributors and managers: scrum masters, project managers. In practice their only job was to translate from the language of doers, developers for example, into the language of managers. I myself, when I was at an earlier stage of my career and thought that being a manager was the natural progression, would frequently hear that I was too technical to become a manager. And now we suddenly end up in a situation where there are barely any capable translators. The translators themselves do not understand what is going on, and they repeat the slop one after another.

Agentic AI is fundamentally different from what one was dealing with before.

First of all, there is the speed and the scale at which it produces solutions, though not necessarily the quality. In reality LLMs are not great for repetitive tasks at high volume. They reason for a long time, they burn a lot of tokens, and they are non-deterministic, so every outcome can be different. Send 100 identical requests to a GPT model with the seed parameter set and you can get 24 to 33 different answers back, because the result depends on how busy the server was. What they are great for is writing the code and setting up the pipelines that will run those repetitive tasks.

It also does not mean that the quality of the output of an LLM is better than that of other methods. Yes, it might be better at paraphrasing or summarisation, but it might be worse at something as trivial as named entity recognition. What LLMs do is scale the error as well as the success.

And the scale is beyond anything we have seen before. One needs to experience it to see how fast it can ship a website or a dashboard. And like a car, it cannot move without a road. You very quickly see the limits of an agent mid-implementation, when a certain API is not available, when Google Drive is blocked, when there is no connection to the internal data, and eventually you start feeling like you own a Ferrari that is stuck in the garage.

The consulting companies like talking about guardrails and human in the loop. And here again is where they completely misunderstand the scale and the speed of agentic AI. Guardrails frequently end up blocking certain APIs, disallowing certain models, or gatekeeping access to AI tools from the employees they think might not be able to handle them. The human in the loop is presented as some kind of ultimate guardrail. And in fact it is not.

At best, that poor human acts as a liability sponge but barely adds more safety.

Human in the loop is important, but no human can keep up with the volume and the speed that LLMs produce. Furthermore, more often than not the human simply does not have the expertise to evaluate the output. Let us imagine a financial analyst wants to build an interactive dashboard. He asks Claude Code to code it, and the agent asks him every other moment to validate the code. He has no idea, so he accepts everything. In half an hour a beautiful dashboard is ready, and the financial analyst is happy to use it and even share it with colleagues.

Not rarely, the financial analyst will be prohibited from even building the dashboard, and from using Claude Code at all, because he cannot validate the code. If he does manage to build it, even more frequently he will be required to fill out a number of forms, to ask the development team to review it, perhaps to run some penetration tests and whatnot. And thus something he built in half an hour turns into half a year of bureaucracy.

This approach is essentially like making him drive a car no faster than 6.4 kilometres an hour, in a circle. And the human in the loop is the man walking 55 metres ahead of it, waving a red flag.

The solution could actually be very easy here, but I pretty much never see it in any recommendations. Instead of the omnipresent loop with a human, use LLMs to build deterministic pipelines for validating the solutions. Models have become genuinely good at this: GLM-5.3 currently leads the CyberGym vulnerability-discovery benchmark, on Zhipu's numbers, and models in general are strong at reviewing code against explicit criteria. The Agent Skills standard lets you create reusable skills that set up templates for project provisioning, so that every dashboard built by a non-developer from now on follows the same best practices your developers defined. Your cloud configurations, your app deployments, your code vulnerabilities: all of it can be monitored by an LLM orchestrator with a collection of scripts.

To match the scale and speed of AI, you need AI and not a human liability sponge.

Unfortunately, this is not something one can put on slides or explain abstractly. This is something one needs to live through.

Thus I would recommend it to every single manager who is in some way responsible for AI decisions: whether it is which subscriptions to buy and for whom, or how to define the strategy for the whole company. Take one day. Get yourself a ChatGPT Pro or a Claude subscription and try to build something you actually wanted to have, be it a dashboard, a website, an agent that processes your emails or manages your calendar. Try different models, do not obsess over prompt engineering, and see how far you get. And then take your company computer, try to reproduce it there, and see how far you get that time. That is where you will see whether you have the roads and the petrol stations, or whether you simply bought a Ferrari that cannot go anywhere.

As a starter you can use the AI Realist guide for building a website with Fable, though by now any frontier model, whether Opus 5, GPT-5.6 Sol or Kimi K3, will get you there.

And for those in Munich or elsewhere in Germany, AI Realist is organising an on-site Agentic AI Masterclass, exactly for those who need to make decisions about AI but have not had much hands-on experience. In this masterclass we will build a working prototype that you can take home. At the same time we are going to push agents as far as they will go: whatever you decide is needed, we will try to implement it there and then, and we will see whether there is already a road our car can drive on, or whether we need to build the road as well. We will also focus on minimising error at scale. You will learn not only to build with agents but to build agents, the ones you can take home too. And most importantly, you will see what an agent actually is.

Early bird tickets are at 100 Euro discount till the 10th of September

And I guarantee you that after this day, whenever you see a slide like the one from Gartner above, you will be shaking your head and saying: this slide was built by someone who has never worked with agentic AI.

Share

Read the original on msukhareva.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.