It must have been in September 2024 that I wrote the last section of the book AI for beginners 2.0, for it was shortly after the first reasoning model was published (o1, by OpenAI). Along with its predecessor it has sold over 11,000 copies, and is also translated to Danish and Norwegian – though not (yet) to English.
I wanted to reflect a bit on what has happened in the AI world since then.
In summer 2024 there was talk about AI hitting a plateau. Multiple AI labs reported not getting as much return than they had hoped for when training bigger models, which fueled the view that AI is slowing down.
However, the so-called reasoning models started a new era of AI advancements. They are LLMs that are trained to produce intermediate reasoning steps before answering questions or attempting to solve the problem at hand. This allowed AI to make more elaborate plans, discover flaws in its reasoning, and generally behave much more robustly on large or complex tasks. The result was a new phase in AI advancements, and also AI that was much better at making (and sticking to) plans.
The S curve of the non-reasoning LLMs might have plateaued, but there was a new S curve making the technology steam ahead faster than ever.
Not even the companies creating new AI models know what they will be able to do, and it is difficult to put “general” AI capabilities on a scale to see how top performance improve over time. However, an organization called METR has created the arguably best study on this.
METR uses a set of real-world software engineering tasks to measure AI capability. The tasks have been benchmarked by actual engineers, or sometimes just estimated by them, and ranked by how long a human professional would need to complete them – from a few seconds to many hours. The AIs are then tested against these tasks, to see how long tasks they manage to complete. Since “100 percent success” create brittle results, they report how long tasks the AI can complete 50 percent of the time (with an alternative metric for 80 percent pass).
The result showed that the length of tasks that AIs can complete are doubling every seven months. However, when looking at only reasoning models, the doubling time is instead four months.
When this is being written (early March 2026), the top performing model can complete half of the software engineering tasks that take professionals 12 hours.
In 2025, we have seen AI systems produce research hypotheses matching those of entire teams of researchers, and solving mathematical problems listed as unsolved for decades. In early 2026, we also saw AI producing novel results in theoretical physics.
Understanding the pace of AI progress is the single most important part of understanding how AI is impacting the world. We humans are notoriously bad at grasping exponential growth. For linear growth, the next step is insignificant compared to all the previous ones. For exponential growth, the next step makes all the previous ones insignificant.
The exponential increase of AI progress is staggering. I have followed the updates on the METR curve, and have frequently thought that the progress cannot possibly continue in the same way. Yet it has. Maybe it will for yet some time.
In the wake of reasoning models, there were AI agents.
We had AI agents before the reasoning models, too, but they easily lost track of what they were doing and were generally unreliable. With increasingly powerful and reliable reasoning models, AI agents picked up pace.
There is no commonly accepted definition of an AI agent, but I prefer to think of them as AI systems that can:
make a plan for achieving a goal,
use digital tools (such as web services or local applications), and
track how they’re progressing through their plan.
One particular type of tool that is useful for AI agents is the ability to write and execute computer code, since that is a tool that allows for performing a lot of other digital tasks.
An AI agent that deserves a special mentioning is Claude Code.1 It is a program you install on your computer and run through a terminal window. You can talk to it like a chatbot, but it is capable of running commands in the terminal, and it has scaffolds that lets it interpret outputs from commands it makes, make plans, retry with new approaches when necessary, and much more.
Claude Code can do many different things, but was originally built for coding. During new years holidays 2025/2026, Claude Code in combination with the powerful new model Claude Opus 4.5 sent shock waves over the software engineering world. Now even highly experienced software developers report having changed their habits completely, describing to Claude Code what they want, wait for a while (or start other tasks in parallel), and review key parts of the produced code.
Early 2026 saw the rise of what was eventually called OpenClaw. It is, briefly, an always-on AI agent that you can talk to over services like Telegram and WhatsApp.2 Since it is always on, it can also contact you when it is called for, work while you sleep, and monitor the digital world continuously. It can also follow general long-term goals you give it (“keep improving yourself and surprise me every morning”), and use it to assign itself tasks it sees fit (“get a phone number, make use of voice AI services, and call you in the morning”).
We are still just starting to learn what agents can do, how we should use them, and what they can do when working together.
While software engineering is being revolutionised (or rocked by a “magnitude 9 earthquake” as AI developer Andrej Karpathy puts it), other jobs are still not as affected. There are signs of decreasing numbers of entry-level jobs, but it is possible that this is an effect of post-pandemic normalisations. There are also studies showing that previously strong correlations between productivity and available jobs are being decoupled, which by some are taken as a sign of AI starting to affect the labour market.
Since ChatGPT was launched, there have been talks about an AI bubble that is about to burst. The extreme investments in AI, coupled with difficulties to predict where things are going, have caused some fluctuations on the stock market. In 2026, the stock market has so far swung in the other direction: new and more capable AI models, combined with tools aimed at specific sectors, has caused values for companies within software (particularly software-as-a-service), finance and law to go down. It could be a sign that markets are starting to take AI as knowledge worker seriously, but it could also be something else.
The geopolitics of AI is very similar to how it was in September 2024: The US is in the lead, and strives hard to keep China on the second place.
Something that surprised many was when Donald Trump lifted the export ban on advanced chips to China, and allowed exports of the second best generation of chips (against the advise of every expert I’ve heard). The access to advanced computer chips is still the dominant bottleneck for AI, but energy supply is starting to constrain the AI companies – there are already data centers that are partially unused due to lack of power. China has a much better situation when it comes to power supply, but has been severely limited by access to chips. This is about to change, as the lifted export bans take effect.
Talking about China, DeepSeek R1 should probably be mentioned. It was a reasoning model launched by a Chinese company with the same name in January 2025 – only months after the first reasoning model was released, and basically as powerful. Notably, it was created with a much smaller budget than any of the big US companies, which caused a lot of talk and ups-and-downs on the stock market.
Looking closer at the general trends in AI, DeepSeek R1 was more or less right on track – which is a sign that China is not far behind the US when it comes to artificial intelligence.
One final section in the geopolitical chapter must be dedicated to the conflict between Anthropic and the US Department of War, that is very much ongoing right now.
Anthropic is along with Google and OpenAI the three leading AI companies in the world. (xAI, now merged into SpaceX, are not quite there.) Anthropic has/had the only AI models allowed in classified networks in the Department of War.3 In late February, the DoW demanded that their terms of use should change, allowing them to use the models for “all legal purposes”. Anthropic, a company founded on the idea of creating safe AI, insisted on keeping clauses that prohibits domestic surveillance and fully autonomous weapons, which the DoW very much did not like. (Maybe there were some personal clashes involved as well.)
It resulted in the DoW classifying Anthropic as a supply chain risk. This doesn’t only mean that they can’t do business with the DoW, but no supplier for the DoW can use Anthropic products in the work they do for the DoW. This kind of measure has for example been used against Huawei, but never before against a US company – let alone one that is allowed in classified networks in the Pentagon. The only reasonable interpretation is that it is a retaliation, not that the DoW actually sees Anthropic as risk. (This is underlined by their AI models being used in the attack on Iran the very next day. For one thing.)
Where this leads is very difficult to say, but it is clear that the US state is actively causing severe economic harm to one of the leading AI companies. It is also clear that all other AI companies are expected to fall in line. When the DoW say “jump”, the expected reply isn’t “only under these circumstances”.
In terms of AI risks and AI safety, there have been a few updates.
There have been a number of suicides and attempted suicides that (allegedly) were triggered or assisted by either AI companions or more common chatbots. Some of these concerned children under 18.
AI has been proved to be a powerful tool in cyber attacks and cyber defense. There are real-world examples of malicious actors using AI systems to automate or vastly accelerate advanced cyber attacks, as well as AI identifying previously unknown software vulnerabilities in order to patch them.
The Trump administration has made efforts to restrict and even prohibit individual states from passing laws that would slow down AI development. So far they have not succeeded, but it is clear that the administration’s stance towards AI is “accelerate” rather than “safety first”.
There have been a number of new studies showing that AI models are capable, and sometimes prone to cheat, deceive or blackmail to avoid being shut off, replaced, or in other ways prevented from reaching goals they deem important.
The phenomenon of “evaluation awareness” (also called “situational awareness”) has become common. It means that AI models often can guess or deduce when they are being evaluated, and adapt their behaviour to what they think is expected.
AI is increasingly being used in the design of the next AI model, and to design better scaffolding to make AI more capable. As is expected with exponential growth, acceleration itself is accelerating.
In summary, the last 18 months have been more or less on track with the exponential growth we’ve seen for a number of years. AI advancements continues to be both staggering and expected.
The big AI companies report that they “see no wall” – meaning that current methods of improving AI most likely hold for at least six more months. If they hold for 12 or 18 months, we should expect truly revolutionising and disruptive capabilities, not only in coding.
It is worrying that general awareness of AI advancements remains quite low. I hope that AI for beginners helps changing this, if only a tiny bit.
Claude Code is not the first code-writing AI agent, but it is the tool that many seem to have converged on, and the tool that made a new way of coding take off. For alternatives, check for example out OpenAI Codex and Google CLI.
Again, OpenClaw is not the first always-on AI agent, but it is the one that really took off.
Previously named the Department of Defense, but this was changed by Trump. Go figure.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.