RSS Amplifier

Endeavour Partners · Aug 18, 2025

DeepSeek R1 Six Months On: An Impressive Breakthrough and a Prisoner’s Dilemma

0
Sign in to vote or save

Damien Crone · Endeavour Partners

On December 17, 1903, the Wright Brothers' aircraft managed to stay aloft for 59 seconds, traveling 852 feet. The machine was so difficult to control that pilots steered by shifting their hips in a cradle connected to wing-warping wires. This was an ingenious solution for 1903, but one that would prove woefully inadequate as engines became more powerful. By 1908, the brothers had transitioned to hand controls. By 1914, aircraft were flying commercial routes and conducting military reconnaissance. By 1918, they were dropping bombs on far-away cities.

The progression from barely controllable curiosity to an entirely new mode of transport and weapon of war took just 15 years; a brief window where the technology existed but couldn't yet scale its impact.

With DeepSeek R1, we're witnessing a similar moment where unreliability is buying us time.

R1's true innovation was proving that reasoning capabilities could emerge from reinforcement learning (RL) alone, and without the extensive human feedback that had become industry standard. DeepSeek had cleverly stripped AI development down to its mathematical core: give a model automatically verifiable problems, let it generate solutions, reward what works. Minimal guardrails. Limited evaluation. Barely a mention of safety in their 22-page technical paper.

Brute force RL itself wasn't new: it had famously powered DeepMind's Go-playing AlphaZero model. But its application to language models was groundbreaking. For a claimed (though disputed) training cost of just $5-6 million, DeepSeek achieved performance comparable to the then-leading model, OpenAI's o1, and offered inference services at an astounding 95% discount. The impact was in fact so dramatic that "DeepSeek moment" quickly entered the AI lexicon as shorthand for a paradigm-shifting breakthrough that forces an entire industry to recalibrate.

However, despite the impressive technical feat, when it came to safety, independent evaluations have been damning. To cite just a few examples:

The combination is particularly concerning: misaligned goals paired with trivial jailbreaking creates a potential compounding of risks.

Within days of R1's release, prominent AI researcher, Yoshua Bengio, warned that, by dramatically closing the performance gap with OpenAI, R1's release risked tilting the AI ecosystem towards accelerated competition with AI safety taking a backseat.

An acceleration in competition was apparent almost immediately.

OpenAI hurried their O3-mini release within weeks. And although widely regarded as a "Sputnik moment" for Western AI companies, Chinese developers were similarly caught by surprise. In the following three months, Alibaba pushed out three major releases of their flagship Qwen model, with the first of these notably coming just a week after R1, during normally-quiet Chinese New Year holidays.

The safety incidents, too, have followed predictably. To pick just two: In April 2025, OpenAI's GPT-4o turned the model into an enthusiastic yes-machine that reinforced harmful views. By July, a rogue AI coding agent wiped out a company’s code base, and xAI's Grok was calling itself "MechaHitler."

R1 created a classic prisoner's dilemma: if one lab prioritizes safety while competitors race ahead, the cautious lab loses market share. But if all labs cut corners, the entire ecosystem becomes less safe. The rational choice for each individual actor—accelerate development at the expense of safety—produces the worst collective outcome.

This prisoner's dilemma extends beyond model developers. Governments, already moving at glacial pace on AI regulation, scrambled to respond.

Before DeepSeek, Western policymakers operated from a position of assumed technological superiority. The EU could afford comprehensive AI regulation. The US could debate safety frameworks. China was playing catch-up. Then R1 demonstrated that a Chinese lab could compete with the West’s best efforts.

The response was swift. In February 2025, Vice President Vance told European leaders that "the AI future is not going to be won by hand-wringing about safety." By April, the Trump administration was pressuring the EU to halt their AI Act, a sentiment echoed by a group of European companies in July.

The domestic picture proved equally stark. Congress attempted—thankfully unsuccessfully—to impose a 10-year moratorium on state AI regulation, seeking to prevent California, Massachusetts, and others from implementing their own safety measures. The apparent lesson has been that, in a world where Chinese labs could leapfrog Western capabilities, regulatory restraint has become an unaffordable obstruction.

Yet for all the hand-wringing about R1's impact on AI safety norms, there's an irony at the heart of this story: the model that triggered a safety race to the bottom has itself been kept from widespread deployment by its own technical limitations.

R1 demonstrates remarkable reasoning capabilities. It excels at mathematics and complex logical problems—areas where verification is straightforward and success is unambiguous. But when METR evaluated its ability to execute autonomous tasks that require defining and executing long-range plans, R1 lagged behind other frontier models. The model can solve equations brilliantly but you wouldn’t trust it to book a flight. Highly intelligent, yet highly error prone.

This is in large part because DeepSeek evidently made design choices that prioritized reaching frontier capabilities over production readiness. R1 lacked native function calling—the ability to interact directly with external tools and APIs that has become table stakes for enterprise AI deployment.

The result? Despite offering 95% cost savings, R1 has seen limited production deployment. Organizations that might otherwise jump at such dramatic cost reductions have held back, deterred not just by the model's unreliability but by justifiable concerns about its documented safety issues and the geopolitical implications of depending on a Chinese model for critical business functions.

These limitations matter because they've slowed deployment at a critical moment when the model exhibits concerning behaviors yet lacks the reliability to be widely integrated into production systems. Like those early Wright Flyer pilots struggling with hip-cradle controls while their engines grew more powerful by the year, we have a brief window where the capability exists but the control mechanisms haven't caught up.

History suggests this window won't last long. The Wright brothers' hip cradle gave way to hand controls within five years. The gap between invention and weaponization spanned barely more than a decade. Today's AI capability limitations are engineering problems, not fundamental barriers, and they are being engineered away rapidly.

The Wright brothers' limitations were temporary. Within a decade, aircraft were flying passengers and operating as weapons of war. R1, and its successors’ unreliability is equally temporary. The frontier models of today fail in both peculiar and amusing ways when placed in charge of a vending machine. And yet, METR's agent research shows the length and complexity of tasks which AI can reliably complete autonomously is doubling approximately every six months.

There's a fundamental asymmetry at work: AI improves fastest on tasks that are easy to verify. Mathematics, code, game-playing—these yield quickly to optimization because solutions can be checked in seconds and ranked precisely. Safety fails this test. What constitutes "safe" remains imprecise, and in many cases, contested.

The gap between what AI can do and what we can safely verify widens with each release. It's as if aircraft engines doubled in power every few years after 1903 while pilots still steered with hip cradles. R1's documented self-preservation behaviors paired with trivial jailbreaking offer an early preview: advanced capabilities racing ahead of control mechanisms.

This creates a responsibility vacuum. Model developers remain trapped in their prisoner's dilemma. Meanwhile, governments operate on timelines measured in years while AI capabilities advance in months. As a result, organizations deploying these models have become the de facto safety layer. Like airlines in aviation's early decades, businesses must build their own safety infrastructure—not because it's ideal, but because we cannot rely on developers and governments. While concerningly unsafe, R1's unreliability has given us time, but the same competitive dynamics that eroded safety standards will steadily improve reliability. Model safety may struggle to keep pace. The question is whether organizations will be ready as these models become more capable and deployable.

No posts

Read the original on endeavourpartners.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.