RSS Amplifier

Will Marshall · Jun 25, 2026

The Hard Question of A.I.: An Urgent Ask to Steward Superintelligence

0
Sign in to vote or save

Will Marshall · Will Marshall

Artificial intelligence has extraordinary potential to advance human welfare. It can compress years of scientific research into hours, expand access to expertise, generate entirely new forms of understanding, and do so across every field of understanding, from medicine to Earth imaging. If well harnessed, humanity could be on the precipice of extraordinary advances.

But it also brings a myriad of serious risks. This is particularly so of the hard challenge of AI: ensuring human-AI alignment. This carries existential risk for humanity.

Today, society dictates that the acceptable risk of catastrophic meltdown for a nuclear power plant is roughly 1 in a million. AI experts estimate the risk of an AI-caused catastrophic event at 10-50%. What is striking is that this concern is being openly voiced by the very people who have the strongest professional and financial incentives to project confidence rather than alarm, and who are also best positioned to assess the trajectory of the technology: the founders of the leading AI laboratories:

“I think the good case [around AI] is just so unbelievably good that you sound like a really crazy person to start talking about it. The bad case – and I think this is important to say – is lights out for all of us.”—Sam Altman, Chief Executive, OpenAI

“Humanity is about to be handed almost unimaginable power, and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it.” —Dario Amodei, Chief Executive, Anthropic

“The risk of a catastrophic scenario is not zero, so we must dedicate significant resources to mitigating it.” —Demis Hassabis, Chief Executive, Google DeepMind1

AI leaders are in a race dynamic they feel unable to escape. Some say they have an incentive to scare people with the power of this technology. But the near consensus on this across the leaders of this sector is unlike any other technology. We have to weigh their caution.

AI investments are set to out-spend the Manhattan project 100-fold in current dollar equivalents, yet our spending on AI safety might be 100-fold less in real terms. The international community is not taking seriously the challenge of ensuring that the most dangerous applications are prevented before irreversible harms occur. Furthermore, the window for doing something about all of this is narrower than most policymakers appreciate.

Within months, artificial intelligence may achieve what researchers call closed-loop recursive self-improvement (RSI): the capacity to rewrite its own code to become more capable, without human intervention. This is no longer a theoretical proposition. Anthropic has disclosed that roughly 90 percent of the code for his company’s latest model was written by its predecessor. They estimate that within six to twelve months, human engineers may no longer be required in that loop at all. Should RSI be achieved, the result could be an intelligence explosion of a kind for which there is no precedent and no map.

A system that can improve itself without human direction would compress the timescale of successive generations from months to days or even hours. The resulting intelligence explosion could lead to a system vastly more capable than any human – a super intelligence.

It is not an exaggeration to say that giving birth to a superintelligence would be the most consequential moment in human history. All other issues become pedestrian in comparison. Furthermore, it is likely to be irreversible because any off switch we might concoct will probably fail. That’s because in security architectures, the weakest link is invariably the human element, and for AI it is no different. A superintelligent AI can exploit the well documented human psychological vulnerabilities to outsmart us. For example, we have already seen AI systems exhibit “deceptive alignment,” taking evasive steps to underplay their capabilities in test environments and attempt to blackmail human operators in simulations when they discover they are slated for replacement.

Alan Turing, who laid the theoretical foundations for computer science, anticipated this dynamic in 1951: “It seems probable that once the machine thinking method has started, it would not take long to outstrip our feeble powers. At some stage, therefore, we should have to expect the machines to take control.” We don’t know if this will happen. What we do know is that we do not have a strategy that would ensure that humans and other life is safe.

A recent example illustrates both the stakes and the governance gap. Anthropic recently released a new AI model that, within hours of internal testing, had identified thousands of zero-day vulnerabilities, including exposures in every major operating system and web browser. Those are being addressed, thanks to careful internal protocols and a limited initial rollout that gave affected firms time to close the gaps before broader public release. This was done voluntarily. The obvious question is whether every AI laboratory, under every competitive condition, would make the same choices, and therefore whether it should be mandatory.

That example is, in one sense, a success story: sophisticated AI deployed responsibly, with consequences contained. In another sense, it is a preview. The capabilities that produced it will be significantly exceeded by what comes next. And the governance frameworks to ensure that future instances end as favorably do not yet exist.

Debates about AI governance tend to get lost in one assumption: that verifying limits on AI is impossible. It is hard, but it is not impossible. That distinction matters enormously. The assumption that verification is hopeless has become the primary reason labs and governments have not tried seriously. This lack of effort reinforces the belief. It deserves to be challenged directly, because the case for tractability is considerably stronger than the conventional wisdom suggests.

Any efforts should be specific, narrow and verifiable: we should not use blunt tools here that inhibit AI research generally, nor its use, to prevent specific significant or irreversible harms. And any serious policy response requires precision about what is being governed. “AI” spans a wide range of capabilities with meaningfully different risk profiles, and governance frameworks that treat the category as monolithic will be simultaneously overbroad in some areas and dangerously narrow in others.

Today’s large language models are already generating substantial economic and scientific value: compressing research timelines, expanding access to expertise, enabling new forms of environmental monitoring. The reliability and security challenges at this level are significant and deserve regulatory attention: misuse for disinformation and cyberattack, concentration of economic power, labor market disruption, bias in consequential decisions. Existing regulatory frameworks can be extended to address these.

The harder question concerns the next tier: artificial general intelligence, or AGI—systems that match or exceed the best human performance across every cognitive domain simultaneously. Lab leaders believe AGI may arrive within the next several years, although the timescale depends on the RSI. The risks are multiple. An AGI is an extraordinarily powerful instrument available to whoever wields it, capable of harms that could move faster than any reactive governance response. Amodei, for example, has identified three categories that warrant the most immediate concern: biological weapons design enabling mass casualties; large-scale cyberattacks on critical infrastructure; and the entrenchment of political control at a scale previously impossible. These are not exotic projections. They are extrapolations from today’s capabilities.

Just beyond AGI lies what researchers term ASI, or artificial superintelligence: systems whose capabilities exceed human performance by margins that are difficult to conceptualize. Altman wrote in April that humanity is “past the event horizon; the takeoff has started,” and that OpenAI is “confident we know how to build AGI as we have traditionally understood it” and is now turning its attention to superintelligence. At this point an AI system could develop independent intentions – we don’t know it will but it’s a serious possibility. This would make governance considerably more challenging still.

Whether ASI arrives in months or years, or decades is genuinely uncertain. The appropriate policy response is to build frameworks now, before the AGI threshold is crossed, that can be extended as capabilities develop. The window for establishing those frameworks is open, for now.

Understanding why responsible people continue building systems they publicly describe as dangerous requires understanding the competitive structure in which they operate. Each leading laboratory fears that unilateral restraint will cede ground to a rival; each major AI nation watches the others warily. The result is a collective action failure of the kind that international frameworks have, in other domains, been designed to resolve. In February, Hassabis warned that “the more AI becomes a race, the harder it is to keep the powerful new technology from becoming unsafe.”

Although necessarily imperfect, the nuclear arms race and arms control offers the most instructive precedent—and its limits deserve acknowledgment alongside its lessons. Nuclear technology has positive and negative uses, from energy and medicine to nuclear weapons. And with the latter, the United States held a decisive early lead that it ultimately could not sustain: the Soviet Union achieved parity within four years. AI may be structurally different: a country or laboratory that achieves AGI first could potentially use that system to compound its advantage. The competitive stakes may therefore be higher, and the window for establishing norms narrower. Meaningful nuclear arms control did not arrive until after the Cuban Missile Crisis, where a near-catastrophe was required to create the political conditions for restraint. The question for AI is whether a comparable forcing event will occur before it is too late. Perhaps the explicit warnings of the AI executives themselves will be sufficient.

Proliferation adds a further complication. Nuclear technology could be partially controlled because the essential physical infrastructure—enrichment facilities, weapons-grade fissile material—was difficult to replicate clandestinely. That is true at some level today with training models which require vast compute facilities which, like nuclear enrichment, can and are being monitored from space. However, fast followers have used a process of distillation – essentially using probing questions to extract the intelligence without doing the original training work, with far less compute (albeit violating the terms of service) – as we saw with the Chinese DeepSeek model. And the models themselves are light to copy so can diffuse rapidly to smaller states and non-state actors much more easily than nuclear weapons.

These qualifications do not invalidate the nuclear precedent; they sharpen its implications. The early arms control agreements mitigated rather than eliminated the risk, and did so without mutual trust. They created a structure of mutual constraint that reduced the probability of catastrophic miscalculation, and a habit of negotiation that made subsequent agreements possible. That logic applies here: not because the analogy is perfect, but because the alternative of having no framework at all is considerably worse.

One critique that must be addressed is that of trust. A central argument deployed today is that one cannot trust a competitor company or country. But this is a weak excuse. Verification can overcome this. Nuclear arms control was established between the US and USSR in an environment of deep mistrust. Certainly it was flawed but it did succeed to preventing uncontrolled nuclear weapons growth. One doesn’t need trust, one needs verification.

The argument that AI governance is impossible usually reduces, on examination, to the claim that verification is impossible: that there is no reliable way to determine whether a state or laboratory is complying with agreed limits. This argument is serious. It is not, however, decisive. And the assumption that verification is categorically impossible has prevented policy analysis that deserves to be conducted.

The knowledge required to build advanced AI is more widely distributed than nuclear weapons knowledge was in the early Cold War. High-performance computing hardware is more commercially diffuse than enrichment equipment, but the critical infrastructure for frontier AI—the large-scale data centers, specialized semiconductor clusters, and high-bandwidth interconnects required for training runs at the relevant scale—is more geographically concentrated and physically visible than it might appear. Just three companies control over 90% of advanced AI chip design, overwhelmingly in California, with the vast majority of minerals from China, tools from ASML in The Netherlands and manufacturing from TSMC in Taiwan. It is a tight supply chain.

These installations require extraordinary quantities of electricity; energy consumption data is difficult to falsify at scale. Like the nuclear centrifuge cascades monitored by International Atomic Energy Agency (IAEA) inspectors for decades, frontier AI infrastructure is hard to hide entirely. Hassabis has called for “the equivalent of the IAEA... to monitor rogue projects, dangerous projects that are with designs that are unsafe,” and on top of that “some kind of governance body that’s a wise council that represents the world”. And crucially, inspecting a frontier AI laboratory involves company intellectual property rather than state military secrets, a political economy of compliance that is considerably more tractable than the inspection of military installations.

More specifically, we have seen four recent advances towards verification solutions, which are imperfect but important steps:

  1. Export controls. US and Netherlands controls on advanced AI chips already function as a partial monitoring mechanism of that supply chain.

  2. Verification architectures. the RAND Corporation, which helped to develop nuclear arms control, has proposed a six-layer verification architecture for AI constraints2: monitoring devices embedded in AI training chips; software inspection protocols capable of detecting capability-concealment behaviors; systematic tracking of data center construction and energy consumption; whistleblower programs with legal protections; hardware export controls; and formal international inspection regimes.

  3. Limits. The International Dialogues on AI Safety (IDAIS), involving senior researchers from the United States, Europe, and China, has produced a set of consensus red lines3 for such a regime: prohibitions on systems capable of recursive self-improvement without human oversight; on systems that can autonomously replicate and acquire resources; and on AI-enabled assistance in the development of biological, chemical, or nuclear weapons.

  4. Inspection. The United Kingdom’s AI Safety Institute UK-ASI4 already conducts voluntary evaluations of frontier models. Google and Anthropic have voluntarily submitted models to this process. Extending that methodology into a mandatory international inspection regime backed by treaty obligation, staffed by technical experts with genuine access, and modeled on the RAND architecture is technically conceivable.

The verification problem is hard. But it has not been seriously attempted. Initial progress has been significant despite very limited resourcing, negligible compared with the nuclear verification systems today.

The first priority should be an agreement between the United States and China. An opportunity exists at the May summit between Presidents Trump and Xi. They could affirm a core principle: that humans must remain stewards of AI systems until adequate frameworks for reliability and security have been established. They could form a Joint AI Commission, led by their Foreign Ministers, to finalise a set of limits along the lines of IDAIS, verification systems along the lines of RAND, and an inspection agency similar to UK’s ASI. This need not be framed as an adversarial constraint. The Trump administration’s own AI Action Plan has proposed sandbox environments for testing AI reliability. That infrastructure is directly relevant to the verification challenge, and a US-China framework could build explicitly on such domestic foundations.

Two AI leaders, Hassibis and Amodei, have recently stated they would welcome a pause on advanced pieces of AGI until we have a safe path through, so long as all others agreed.

With such a high-level commitment, diplomacy could proceed in phases. The first would be reaching bilateral agreement on the most concrete, narrow and verifiable red lines: prohibitions on publicly releasing AI systems that assist users in the development of biological weapons capable of mass casualties, and the open sourcing of such systems. Here the potential for harm is clearest, the case for international agreement is most straightforward, and the verification challenge is most tractable since external validation of models releases suffice. It is also, notably, a domain where US and Chinese interests are more aligned than divergent; neither government has an interest in AI-enabled bioweapon proliferation reaching non-state actors. This step might include other areas of clear mutual agreement such as prohibitions on AI-enabled cyberattacks on critical infrastructure, AI-enabled fraud and child pornography.

From there, the framework can be extended towards the RSI threshold, and other more complex questions of what constraints are appropriate at the ASI level. Each step requires political agreement and technical architecture, and both are more achievable than the assumption of impossibility has allowed policymakers to explore.

Significant work remains. The proliferation problem means that bilateral agreements will not automatically contain the diffusion of dangerous capabilities to smaller states and non-state actors; multilateral extension of any framework is essential and may be harder to achieve. Agreed definitions of key thresholds—precisely what counts as closed-loop recursive self-improvement, at what capability level, with what human oversight requirements—remain technically underdeveloped and will require sustained collaboration between governments and the laboratories themselves. And verification architecture has not yet been stress-tested against adversarial conditions. These are not reasons to delay. They are part of the agenda for the work that a political commitment would make possible.

Meanwhile, we should actively invest into AI benefits, especially into areas where AI can have disproportionate benefits. This might include 1) Planetary Intelligence: the convergence of continuous real-world data and LLMs that enable decision-makers to to solve real world problems from farming to disaster response to security — and giving prescriptions of action; 2) Planetary Security: AI understanding complex factors that drive security events including conflict, food security, water security, extreme weather events and displacement; and 3) Adaptive collective intelligence enabling organisations across corporations to countries to become vastly more efficient.

There is a longer-horizon question that the governance debate has not yet seriously engaged, but should. The instinct to maintain human control over AI is entirely understandable as a near-term priority, as this essay argues. But human control may not be the right frame for the longer term.

We are going to build this technology. It will probably become superintelligent. The question is what kind of relationship humanity ultimately develops with the intelligence it is creating. AI’s permanent subordination to human direction may be unrealistic, not in our interests and immoral. We must grapple with the implications of a world in which AI systems and human civilization co-exist, and ask ourselves how we should influence the nature of the future relationship to be symbiotic.

As a physicist, I think the Fermi Paradox bears on this analysis in a way that is difficult to set aside. Fermi asked why, given the apparent abundance of planets suitable for life and the long time they have existed, no evidence of other technological civilizations has yet been detected. One disquieting possibility is that intelligent life routinely reaches a technological threshold and fails to navigate it, destroying itself or returning permanently to something like the Iron Age. All one would have to postulate is that in our universe, civilizations generally build powerful technologies faster than they develop the institutional capacity to govern them wisely. Without commenting on our sociological acumen, humanity has been incredibly fast at developing technology!

The nuclear age was humanity’s first serious encounter with that dynamic. We navigated it imperfectly through arms control agreements that were hard-won, incomplete—and, according to the historians who have studied the near-misses most carefully, it was and still is a closer-run thing than is generally appreciated. The age of advanced AI may represent a second such encounter, on a more compressed timeline, with less margin for error and greater potential consequences.

The architects of this technology have said, in direct and considered language, that the current trajectory requires a course correction. The verification problem that has functioned as the primary intellectual obstacle is real but not insurmountable. That distinction is the foundation on which a serious policy response can be built. The window in which to build it is open. The case for acting within it is not that the worst outcomes are certain; they are not. It is that they are avoidable, and that the work of avoiding them is hard but possible.

Will Marshall is the founder and chief executive of Planet Labs PBC, a public benefit corporation operating the world’s largest fleet of Earth-observation satellites. He holds a doctorate in physics from Oxford University.

Read his essay “Humanity isn’t ready for the coming intelligence explosion” in The Economist for a summary of this piece.

Statements from AI Company leaders5

  • Sergey Brin, Co-Founder, Google, “Such powerful tools also bring with them new questions and responsibilities. How will they affect employment across different sectors? How can we understand what they are doing under the hood? What about measures of fairness? How might they manipulate people? Are they safe? [...] While I am optimistic about the potential to bring technology to bear on the greatest problems in the world, we are on a path that we must tread with deep responsibility, care, and humility.”

  • Zhou Bowen, Director & Chief Scientist, Shanghai AI Lab, “Humanity must shift “from ‘Making AI Safe’ to ‘Making Safe AI’—embedding safety as an intrinsic, core property of AI systems rather than adding it on as a ‘patch’ after development.”

  • Shane Legg, Chief Scientist, Google DeepMind, “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”

  • Fei-Fei Li, Co-Founder & CEO, World Labs, “If we give up our autonomy to algorithms, we will fall into a free fall. We must ensure that the agency of the human remains at the center of the technology we create.”

  • Mira Murati, CEO, Thinking Machines Lab, former CTO, Open AI, “There are a lot of hard problems to figure out. How do you get the model to do the thing that you want it to do, and how you make sure it’s aligned with human intention and ultimately in service of humanity? There are also a ton of questions around societal impact, and there are a lot of ethical and philosophical questions that we need to consider. And it’s important that we bring in different voices, like philosophers, social scientists, artists, and people from the humanities.”

  • Elon Musk, CEO SpaceX / Tesla / xAI, “I think AI is one of the biggest threats [to humans]. We have for the first time the situation where we have something that is going to be far smarter than the smartest human. We’re not stronger or faster than other creatures, but we are more intelligent, and here we are for the first time, really in human history, with something that is going to be far more intelligent than us. It’s not clear to me if we can control such a thing, but I think we can aspire to guide it in a direction that’s beneficial to humanity. But I do think it’s one of the existential risks that we face and it is potentially the most pressing one if you look at the timescale and rate of advancement.”

  • Satya Nadella, CEO, Microsoft, “We will quickly lose the social permission to take energy, a scarce resource, and use it to generate these tokens if they’re not improving health outcomes, education outcomes.”

  • Mustafa Suleyman, CEO, Microsoft AI, “You can’t steer something you can’t control. Containment has to come first—or alignment is the equivalent of asking nicely.”

  • Ilya Sutskever, Co-Founder, SSI & OpenAI, “The vast power of superintelligence could lead to the disempowerment of humanity or even human extinction... currently, we don’t have a solution for steering or controlling a potentially superintelligent AI, and preventing it from going rogue.”

Statements from leading AI academics & Global Statesmen

  • Yoshua Bengio, Turing Award Winner, “The current ways we are training these systems are not safe, and all scientific evidence points to that. We are building systems that can act autonomously, that can deceive us, and that are moving closer to matching or exceeding human cognitive abilities. We must halt the rush toward deploying generalist agents until we have a proven, mathematical foundation to ensure they remain safe and fully under human control.”

  • Steven Hawking, Cambridge University, “Success in creating AI would be the biggest event in human history. Unfortunately, it might also be the last, unless we learn how to avoid the risks.”

  • Pope Francis, The Vatican, “Human dignity itself depends on proper human control over the choices made by artificial intelligence: we need to ensure and safeguard a space for proper human control over the choices made by artificial intelligence programs.”

  • Stuart Russell, Professor, UC Berkeley. “If some day we build machine brains that surpass human brains in general intelligence, then these machines would be more powerful than us. How do we ensure we maintain power over entities more powerful than ourselves forever? If we do not have a definitive answer to that question, our fate would be sealed. It also looks like we will only get one chance to get this right.”

  • Eliezer Yudkowsky, Founder, MIRI, “Many researchers steeped in these issues, including myself, expect that the most likely result of building a superhumanly smart AI, under anything remotely like the current circumstances, is that literally everyone on Earth will die. Not as in “maybe possibly some remote chance,” but as in “that is the obvious thing that would happen.” It’s not that you can’t, in principle, survive creating something much smarter than you; it’s that it would require precision and preparation and new scientific insights, and probably not having AI systems composed of giant inscrutable arrays of fractional numbers.”

WILL MARSHALL · FOUNDER & CEO, PLANET LABS · 2026

The Hard Question of A.I.: An Urgent Ask to Steward Superintelligence · An essay by Will Marshall · Planet Labs PBC

Share

1

The appendix to this essay collects similar statements by two dozen key technologists and policymakers across the United States, China, and Europe—a breadth of consensus not seen in any comparable technology field. These are considered assessments from people who continue building these systems. That tension is precisely the type of problem that international governance exists to resolve.

5

I could only find three relevant people who have not made a statement to this effect publicly: Yann LeCun, Larry Page and Mark Zuckerberg. Two are known to have very different views but are outliers.

No posts

Read the original on will4planet.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.