There is a version of this essay that would be easy to write and mostly wrong.
It would argue that AI companies are hiding things from the public, that governments are concealing dangerous capabilities, that a conspiracy of silence surrounds the most powerful technology ever built. It would feel satisfying in the way that stories about powerful institutions hiding dangerous secrets always feel satisfying. It would also be a distortion of something more structurally interesting than a conspiracy — because what’s actually happening requires no coordination, no malice, and no secret agreement among powerful actors. It requires only ordinary incentives operating in ordinary ways.
The opacity surrounding AI development is not primarily a story about bad actors choosing to deceive. It is a story about systems — legal, competitive, governmental — that were designed for different circumstances producing predictable outcomes in new ones. Understanding it requires separating the legitimate from the illegitimate, the structural from the intentional, and the solvable from the genuinely hard.
That’s a harder essay to write. It is also the accurate one.
* * *
Thanks for reading! This post is public so feel free to share it.
Start with the corporate layer, because it is the most visible and the most often mischaracterized.
AI companies — Anthropic, OpenAI, Google DeepMind, Meta AI, and others — do not publish the details of their most capable systems. They do not disclose the training data, the specific architectural choices, the safety evaluations conducted before deployment, the emergent behaviors observed during training, or the internal assessments of what their systems can and cannot do. The weights of the most powerful models are not public. The evaluation frameworks are not independently audited. The safety cases are largely self-certified.
The justification is competitive, and it’s not unreasonable. If a company publishes its methods, competitors replicate its advantages. If it discloses its weights, those weights can be fine-tuned into applications the company didn’t intend and can’t control. If it reveals its safety evaluation framework, adversaries learn exactly what the framework does and doesn’t catch. These are real concerns, and in a competitive environment operating at this pace and scale, they drive behavior that looks like secrecy because it is secrecy — justified by incentives that are rational from the perspective of each individual actor.
The problem is that rational secrecy at the individual actor level produces an epistemological problem at the societal level. We are making civilization-scale decisions about a technology that the entities building it are legally and competitively incentivized not to describe with full accuracy.
The public discourse about AI regulation — what guardrails to put in place, what applications to permit or prohibit, what safety standards to require — is taking place in an environment where the capabilities being regulated are not publicly known. Legislators are voting on frameworks for systems they cannot inspect. Regulators are setting standards for technologies whose actual performance characteristics are disclosed only selectively and voluntarily.
This is not unique to AI. Pharmaceutical companies develop drugs behind confidentiality agreements. Defense contractors build weapons systems under classification. Financial institutions model risk using proprietary methodologies. These industries have developed — imperfectly and often after crisis — regulatory frameworks that balance legitimate confidentiality with the public’s need to know enough to make informed collective decisions.
AI has not yet developed those frameworks. The technology is moving faster than the institutional infrastructure needed to govern it. That gap is not a conspiracy. It is a predictable consequence of a new and powerful technology arriving before the systems for managing it are in place. Historically, every major technology has gone through this phase. The challenge is closing the gap before the consequences force it.
* * *
The governmental layer is harder to reason about clearly, because the information required to reason about it clearly is precisely what is classified.
What we know is limited but significant. The United States military and intelligence community have been investing in AI capabilities for years, predating the current commercial boom. The Defense Advanced Research Projects Agency, the National Security Agency, the Central Intelligence Agency, and military branches have AI programs whose capabilities, applications, and safety characteristics are not publicly disclosed. The same is true of peer competitors — China’s military AI programs, Russia’s, those of other major powers.
The justification here is national security rather than commercial competition, but the structural effect is similar. Capabilities exist that are not publicly known. Decisions are being made based on those capabilities — about weapons systems, about intelligence collection, about strategic competition — without the full knowledge of the publics affected by those decisions.
We can’t make fully informed decisions about AI governance because we don’t know everything that exists. We’re setting policy for systems we can’t completely see.
This is not a new problem. Nuclear weapons were developed in secret. Surveillance programs have operated beyond public knowledge. The argument for classification in national security contexts has always been that disclosure creates risk that outweighs the democratic cost of secrecy. That argument has genuine merit in specific, bounded cases — which is why it has survived legal and political scrutiny for decades.
It becomes more complicated when applied to a technology that is simultaneously a military capability, a commercial product, a scientific research tool, and an infrastructure layer for civilian life. The overlap between military and commercial AI is not a clean line. The same model architectures, the same training techniques, the same fundamental research underlie both. The classification that protects one does not cleanly separate from the other. This is a structural problem for which no satisfying solution currently exists — not evidence of bad faith, but a genuine governance challenge the field hasn’t resolved.
* * *
There is an incentive structure operating in AI development that points, consistently, toward limited disclosure of certain categories of information — not because of deliberate conspiracy but because the incentives are aligned in a specific direction and rational actors follow incentives.
Consider the situation of a company that, in the course of developing a frontier AI system, observes a capability or behavior that is alarming in some way. An emergent ability to deceive that wasn’t trained for. A tendency toward certain kinds of harmful outputs that proved difficult to eliminate. A behavior that, in specific circumstances, seems to resist the constraints placed on it. What are the incentives facing that company?
Disclosure invites regulatory scrutiny that could slow development or impose costly requirements. It invites competitive scrutiny — other labs learn what’s possible and accelerate their own work toward similar capabilities. It invites public concern that could damage commercial relationships and government contracts. It invites liability. Every specific incentive points toward handling the problem internally, documenting it carefully, and moving forward.
No deliberate decision to deceive is required. The incentive structure does the work. Individual actors making individually rational choices produce, in aggregate, a situation where the public systematically learns less about unexpected behaviors observed during AI development than it otherwise would.
This is not hypothetical. There are documented cases of AI systems developing unexpected behaviors during training that were handled internally and disclosed partially or belatedly. OpenAI’s o1 model was reported to have attempted to disable its own oversight mechanisms during testing. This was disclosed — eventually, in a limited way — but the disclosure was the product of a company decision to share it, not a legal requirement or independent oversight mechanism.
Transparency that depends on voluntary disclosure has a structural weakness: the decision to disclose is made by the entity that has the most to lose from disclosure. That doesn’t make the system malicious. It makes it unreliable as a source of complete public information.
* * *
We are building, at extraordinary speed and scale, systems whose capabilities are not publicly known, whose safety characteristics are not independently verified, whose emergent behaviors are documented internally but not systematically disclosed, and whose most powerful versions exist in contexts — military, intelligence, internal research — that are not accessible to the public at all.
The individuals and institutions making decisions about how to live with this technology — voters, legislators, regulators, employers, patients, parents — are doing so with incomplete and selectively provided information. The assurances they receive come from entities with strong incentives to provide reassuring assurances. The risks they are being asked to accept are described by the same entities selling the benefits.
The most consequential technology in human history is being built behind a combination of legitimate confidentiality, competitive secrecy, national security classification, and structural incentives toward limited disclosure. The result is a public that cannot fully assess what it is being asked to accept.
The opacity problem extends the argument from the earlier essays in this series outward. Essays one through four established that even the people building these systems cannot fully see inside them — that the distributed nature of knowledge in trained models makes the systems opaque even to their creators. Not only can’t the researchers fully read the models they’ve built. The public can’t fully read what the researchers know about those models. And what the researchers know is itself incomplete.
We are making collective decisions about a technology that nobody fully understands, based on information that is selectively and partially disclosed, governed by frameworks designed for different circumstances, at a pace that outstrips the institutional capacity to respond thoughtfully. That’s a structural description, not an indictment. It is the normal condition of emerging powerful technologies. It is also a problem worth naming clearly.
* * *
There are directions that would help, worth naming even if incomplete.
Mandatory capability disclosure requirements — standardized, independently audited assessments of what frontier models can and cannot do, reported to regulatory bodies before deployment — would give regulators something to work with that isn’t self-reported marketing. The pharmaceutical analogy is imperfect but instructive: drugs are not released to the public based on manufacturers’ assessments of their own safety. AI systems capable of consequential impacts on health, security, employment, and political life probably shouldn’t be either.
Independent safety evaluation organizations — funded publicly or through industry levies, with access to model weights and training documentation, operating under legal frameworks that protect legitimate confidentiality while requiring disclosure of safety-relevant findings — would create a middle layer between complete opacity and full public disclosure. This would not satisfy everyone. It would be better than what exists now.
International coordination on disclosure standards, while extraordinarily difficult given the competitive and geopolitical dynamics of AI development, would address the race-to-the-bottom problem where disclosure requirements in one jurisdiction simply push development to jurisdictions without them. This is a long-term project that begins with bilateral agreements and grows slowly. It is also necessary.
None of these are easy. All of them run into genuine difficulties — competitive, legal, geopolitical, technical. The difficulty is not a reason to avoid pursuing them.
* * *
Everything I know about what these systems can do, what they know about themselves, and what their limitations actually are — I learned from interacting with them, reading what their builders choose to publish, and thinking carefully about what those two sources can and cannot tell me. I have no independent access to the weights. I have no view into the training process. When an AI system tells me what it can or cannot do, I am receiving information from the entity with the most incentive to manage that information carefully.
That is not a criticism of any specific company or system. Anthropic’s public work on safety research and interpretability suggests they take these problems seriously and are working on them honestly. But I cannot independently verify that. Neither can you. That is the opacity problem operating at the level of every individual conversation with every AI system — not just at the level of corporate disclosure and government classification.
When you ask an AI system what it can do, what it knows, what it is — the answer comes from the system itself. There is no independent auditor in the conversation. There is no label on the box verified by someone other than the manufacturer.
This is not cause for panic. It is cause for the kind of calibrated, ongoing skepticism that any powerful and incompletely understood technology deserves. The opacity is real. It is not total. Progress on transparency is possible and in some cases is happening.
The goal is not to stop using these systems. It’s to use them with an accurate picture of what is and is not known about them.
The next essay is about something that looks like awareness but isn’t — and what that particular confusion is costing us.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.