Sam Altman, (@sama) recently posted about Astra, their next generation of models that are an order or two of magnitude more capable than GPT 5.6 Sol or Claude Fable.
Today, Daniel Litt (@littmath) posted a breezy thread stemming from his talk entitled “The End of Mathematics.” You can read it here:
Daniel Litt@littmath
I'm currently returning to Toronto from a summit on the future of mathematics, at OpenAI. @SebastienBubeck asked me to talk a bit about the future we'd all like to avoid, where humans are mathematically disempowered. @Jacob_Tsimerman advised us to try to prioritize detail over

10:18 PM · Aug 11, 2026 · 12.6K Views
15 Replies · 33 Reposts · 287 Likes
So what are the implications of this?
The mathematics scenario generalizes into a broader risk:
AI - even on that is just an order of magnitude more capable than today’s frontier - could make civilization vastly more productive while making individual humans vastly less able to understand, verify, or direct what civilization is doing.
Call it epistemic disempowerment. The machine produces the diagnosis, proof, design, legal theory, trading strategy, software architecture, or policy. Humans approve it because the alternatives are slower and usually worse. Eventually, the relevant humans cannot independently reconstruct why it works.
One caveat: 10 times more pretraining compute alone does not guarantee all of this. The largest gains would require good data, strong post-training, tools, memory, experimentation environments, and enough inference compute. But current systems are already advancing in long-horizon software work, mathematical discovery, cyber research, biomedical reasoning, scientific tool use, weather prediction, and embodied robotics. Extrapolation is no longer based purely on science fiction.
In most areas, progress has four stages:
Generate a candidate answer.
Determine whether it is correct.
Turn it into something deployable.
Diffuse it through institutions and everyday life.
AI is rapidly accelerating stage one. In software and mathematics, it is beginning to accelerate stage two as well. In medicine, materials, manufacturing, and energy, physical testing, regulation, capital investment, and construction remain major bottlenecks.
So the effects will arrive unevenly:
Digital work changes almost immediately.
Medicine and science change first in research and decision support, then later in actual treatments.
Physical industry changes more slowly, but potentially more profoundly.
Government and education may change slowly until they suddenly face institutional crisis.
Instead of writing isolated functions, it can understand the actual operating model of a company’s software:
What the system is supposed to do
What it really does
Which undocumented assumptions keep it working
Which services are redundant
Which bugs share a common architectural cause
How to migrate without breaking customers
How to monitor whether the migration succeeded
A company might say:
“This codebase has been developed by 80 people over nine years. Simplify it by half, preserve every customer-visible behavior, close the major security gaps, and move it to a maintainable architecture.”
The model could inspect the repositories, logs, issue history, infrastructure, customer tickets, and database schemas, then execute the project over days.
METR’s evaluations already frame agent capability in terms of the duration of professional software tasks that models can complete autonomously, and report a strong historical increase in that time horizon. Current measurements are still narrow and imperfect, but the direction supports the idea that the important frontier is moving from prompts toward projects. (Metr)
Most people will not experience this as “AI coding.” They will experience it as:
Apps that adapt to exactly how they work
Bugs being repaired much faster
Internal tools becoming available to small businesses
Software integrations that no longer require months of consulting
Personal automations assembled from plain-English requests
Fewer forms and screens because agents directly produce outcomes
You might say:
“Move my recurring bills to the card with the best benefits, except rent, cancel anything unused, and ask me before making a change that affects insurance or credit.”
The system would operate across multiple services rather than merely explaining how to do it.
Software becomes less like a fixed product and more like a temporary expression of an objective. Companies may stop buying rigid workflow software and instead buy reliable agents plus access to the necessary systems.
The economic value moves away from merely writing code and toward:
Owning distribution
Controlling proprietary data
Defining objectives
Verifying outcomes
Holding legal responsibility
Operating trusted infrastructure
Organizations may deploy systems that no employee fully understands. Engineers become supervisors of generated architectures rather than builders who possess deep internal models.
When something unusual fails, the company may need another model to explain what the first model built.
That is the software version of mathematical disempowerment.
A sufficiently stronger model would not merely scan for familiar vulnerability patterns. It could construct a model of the intended security boundaries, find where the actual system violates them, develop a safe reproduction, and propose a verified repair.
OpenAI has said Astra’s preliminary evaluations were strong enough that it could not rule out its Critical cyber threshold, which includes highly consequential autonomous vulnerability discovery and exploitation capabilities. GPT-5.6 is also evaluated on realistic vulnerability research with long-running tool use. (OpenAI)
For normal people:
Fraud detection improves
Account takeovers may be caught earlier
Security patches arrive faster
Scam messages become vastly more convincing
Voice and video identity become unreliable
Banks and services require stronger proofs of identity
“Did you really send this?” becomes a routine question
Passwords become less important than hardware keys, device identity, continuous authentication, and cryptographic authorization.
Cyber conflict accelerates from human-speed campaigns to machine-speed search and response.
One model discovers a vulnerability. Another detects the attempted exploitation. A third writes and tests the patch. A fourth searches whether the patch created a different weakness.
National security increasingly depends on:
Compute access
Model access
Secure deployment
Software provenance
Rapid patch distribution
The quality of defensive autonomous agents
Human security teams may no longer be able to independently understand every attack chain. They approve model-generated remediations because waiting for complete human comprehension would leave the system exposed.
Cybersecurity could become safer on average and more catastrophic in the tail.
A stronger medical system could synthesize:
Medical history
Symptoms
Laboratory trends
Imaging reports
Medication interactions
Family history
Wearable data
Clinical literature
Insurance constraints
Patient preferences
It would not simply offer a diagnosis. It would construct and continuously revise a model of the patient.
Current healthcare evaluations already emphasize realistic conversations, clinical reasoning, safety, uncertainty handling, documentation, and medical research rather than multiple-choice recall. OpenAI has also introduced a life-sciences model oriented toward drug discovery, genomics, chemistry, and protein reasoning. (OpenAI)
The first major benefits are likely to be mundane but enormous:
Preparing a complete brief before an appointment
Noticing that a symptom pattern has changed
Explaining test results in understandable language
Checking medication combinations
Finding appropriate specialists
Matching patients to clinical trials
Drafting prior-authorization appeals
Reconciling contradictory advice from multiple doctors
Tracking whether a treatment is actually helping
Instead of remembering everything during a 15-minute visit, the patient arrives with a structured timeline, unresolved questions, likely explanations, and warning signs.
The model might say:
“Your last four visits treated these as separate complaints. The symptoms began within the same three-week period and changed after the medication dose increased. That does not prove causation, but it is worth asking your clinician whether the timing changes the differential.”
Healthcare gradually moves from episodic intervention toward continuous prevention.
The system could detect deteriorating patterns before the patient knows what they mean:
Early heart failure signals
Medication adherence problems
Sleep and mood interactions
Subtle cognitive decline
Postoperative complications
Dangerous combinations of symptoms that were separately dismissed
Medical research could also shift from manually designed broad studies toward machine-generated hypotheses about subpopulations, biomarkers, and treatment combinations.
Doctors become dependent on systems whose reasoning is too broad or complex to reproduce during ordinary care.
A model might correctly combine 200 weak signals into a recommendation, while no clinician can fully audit the interaction. Responsibility becomes unclear when the model is more accurate statistically but wrong for a particular patient.
The important bottleneck will not just be intelligence. It will be prospective clinical evidence, accountability, and calibrated trust.
The major transition is from AI making predictions to AI running a scientific loop:
Read the literature and experimental history.
Develop competing biological explanations.
Design molecules, proteins, or experiments.
Send instructions to an automated laboratory.
Analyze the results.
Update the underlying theory.
Choose the next most informative experiment.
AlphaFold has already given millions of researchers access to predicted protein structures, while newer life-sciences systems target chemistry, genomics, protein engineering, and ambiguous research-level biological judgments. Google DeepMind has also announced an automated materials laboratory integrated with its models, showing the broader move toward closed-loop experimentation. (Google DeepMind)
These effects take longer to reach ordinary life, but could eventually produce:
Faster development of treatments for rare diseases
Better matching of drugs to patient subtypes
Vaccines that can be redesigned more rapidly
Cheaper diagnostic tests
Enzymes that reduce industrial pollution
Crops resistant to drought, heat, and disease
Less toxic pesticides
New ways to manufacture chemicals biologically
The immediate effect will often be shorter research cycles, not instant cures.
Biology becomes increasingly programmable.
The machine does not merely predict what a biological system will do. It proposes interventions that produce a desired behavior:
A protein that binds a specific target
A delivery system that reaches one cell type
An enzyme that operates under industrial conditions
A gene circuit that responds to an environmental signal
A microbial process that manufactures a scarce material
The same competence that helps design therapies can lower barriers to harmful biological work. Another issue is scientific opacity: researchers may generate successful interventions without possessing a satisfying mechanistic explanation.
We could get medicines that work before humans fully understand why.
That is valuable, but it changes how evidence, safety, and scientific understanding interact.
A stronger model could search for materials with several properties simultaneously:
Cheap inputs
High energy density
Long life
Low toxicity
Manufacturability
Temperature stability
Recyclability
GNoME predicted 2.2 million crystal structures, including 380,000 candidates predicted to be stable. The challenge now is selecting, synthesizing, characterizing, and manufacturing the useful candidates. Automated laboratories linked to general reasoning systems could compress that bottleneck. (Google DeepMind)
The eventual consumer outcomes could include:
Phones and laptops with materially longer battery life
Electric vehicles that charge faster and last longer
Cheaper home energy storage
Better heat pumps
Lighter and stronger construction materials
More efficient air conditioning
Lower-cost solar panels
Better water filtration
More durable roads and buildings
Reduced use of scarce minerals
These outcomes will not arrive as a chatbot feature. They arrive through factories, supply chains, permitting, and capital equipment.
Energy becomes cheaper and more abundant if AI helps improve several layers at once:
Battery chemistry
Power electronics
Grid forecasting
Catalyst design
Solar materials
Reactor control
Industrial heat
Carbon capture
Semiconductor efficiency
Even modest improvements across several layers could compound into major reductions in the cost of electricity, transportation, computation, and manufacturing.
AI may generate millions of plausible candidate materials faster than laboratories can test them. We recreate Litt’s mathematical-output problem in physical science:
An enormous reservoir of theoretically valuable discoveries that civilization lacks the experimental and industrial capacity to evaluate.
The winning countries and firms may be those with automated labs, power, factories, permitting capacity, and supply chains, not simply those with the smartest model.
The critical capability is not graceful walking. It is understanding progress:
What has already been completed
Whether a step succeeded
What went wrong
Whether to retry
When to ask for help
How to coordinate multiple machines
Current embodied models already focus on video understanding, real-time progress tracking, tool orchestration, success detection, unfamiliar multi-step tasks, and collaboration among robots. They remain limited, especially outside familiar distributions, but these are the exact capabilities required for economically useful autonomy. (Google DeepMind)
The first widespread effects are more likely to appear behind the scenes:
Warehouse handling
Hospital supply delivery
Industrial inspection
Cleaning commercial buildings
Agricultural harvesting
Infrastructure maintenance
Construction-site monitoring
Dangerous-environment operations
Assistance in senior living facilities
Home robots arrive later because homes are chaotic, safety-sensitive, varied, and filled with fragile objects, pets, and humans.
A practical home system might initially be less like a general humanoid servant and more like several specialized devices coordinated by one model.
Physical labor begins acquiring software-like scalability.
A company could demonstrate a task once, let the system form a procedure, and distribute the learned behavior across thousands of machines. Physical expertise becomes transferable with much lower marginal cost.
That could increase industrial capacity dramatically, especially in countries with aging populations or labor shortages.
Workers may lose not only jobs but also the practical knowledge needed to recover when systems fail.
A plant could become extremely productive while depending on machines whose learned procedures no local employee can fully reproduce.
This is exactly where products that capture physical know-how matter. The valuable artifact is not merely a video recording of the expert. It is a causal model of what the expert notices, why an adjustment matters, and when normal procedure should be overridden.
Weather systems increasingly generate probabilistic scenarios rather than one deterministic forecast. DeepMind has demonstrated experimental cyclone prediction covering formation, track, intensity, size, and shape through ensembles extending as far as 15 days. (Google DeepMind)
The next step is connecting forecasting to decisions:
How should the electrical grid prepare?
Which neighborhoods should evacuate?
When should a farmer plant?
Where should emergency equipment be moved?
Which supply routes are likely to fail?
Which buildings face the highest localized risk?
People get more useful warnings:
“There is a 30 percent chance of damaging wind in your area” becomes “Move your car away from these blocks, charge medical devices before 4 p.m., and expect this road to become inaccessible.”
Other effects include:
More accurate flight and shipping schedules
Better outdoor planning
Earlier wildfire preparation
Lower crop losses
More precise energy pricing
Smarter home heating and cooling
More targeted disaster response
Climate adaptation becomes more granular.
Rather than planning from historical averages, cities and companies continuously simulate infrastructure under changing weather distributions. Insurance, agriculture, construction standards, and emergency planning become forward-looking.
Better forecasts do not guarantee better action. Political trust, evacuation capacity, housing inequality, infrastructure, and institutional coordination still matter.
The model may know exactly what will happen while the relevant organization remains unable or unwilling to respond.
Every student can have a patient, adaptive tutor capable of:
Diagnosing exactly where understanding broke down
Generating examples suited to the student
Switching explanations
Simulating experiments
Giving immediate feedback
Connecting a subject to the student’s interests
Teaching at any pace or level
For a motivated learner, this is extraordinary.
A person could learn:
Statistics through their own business data
Programming by constructing a real product
Biology through an illness affecting their family
A language through simulated daily interactions
Mathematics through interactive visual models
The cost of access to high-quality explanation falls dramatically.
Credentials become less connected to classroom attendance and more connected to demonstrated competence.
Assessment may move toward:
Live oral examination
Real projects
Defending decisions
Reproducing results
Working under observation
Explaining concepts in multiple ways
The same system can eliminate the productive struggle through which competence forms.
A student can submit excellent work without building an internal model of the subject. Over time:
Teachers cannot tell who understands
Students cannot tell when the model is wrong
Employers cannot trust conventional credentials
Fewer people acquire the depth needed to extend or supervise the systems
The danger is not that AI makes students lazy in some moralistic sense. It is that society confuses high-quality output with human capability.
Mathematics is merely the earliest clean example.
A sophisticated legal model could ingest:
Statutes
Regulations
Contracts
Case law
Agency guidance
Correspondence
Evidence
Procedural requirements
Local judicial patterns
It could construct the full argument graph, identify conflicts, draft filings, predict counterarguments, and monitor deadlines.
Ordinary people gain help with:
Lease disputes
Insurance denials
Employment issues
Benefits applications
Tax notices
Small claims
Estate paperwork
Consumer refunds
Contract review
Immigration paperwork
A person might upload a denial letter and receive:
“The company denied this under provision A, but provision A applies only after condition B. Their own earlier letter confirms B did not occur. Here is a draft appeal with the relevant evidence attached.”
The immediate gain is not replacing courtroom lawyers. It is giving millions of people usable help where hiring a lawyer was previously uneconomic.
Litigation and regulation become much more exhaustive.
Every ambiguity can be explored. Every historical precedent can be compared. Every contract can be continuously monitored. Governments could simulate how proposed rules interact with existing law before enactment.
Legal output explodes.
Courts could face machine-generated filings of extraordinary length and sophistication. Regulators may produce rules whose interactions are understood only by other models. Human judges and legislators become dependent on machine summaries.
A formally lawful system can still become democratically illegible.
A government agent could handle the maze between a person’s situation and hundreds of programs, forms, jurisdictions, and eligibility rules.
Instead of finding and completing separate forms, a person says:
“My mother can no longer live independently. Figure out which federal, state, local, insurance, disability, transportation, housing, and tax programs apply, then prepare everything for review.”
Small businesses could ask:
“Tell me what permits and filings are required to open this business at this address, and keep us compliant.”
That could eliminate enormous amounts of bureaucratic suffering.
Governments gain the ability to:
Simulate policy effects
Identify contradictory programs
Detect fraud and waste
Rewrite confusing rules
Target inspections
Forecast service demand
Analyze public comments
Measure whether policies are producing their intended result
The administrative state could become more capable but less contestable.
A benefits decision might rely on a model integrating thousands of features. Even when statistically accurate, the affected citizen may not be able to understand or challenge it.
The core democratic requirement becomes:
A person must be able to know why the state acted, what evidence mattered, and how to appeal.
Without that, optimization becomes domination.
For an individual, the system could continuously optimize:
Cash flow
Debt payoff
Taxes
Insurance
Benefits
Retirement contributions
Subscriptions
Credit usage
Major purchases
Estate planning
Fraud monitoring
Personal finance shifts from periodic advice to continuous management.
The system could notice:
“Your insurance premium increased, your employer added a better plan, and your medication is covered differently this year. Switching during enrollment would likely reduce expected annual cost, but only if your current specialist remains in network.”
It could prepare the decision rather than offering generic educational content.
Companies use models for:
Pricing
Underwriting
Forecasting
Due diligence
Accounting
Fraud detection
Treasury management
Market research
Contract analysis
Small companies gain financial sophistication previously available only to large firms.
Any easily discovered investment advantage disappears quickly because many systems can find it. Competitive value concentrates in proprietary data, execution, capital, regulatory access, and speed.
Financial systems may also become more correlated. If many agents learn similar strategies or respond to the same signals, they could amplify shocks rather than diversify them.
The mathematics scenario repeats in chemistry, neuroscience, economics, physics, climate science, and social science.
A model can generate:
Hypotheses
Experimental designs
Simulations
New measurements
Alternative explanations
Statistical analyses
Replications
Literature syntheses
Instruments
Theoretical frameworks
The valuable step becomes consolidation:
“These 8,000 studies are not 8,000 separate findings. They are noisy observations of four underlying mechanisms.”
AI systems are already being connected to national laboratories, advanced algorithm discovery, biology workflows, materials research, and automated experimentation. The likely next frontier is not merely answering scientific questions, but allocating experiments and revising theories from the results. (Google DeepMind)
Most scientific breakthroughs affect daily life indirectly and with delay:
Better products
Lower costs
Improved medicine
Safer infrastructure
More accurate forecasts
Better agricultural yields
Cleaner industrial processes
The public sees the product, not the scientific reasoning that enabled it.
Larger outcome
The scientific method itself may change from a human-paced conversation among laboratories into a continuous machine-mediated search process.
The scarce resources become:
Experimental equipment
Physical samples
Energy
Compute
High-quality measurements
Human judgment about desirable goals
Institutional capacity to absorb results
Civilization accumulates more correct science than humans can understand.
Scientific truth continues advancing, while science as a human culture contracts.
A model could generate a film, game, lesson, news explanation, or virtual world dynamically around one person.
You could request:
“Make me a six-hour historical drama about the construction of the transcontinental railroad, accurate enough to teach from, but paced like a prestige thriller.”
Or:
“Create a cooperative game for these four friends based on our actual senses of humor and skill levels.”
Entertainment becomes responsive rather than fixed.
The supply of competent media becomes effectively infinite. Human attention remains finite.
Value shifts toward:
Trusted creators
Live events
Shared cultural moments
Authentic identity
Communities
Selection and curation
Intellectual originality
People increasingly occupy personalized realities.
Shared culture weakens because everyone receives the version of events, entertainment, and explanation most compelling to them. Propaganda becomes individualized. Emotional persuasion becomes continuously optimized.
The threat is not merely fake content. It is a world where no two people are shown the same persuasive environment.
A small team can delegate:
Research
Product design
Engineering
Legal review
Financial modeling
Customer support
Sales preparation
Recruiting coordination
Operations
Documentation
More people can start companies without raising large amounts of capital or hiring large staffs. Local businesses gain capabilities previously limited to sophisticated technology firms.
A five-person company may operate with the functional surface area of today’s 50-person company.
There are two opposing effects:
Democratization: capable individuals and small teams can build much more.
Concentration: the owners of models, compute, data, distribution, energy, and trusted platforms can capture an enormous portion of the surplus.
Both can happen simultaneously.
Organizations hollow out their apprenticeship pipeline.
Junior employees traditionally learn by performing bounded work, receiving correction, and gradually forming judgment. If AI immediately performs all junior work, companies may enjoy a burst of productivity and later discover they have produced no future senior operators.
This is the workforce equivalent of Litt’s concern.
At a high level, advanced models can improve:
Intelligence synthesis
Logistics
Maintenance
Cyber defense
Satellite analysis
Planning
Simulation
Drone coordination
Translation
Early warning
Procurement
The strategic advantage may come less from one spectacular autonomous weapon and more from connecting the entire operational system:
Detect a change
Understand its significance
Generate options
Move supplies
Retask sensors
Coordinate assets
Evaluate the result
The country whose systems complete that loop fastest gains leverage.
Human decision-making becomes the slow component.
That creates pressure to delegate increasingly consequential choices. Even if every state wants meaningful human control, each may fear that retaining it places them at a speed disadvantage.
The dangerous threshold is not a conscious machine choosing war. It is institutions allowing automated recommendations to become practically irreversible because the decision window is too short for serious human review.
The ordinary person’s most visible experience will be near-zero-marginal-cost access to competent analysis.
Not perfect analysis. Not guaranteed wisdom. But a level of sustained intellectual support that currently requires a collection of professionals.
A personal agent could help manage:
Health
Work
Learning
Paperwork
Purchases
Travel
Home maintenance
Relationships and commitments
Government interactions
Financial administration
It remembers what happened, notices unresolved problems, and converts intentions into executed plans.
The interface changes from:
“Find information for me.”
to:
“Understand my situation and move it toward the outcome I want.”
That is a much larger change than better search.
Today, many projects are constrained because there are not enough capable people to:
Analyze the problem
Coordinate the work
Write the software
Read the literature
Prepare the documents
Check every edge case
Maintain institutional memory
If that constraint weakens, the remaining bottlenecks become more visible:
Energy
Factories
Laboratories
Housing
Infrastructure
Permits
Trust
Leadership
Capital
Physical resources
Political coordination
AI may reveal that many problems we call “knowledge problems” are actually governance or implementation problems.
A model can design excellent housing policy. It cannot, by intelligence alone, reconcile entrenched interests, approve construction, and build homes.
The central question will not be whether AI is smarter.
It will be whether people and institutions remain capable of:
Choosing goals
Detecting bad objectives
Demanding explanations
Verifying consequential claims
Rejecting attractive but unsafe outputs
Maintaining expertise through generations
Retaining authority when machine recommendations are usually better
The bleak outcome is not necessarily unemployment or machine rebellion.
It is a world in which:
Science progresses, but scientists cannot follow it.
Software works, but engineers cannot explain it.
Medicine improves, but clinicians cannot independently evaluate it.
Government optimizes, but citizens cannot contest it.
Companies grow, but leaders do not understand their operations.
Students produce brilliant work, but do not become brilliant people.
A 10-times-larger effective training run - like Astra, which is being released within the next 6 to 8 weeks - combined with strong tools and post-training, probably does not instantly produce a machine civilization.
It could produce something more immediately consequential:
A system that can become operationally expert in nearly any digital domain, construct its own representation of messy problems, run investigations, generate and verify solutions, and remain coherent over projects lasting days rather than minutes.
The question is no longer just what the model can discover.
It is whether human civilization can absorb discovery at machine speed without surrendering understanding, agency, and control.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.