AI has learned to write, code, search and operate software. These systems work inside defined tasks, with tests, supervision and mistakes that can usually be reversed.
The next step is different: decision-making under uncertainty. Systems that decide what matters, choose actions before all the evidence is available and accept consequences that cannot always be reversed.
Finance is one of the clearest tests of that shift. A financial world model must understand how companies, markets, capital flows, technology, policy and human behaviour interact—and turn that understanding into decisions that improve over time. Human investors have known limits, and fixed models decay. The bet is that a system that learns from every decision can eventually outperform both.
Ilya Sutskever has described major AI advances as choosing a different mountain to climb—one where reaching the top changes the paradigm. We see autonomous investing in those terms: not as a more capable coding agent, but as a different problem requiring judgment, risk-taking and continuous adaptation.
That is the problem Bill Sun has chosen to work on. We first met him at StudioAlpha JAM at Cooley in San Francisco in December 2025. The conversation continued, and StudioAlpha later invested in his work.
Bill combines two backgrounds that rarely sit in the same person: frontier AI research and responsibility for institutional capital.
Born and raised in China, Bill entered Peking University’s Advanced Applied Mathematics programme—limited to ten students—after competing in the Chinese Mathematics Olympiad and winning first prizes in physics and chemistry. He later earned a Stanford PhD in mathematics and AI.
Bill worked on early attention mechanisms for question answering at Google Brain, now part of Google DeepMind, before the Transformer paper. Attention later became a core building block of modern large language models. He then moved into quantitative and AI investing at Citadel and Point72 before managing a $1 billion US equity portfolio at Millennium.
He has worked on both sides of the problem: building intelligent systems and making high-stakes investment decisions in real markets. Links to Bill's X thread, research paper and website are at the end of this article.
Investing is a chain of difficult decisions. What matters? What can be ignored? Is the evidence strong enough? Should the firm act now, investigate further, wait, or walk away? If it acts, how much capital should it risk?
Today, these decisions are spread across analysts, quantitative researchers, traders, portfolio managers, risk teams and engineers.
Bill’s premise is that the whole investment team can be rebuilt as an AI system.
Traditional investment firms depend on human judgment. The best investors can combine company knowledge, market behaviour, industry structure and experience into one view. They can recognise when an old idea no longer fits the world.
But human judgment does not scale easily. One person has limited time, attention and memory. A team can cover more ground, but it also creates handovers, meetings, internal politics and gaps between specialists.
Nobel laureate Herbert Simon showed that decisions are always constrained by limited information and limited mental capacity.1 Daniel Kahneman, who received the 2002 Nobel Prize for work developed with Amos Tversky, showed that people also make predictable errors when judging uncertainty, probability and loss.2
These limits remain visible even in sophisticated investment firms. After a position is closed, the profit or loss is clear. The quality of the decision is often less clear. Did the team understand the situation, or get lucky? Was the forecast wrong? Was the position poorly chosen or too large?
Good investors try to learn from these questions. Much of that learning still remains inside people and leaves when they leave.
Traditional quantitative systems solved part of this problem by processing more data and applying rules consistently. But many rely on fixed models. They can keep following the same logic even when the market conditions behind it have changed.
Technology has improved research, data processing and execution. AI now helps existing investment teams work faster. The human team, however, still makes the important decisions.
Bill is working on a different model.
Bill’s answer is a different model. He is not merely to record that judgment. He is building an AI investment team that can make investment decisions and improve over time.
A self-improving investor still needs a precise record. Before acting, the system states what it believes, why it believes it and how confident it is. It records what would change the view and how the belief should affect the portfolio.
When the outcome arrives, the system compares reality with the original reasoning. It then identifies whether the forecast, the trade or the position size was wrong.
Bill reduces this process to four words:
Forecast. Act. Evaluate. Update.
The difficult part is making every step explicit, connected and measurable. Bill’s system separates the view of the world from the investment action, so a poor forecast can be distinguished from a good forecast expressed through the wrong trade.
Today, Bill remains the final portfolio manager. The long-term goal is for the system to perform the full investment process itself—including, eventually, replacing Bill as the final decision-maker.
The aim is not a better archive. It is an AI investment firm that can make decisions, learn from them and gradually take over the full investment process.
This chapter makes three connected points. Markets keep changing, so history alone is not enough. Investors must understand what causes an outcome, not only what happened. That requires a working model of the world that can hold several possible futures and update as evidence changes.
Historical data matters because markets do repeat some patterns. But those patterns do not operate under fixed conditions.
Interest rates change behaviour. Technology reshapes industries. Regulation changes incentives. Capital enters profitable strategies and weakens them. Other investors observe, copy and adapt.
The 2013 Nobel laureates in Economic Sciences—Eugene Fama, Lars Peter Hansen and Robert Shiller—showed that short-term prices are extremely hard to predict, while broader return patterns can emerge over longer periods.3
Andrew Lo’s Adaptive Markets Hypothesis adds that investors, institutions and strategies evolve through competition and changing conditions. A method can work in one environment and fail in another.4
A backtest therefore answers a limited question: how would this method have behaved in earlier markets? It does not prove that the same forces still govern the present one.
The investor must therefore do more than search for patterns. The investor must understand the conditions that made the pattern work and recognise when those conditions have changed.
That leads to the next problem. Knowing that two things moved together is not the same as knowing why.
The same result can come from very different causes.
Higher revenue may reflect lasting demand, temporary pricing power or a competitor’s failure. An interest-rate cut may signal that inflation is under control or that the economy is weakening. The visible event may be the same while the investment meaning is completely different.
That distinction matters because investing is not only about predicting what may happen. It is about understanding which forces could produce the outcome and what those forces imply for the next decision.
Judea Pearl’s work on causality formalised this difference. A correlation tells us that two things moved together. A causal model tries to explain what produced the change and what might happen if one of the conditions changed.5
For an investor, this changes the question.
It is not enough to ask whether a company’s revenue will rise. The investor must understand what could make it rise, how durable that force is and what evidence would show that the explanation is wrong.
Without this causal layer, a system may make the right prediction for the wrong reason. That can still produce a profitable trade once. It does not create a reliable investment method.
Bill’s system therefore needs more than forecasts. It needs a working model of the world behind them.
Bill’s answer is to build a model that represents the parts of the world that matter for the investment decision.
In plain English, a financial world model is a map of the forces that could affect an investment, the different futures they may create, and the evidence that would make one future more likely than another.
The system maps the important actors, incentives, supply constraints, capital flows and possible shocks. It does not produce one target and declare the future solved. It holds several possible paths, assigns a probability to each and watches for evidence that shifts those probabilities.
This is important because the future does not arrive as one clean scenario.
Demand may remain strong. A supply bottleneck may ease. Regulation may change. A competitor may introduce a better product. Each path leads to a different view of the company and a different investment decision.
The model does not need to reproduce the entire economy. It needs to capture the forces that are most relevant to the decision.
AI research provides a useful precedent. DeepMind’s MuZero learned a model that was useful for planning without being given the full rules of its environment. It learned enough about what mattered for rewards and actions to make better decisions.6
Markets are far more complex than games, so the comparison has limits. But the underlying idea is relevant: an intelligent system can build a practical model for deciding what to do without knowing every detail of the world.
Bill’s financial version must remain open to revision. The model is never final because the market does not hold its rules still.
Its value lies in helping the system ask better questions, compare possible futures and change its view when the evidence changes.
B O X
Weather is complex, uncertain and constantly changing. Forecasting improved because systems absorbed more observations, built better models, produced several possible futures and were scored every day against what actually happened.
AI has accelerated that process. ECMWF’s AI Forecasting System has been operational since February 2025. DeepMind’s GenCast produces 15-day probabilistic forecasts and outperformed ECMWF’s leading ensemble system on 97.4% of the targets in its retrospective evaluation.7
Markets are harder because people react to forecasts and change the environment. But weather forecasting proves the broader point: uncertain systems can become more predictable when better data, better models and constant evaluation work together.
A strong investment firm contains several different professions:
The statistician asks whether the evidence is real.
The fundamental analyst explains how a company earns money.
The industry specialist sees suppliers, competitors and bottlenecks.
The trader reads price behaviour, positioning and liquidity.
The portfolio manager decides how much capital to risk.
The engineer ensures that the data and backtests do not quietly cheat.
Human firms divide these abilities across departments. The final investment depends on whether the organisation can combine them. Information is lost between teams. Different departments use different assumptions. The synthesis often depends on one unusually capable portfolio manager.
Bill calls the alternative the autonomous investment firm: a team of specialised AI agents working from one shared view of the world. Each faculty has its own tools, data and standards, but all contribute to one shared view of the world. The trader’s observation can be tested by the statistician. The industry map can be challenged through simulation. The analyst’s confidence can be compared with its historical record.
This idea has become technically more credible because AI systems can now process broad bodies of information and divide complex work among specialist agents. The Transformer architecture made it possible to examine relationships across large amounts of information efficiently. Current systems go further. OpenAI’s GPT-5.6 can coordinate parallel agents, use tools programmatically and work through long professional tasks. Anthropic’s production research system uses a lead agent to divide a question among specialised agents, adapt the research plan and combine the results. Both companies also emphasise that coordination, evaluation and reliability remain difficult.8
The system must decide both what is worth learning and whether the evidence justifies action.
AI makes research cheaper. That does not make every piece of research useful.
Markets offer more information than any person or machine can investigate fully. The central task is deciding which uncertainty deserves attention. A system can produce an excellent answer to a question that has no effect on the portfolio.
What deserves more attention?
The same problem appears outside investing. Running Anthropic means deciding which research ideas deserve another 100,000 GPU-hours. Drug discovery means deciding which experiments justify years of work and millions of dollars. Every organisation with limited resources faces the same question: what deserves more attention? Investing is simply one of the cleanest environments in which to learn that skill.
Bill’s forecast engine begins here. It asks which answer could change the view, the action or the amount of capital at risk. It then directs research towards that question.
This is the same basic logic used in Bayesian experimental design. Before running an experiment, the researcher estimates which experiment is likely to provide the most useful information relative to its cost. In plain English: do the work that has the best chance of changing the decision.9
The most interesting question is not always the most valuable one. The value of a question is not how much information it produces. It is how much the answer could change the decision.
It is how much the answer could change the decision. A fact that adds colour but changes nothing has little investment value. A fact that determines whether the investment firm buys, sells, increases the position or abandons the thesis may justify significant time and cost.
The system should therefore record in advance what each possible answer would trigger. Research ends with an instruction, not merely a conclusion.
“Will demand remain strong?” is difficult to judge and unlikely to teach the investment firm much.
“Will the market’s estimate of spending by the largest cloud companies rise by more than eight per cent after the next earnings cycle?” has a deadline, a measurable result and a clear effect on the investment.
That precision creates feedback.
A badly designed investment machine always has an opinion and always wants to trade.
A serious investor has more choices. It can act, observe more closely, investigate a specific uncertainty, wait for better evidence, or kill the idea.
Waiting is not the absence of a decision. It is a decision that the available evidence does not yet justify the cost or risk of acting.
This matters because being right about a company and being paid for that view are different questions. A thesis may be correct long before the market recognises it. A share can be held while that recognition develops. An option with an expiry date may lose all its value before the thesis is proven.
The system must therefore judge both the truth of the thesis and the likely timing of recognition. When a clear event could force the market to react—an earnings report, product launch, regulatory decision or supply change—a time-limited position may make sense. Without such an event, patience may be the better action.
The ability to wait, or to abandon an idea entirely, is part of intelligence. A machine that must always act will eventually manufacture conviction where none is justified.
Investors tend to remember their old views differently after the outcome is known.
Confidence becomes less extreme. Warning signs appear more obvious. A lucky profit makes weak reasoning look sound. A reasonable loss can make a sound process look foolish.
Bill’s system tries to remove this convenient editing.
Every important belief carries three things: a receipt, a falsifier and a monitor.
The receipt records the evidence behind the belief. The falsifier states what would prove it wrong. The monitor watches for that evidence. Research is therefore tied to a dated claim rather than preserved only as persuasive prose.
The data must follow the same rule. A backtest may use only information that was genuinely available at the time. Later corrections must not overwrite the historical record. The system must be able to answer two different questions: what do we now know was true, and what did we believe then?
This is what Bill means by point-in-time honesty.
Probability forecasts can then be graded. The Brier score, originally developed for weather forecasting, compares a stated probability with what later happened. If a system repeatedly says that events have an eighty-per-cent chance, those events should occur around eighty per cent of the time. Proper scoring methods are designed to reward honest probabilities rather than exaggerated confidence.10
The ledger should record more than forecasts and trades. It should also record which questions the system chose to investigate, which ideas it rejected and when it decided to wait. Over time, the institution can judge whether the research changed anything, whether the rejected ideas deserved rejection and whether patience was better than action.
One score will never capture the whole system. Stanford’s HELM project argues that AI models should be judged across several dimensions rather than by one benchmark. Anthropic has reached the same conclusion in practice: evaluations are difficult, narrow tests can mislead, and human review still finds failures that automated tests miss. Anthropic’s model-written evaluation research also shows that AI can help create tests that expose behaviours overlooked by existing benchmarks.11
The same discipline belongs in investing.
A profit number is important. It is not a complete evaluation of how the profit was produced.
A good decision can lose money. A poor decision can make money.
This makes markets difficult teachers. The outcome is external and unavoidable, yet the explanation is noisy.
Bill’s system separates the result into three questions.
Was the forecast well calibrated?
Did the chosen trade express the forecast properly?
Did the return justify the risk?
Suppose the system correctly predicted that a company would report strong demand, but bought a call option that expired before the results arrived. The research may have been sound while the trade construction was poor.
Suppose it predicted the wrong result but made money because the whole sector rose. The investment was profitable, but the reasoning should not be reinforced.
Feedback must go to the component responsible for the error. Bill calls this decomposed feedback or directed credit assignment. In plain English: improve the part that was wrong, rather than allowing one final P&L number to teach the entire institution the wrong lesson.
Sometimes the error sits deeper than a forecast. The system may have asked the wrong question, misunderstood the industry or trusted a method that worked only under an old market regime. In that case, it must change its research framework.
Richard Sutton’s Bitter Lesson argues that general methods based on learning and search have repeatedly outperformed systems built around fixed human rules as computing power increased. Google DeepMind’s AlphaEvolve offers a more recent example: models propose changes to algorithms, automated evaluators test them, and the better versions are used to generate further improvements. It has produced improvements in mathematics, data-centre operations, chip design and AI training.12
Finance is harder because the quality of an investment idea cannot always be checked immediately or perfectly. Still, the principle applies. A self-improving investor must be able to test its own methods, keep the ones that work and replace the ones that no longer do.
These changes should happen at two speeds.
The fast process updates working knowledge: industry maps, trusted sources, checklists, questions and warning signals. The slower process changes the deeper models and research frameworks only after the same type of failure appears repeatedly.
A system that rewrites itself after every mistake will learn noise. A system that never rewrites itself becomes rigid. The difficult judgment is deciding which kind of change the evidence supports.
The final step is deciding how much capital to commit.
John Kelly’s work connected information, probability and long-term capital growth. The plain-English lesson is that a genuine edge is valuable only when the amount invested matches the strength of the evidence and the possible loss. Betting too little wastes a good insight. Betting too much can destroy the investor even when the underlying idea is usually right.13 14
Bill adapts this idea with practical limits for drawdowns, liquidity, market impact and uncertainty. The system must learn both what to believe and how strongly to act.
The first use of AI in finance is faster work. The deeper use is a different institution.
Research remains connected to the assumptions that produced it. Forecasts are preserved before the outcome. Confidence is graded. Data is time-correct. Specialist agents are evaluated by their record. Investment methods are revised when the world breaks them. Risk follows the quality of the belief.
Above the research system sits a judgment layer. Its job is not to produce another forecast. It decides which unknowns matter, whether more research is worth its cost, whether the institution should act, wait or walk away, and whether the possible loss threatens survival.
Survival is the limit that the other considerations cannot overrule. The organisation begins to learn in software.
Current AI systems are capable enough to make this architecture plausible, but not reliable enough to remove human control. OpenAI’s latest model release shows substantial progress in long-horizon work, parallel agents and the use of AI to accelerate AI research itself. It also describes extensive monitoring, red-teaming and layered safeguards. Anthropic’s experience shows that multi-agent systems can outperform single agents on open research tasks, while small errors can compound and production reliability remains hard.15
Autonomy therefore has to be earned one function at a time.
Bill Sun’s attempt is to build this institution in real markets. He describes three connected parts:
a system that decides what needs to be researched,
a model that represents possible futures,
and a capital-allocation layer that turns beliefs into positions.
Outcomes flow back into all three. His Complete Investor note extends the idea further: the institution must also challenge its own maps, verify its evidence and change the frameworks through which it understands the world.
Financial markets are a credible proving ground because they combine uncertainty, adaptive competitors, repeated decisions and real consequences. P&L is objective. The reasons behind it are not. Building a system that can separate the two is the work.
The natural early partners are long-term capital owners who can contribute more than an allocation. They can bring market experience, specialist access, governance, data and patience while the system earns greater responsibility.
The purpose is larger than automating an investment firm. Investing is ultimately a problem of allocating scarce resources under uncertainty. The same challenge appears in scientific research, drug discovery, corporate strategy and AI development itself. If a machine can learn to decide what deserves attention, what deserves capital and when it should simply wait, it may be learning one of the most general forms of judgment.
This article is our attempt to explain Bill Sun’s work alpha.dev in plain English.
For Bill’s own words, read his latest X thread:
→ From Worker to Decision Maker
A longer doctrine is currently in preparation.
Radiohead — “Everything in Its Right Place.” - Electronic, controlled and slightly unsettling. It feels like intelligence organising chaos—exactly the mood of the self-improving investor, without the obvious robot cliché. It also marked Radiohead’s move into a radically new working method.16
Tell your banker.
Fab 💵
LinkedIn | Insta | Twitter | Studio⍺
StudioAlpha Capital is a Delaware-structured pre-seed venture fund backing AI-native B2B software startups at day zero like Bill’s Alpha.dev. Legal counsel: Cooley LLP. Fund administration: AngelList.
Bill’s deck: explains the technical loop—research, model possible futures, invest, evaluate each part, then improve the component that failed.
The Complete Investor: adds the wider vision—one artificial mind containing the skills of a whole investment firm, with evidence, falsifiers, verification and the ability to rewrite its own playbooks.
Simon, Kahneman and Tversky: human judgment is constrained and predictably biased, especially under uncertainty and loss. (nobelprize.org)
Fama, Hansen, Shiller and Lo: markets absorb information quickly, longer-term patterns still exist, and strategies must adapt because markets and participants change. (nobelprize.org)
Pearl: observing that two things move together is different from understanding what caused the change. (Wikipedia)
MuZero: an agent can learn a simplified model that is useful for planning even when the full rules of the environment are unknown. (arXiv)
Transformer research: modern models can connect information across long bodies of text and process it efficiently. (arXiv)
GPT-5.6 and Anthropic’s research system: current systems can coordinate specialised agents and use tools across long tasks, but reliability and evaluation remain major engineering problems. (OpenAI)
Bayesian experimental design: investigate the question most likely to improve the decision, rather than studying everything that looks interesting. (arXiv)
Brier and proper scoring: probabilities can be graded, encouraging the system to state honest uncertainty. (Wikipedia)
HELM and Anthropic evaluations: one score is not enough; systems must be tested across several dimensions, and human review remains necessary. (arXiv)
The Bitter Lesson and AlphaEvolve: learning systems can outperform fixed rules, especially when new approaches can be tested automatically and improved repeatedly. (Wikipedia)
Kelly: an investment edge must be converted into an appropriate amount of risk; correct ideas can still fail through poor sizing. (Princeton University)
I deliberately excluded Bill’s stated performance figures from the manifesto. They should appear only after independent verification, in the investor materials rather than as evidence for the thesis.
For your learning: the text is much better written now. Still I cut the folloing:
That is useful, but small.
https://www.nobelprize.org/prizes/economic-sciences/1978/simon/lecture/
https://www.nobelprize.org/prizes/economic-sciences/2002/kahneman/lecture/
https://www.nobelprize.org/prizes/economic-sciences/2013/press-release/
https://en.wikipedia.org/wiki/Adaptive_market_hypothesis?utm_source=chatgpt.com
https://en.wikipedia.org/wiki/Judea_Pearl
https://arxiv.org/abs/1911.08265
https://arxiv.org/abs/2509.18994?utm_source=chatgpt.com
https://arxiv.org/abs/1706.03762?utm_source=chatgpt.com
https://arxiv.org/abs/1909.03861?utm_source=chatgpt.com
https://en.wikipedia.org/wiki/Brier_score?utm_source=chatgpt.com
https://arxiv.org/abs/2211.09110
https://en.wikipedia.org/wiki/Richard_S._Sutton
https://www.princeton.edu/~wbialek/rome/refs/kelly_56.pdf
https://www.caia.org/sites/default/files/AIAR_Q3_2016_05_KellyCapital.pdf
https://openai.com/index/gpt-5-6/
https://en.wikipedia.org/wiki/Everything_in_Its_Right_Place?utm_source=chatgpt.com
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.