I tested writing an investment memo five different ways, from zero-shot prompting in Claude Opus 4.6 with extended thinking to a multi-step agent workflow in Google Opal. The agent won.
The multi-step workflow made a better analysis, had more data-driven insights, and a more opinionated take than the best chat interface on the market could generate in a single pass. Here’s why that matters and how to build your own with a template included.
Chat interfaces are doing too many things at once. When you ask Claude or Gemini to “write me an investment memo on Physical Intelligence,”1 the model is simultaneously figuring out how to approach it, deciding what to research, calling tools, and synthesizing a response. That’s a lot of inferred process in one pass.
It’s also probabilistic. Each run can use different tools, take different paths, and produce varying results. And it can be limiting to whatever tools that specific chat app offers. Google has strong multimodal but weaker connectors. Claude has great connectors but no multimodal output. ChatGPT is trying to do everything. They each have different baselines for intelligence, costs, and speed. More here.
Custom agents solve both problems. Defined the steps. Tool selection at each step. The workflow becomes more deterministic, more reliable, and, more importantly, better at each individual task because no single node is trying to do everything.
Think of the writing process. Nobody sits down and produces a polished memo from a blank page in one pass. First, start with a plan, then research, and then write. Agents work the same way.
This diagram, inspired by Zapier, shows the spectrum of full deterministic automations v. full inference led automations with AI agents.
Automating workflows: full determinism with each step called out
Agentic workflows: AI nodes embedded within a step by step process
AI Agent: AI agent determines each step and tool all to get outcome
So we’ll work to use the best of both: step by step support and embedding AI nodes in the process. This enables more control over the results and more reliability than either one alone. And if you’re new to Google Opal, you can read more here.
My best-performing workflow had three stages:
Plan (outline the memo structure and research directions),
Research (enrich each section with data),
Write (synthesize into a final draft).
Each stage handed off context to the next.
Why this works: the planning node can focus entirely on thinking through what a great memo needs. The research node doesn’t have to worry about structure. It focuses on data gathering and research. The writing node receives a rich brief instead of starting cold. It’s the same reason a chef preps ingredients before cooking.
I tested five approaches head to head.2 Zero-shot in Opus 4.6 extended thinking was my baseline. The multi-step agent — using Gemini 3.1 Pro with better prompts at each node — outperformed it across overall quality, specificity of copy, data-driven insights, and originality of take.
📝 Sample output memo (in footnotes)3
👩⚖️ Eval template with LLM as a judge.4
Prompting inside an agent is different from chatting. You’re writing instructions for a specific task within a larger workflow, not having a conversation.
Three things that make the biggest difference5:
Make information hierarchical.
Use markdown. Headers signal priority.# Rolematters more than## Guidelines. Number your lists. Nest details under sections. The model reads structure as meaning.Assign a specific role.
“You are a startup investor” gets generic output. “You are a General Partner at a $300M deep tech fund evaluating first-round deals, high tech risk tolerance, low market risk tolerance” unlocks a different part of the embedding space. The more specific the role, the more specific the output.Define what success looks like.
Either describe it (”Ask the hardest questions for each section to ensure a more holistic view”) or show it (include an example memo you like as a one-shot). The model can’t hit a target it can’t see.
For example, here’s the planning node that worked best for me
Here’s the planning node prompt that worked best for me:
# Role
You are a General Partner at Khosla Fund who plans thoughtful investment memos that are data driven and persuasive.
# Task
Write an outline plan for investment memo research for {{company name}}.Sections:
1. Executive Summary
2. Founding Team
3. Problem
4. Solution
5. Market Opportunity (TAM)
6. Business Model
7. Competition
8. Diligence Questions.
# Guidelines
1) Each section should include data-driven research directions.
2) Specify what to search for and why.
This can be updated with more specific context and direction to best fit what you want to use this template for.
This workflow is a starting point, like an initial recipe.
Some ideas for improvements:
Pitch deck as input. For memo writing, many teams will have a deck on hand to start the research process. You can use upload file as an input node to kickoff the research workflow.
Multi-modal. Add images, audio, and/or video as part of the
Explore more robust research and analysis. Specific nodes for specific parts of the memo, especially for richer data analysis and modeling, like revenue or market forecasting.
Specific outputs. We kept it text for ease of eval. You can explore Google Docs, Slides, or even custom HTML pages.
Citations. Getting the writing step to have citations.
Try out the memo template. Try out the evaluation LLM as a judge template.
And from there, start tinkering and changing things to make it your own.
If you have ideas, comments, or questions, add them below.
We treat the comments like NYT recipes comment where people share ideas and best practices.
If you enjoyed it, and thinks someone else would, please share.
Made with curiosity by 🌐 Worldwide Studios, a 501c3 CA non-profit hosting AI workshops and hacks.
📆 Check out the full calendar here to join for a future AI workshop.
Claude Opus 4.6 Extended Memo on Physical Intelligence
It’s interesting high level one pager, but it misses specific insights and thoughtfulness for starting to get into an investment memo.
INVESTMENT MEMO: Physical Intelligence (π) Khosla Ventures | March 2026 | Existing Portfolio Company
Thesis Physical Intelligence is building the foundation model layer for robotics — the equivalent of what GPT was for language, but for physical manipulation. The core bet: a single generalist AI brain that can control any robot to perform any task will become the default intelligence substrate for the $280B+ robotics market. We backed this team early and conviction has only increased.
Team — Best in Class The founding team is arguably the strongest assembly of robotics AI talent ever concentrated in one startup.
Karol Hausman (CEO) was a Staff Research Scientist at Google DeepMind and Stanford adjunct;
Sergey Levine (Chief Scientist) is a UC Berkeley professor whose lab pioneered deep RL for robotic manipulation;
Chelsea Finn runs Stanford’s robotic learning lab and is one of the most cited researchers in the field.
Lachy Groom — early Stripe employee, angel investor in Figma, Notion, and Ramp — brings rare operator-investor DNA as co-founder.
Rounding out: Brian Ichter (ex-Google Research, motion planning), Quan Vuong (RL),
Adnan Esmail (hardware engineering). The talent density in this building is extraordinary.
Product & Technical Moat Their flagship model, π0, is a 3B-parameter vision-language-action (VLA) transformer built on PaliGemma that directly outputs low-level motor commands via flow matching. Unlike competitors training narrow task-specific policies, π0 trains across multiple robots and tasks simultaneously — folding laundry, assembling boxes, peeling vegetables, making espresso — achieving cross-embodiment generalization.
Their latest π0.6 model introduced RECAP (RL with Experience & Corrections via Advantage-conditioned Policies), combining demonstration, real-time coaching, and autonomous practice. Results: 90%+ success rates, hours of uninterrupted operation, and doubled throughput vs. imitation learning alone. The MEM (Multi-Scale Embodied Memory) architecture enables tasks exceeding 10 minutes — a major frontier crossed. Key insight: they deliberately use cheap ~$3,500 arms, proving the intelligence compensates for hardware. This is the right abstraction layer.
Traction & Funding
$1.07B raised across 3 rounds; Series B of $600M (Nov 2025) led by CapitalG at $5.6B valuation
Investors: Khosla Ventures, Sequoia, Thrive Capital, Lux, Bond, Redpoint, Jeff Bezos, OpenAI
No disclosed revenue or commercialization timeline — Groom has explicitly told investors he won’t provide one
Active partnerships for real-world deployment (warehouses, homes, factories); partner ecosystem growing
Market & Competitive Landscape The global robotics market is projected to reach $280B by 2034, with general-purpose robots alone potentially a $370B opportunity by 2040 (McKinsey). Key competitors: Skild AI ($1.4B raised, $14B valuation, ~$30M early revenue, commercial-first approach), Figure AI (humanoid-focused), Google/Intrinsic (just folded into Google mainline with DeepMind integration), Tesla Optimus. The strategic fork in the market is research-first (Pi) vs. revenue-first (Skild). We believe Pi’s approach — build the best general model before commercializing — mirrors the LLM playbook that produced the most durable companies.
Key Risks
No revenue path articulated — requires sustained investor patience in a tightening market
Compute costs — Groom says there’s “no upper limit” on compute spend for this problem
Google/DeepMind — Intrinsic’s integration into Google mainline with Gemini Robotics is a serious platform threat
Skild’s data flywheel — commercial deployment generates real-world data that pure research can’t match
Sim-to-real gap — physical world failures carry real consequences; 90% isn’t 99.99%
Recommendation Continue to support. This is a generational company attempt. The team is the best in the world at this problem. The research-first approach carries real risk but also the highest ceiling. The foundation model paradigm has proven correct in every modality it’s been applied to — language, vision, code — and robotics is next. Physical Intelligence is the clearest leader in building that layer. Monitor commercialization signals closely in 2026; if π0.6+ begins generating partner revenue, this validates at multiples of current valuation.
Prepared for internal discussion. Not for external distribution.
Sample output memo from Google Opal multi-step workflow.
What I liked about this memo:
Bottom up TAM analysis versus prior memo citing general industry reports: The “TAM: Assuming a conservative base of 3.5M existing industrial robots and an emerging fleet of 100K humanoids by 2030, a software licensing fee of $10,000/node/year yields an immediate $36B Serviceable Obtainable Market (SOM). Because PI is hardware-agnostic, their TAM is a derivative of all kinetic hardware.”
Product specific insights: “Across tested tasks, $\pi_0$ delivers an ~80% success rate. The nearest open-source baselines—Diffusion Policy (~35%), ACT (~45%), and OpenVLA (<10%)—are not even in the same structural tier.”
Diligence Directives: data velocity and costs, compute economics, and commercial conversion are three areas to follow up on when talking to experts.
Memo output from Google Opal AI Agent:
Investment Memo: Physical Intelligence Date: March 10, 2026 To: Investment Committee, Khosla Ventures From: General Partner
Subject: Series C Investment Recommendation – Physical Intelligence (PI)
The Thesis Physical Intelligence (PI) is building the foundational operating system for the physical world. By shifting the robotics paradigm from bespoke, single-task programming to general-purpose Vision-Language-Action (VLA) foundation models, PI is unlocking the holy grail of automation: zero-shot generalization.
At a $5.6B post-money valuation, PI is not cheap, but venture returns at our scale are driven by category monopolies, not bargains. Foundational LLMs and vertically integrated humanoid competitors (Figure AI at $39B, Skild AI at $15B) have fundamentally re-anchored the exit parameters. To underwrite our standard 10x+ MOIC, we need a credible path to a $50B–$100B+ outcome. PI has the technical moat, the elite talent density, and the hardware-agnostic wedge to capture this entirely. We must lean in.
The Team: A Monopoly on Elite Talent We back technically elite founders who can bend reality. This roster represents arguably the highest density of AI and robotics talent globally.
The Innovators: Karol Hausman (CEO, ex-DeepMind), Sergey Levine (Chief Scientist, UC Berkeley), Chelsea Finn (Research Lead, Stanford), and Brian Ichter (VP Engineering, ex-DeepMind) are the pioneers of deep reinforcement learning and meta-learning. Their collective citation velocities dominate cross-embodiment learning.
The Operator: Lachy Groom (COO, ex-Stripe) provides the hyper-growth operational and capital-allocation engine required to deploy massive compute resources.
The Moat: ****PI’s engineering org bypasses traditional recruiters; they are a direct vacuum for top-tier talent from DeepMind, Stanford, and Berkeley. Incumbents cannot replicate this brain trust.
The Market & The Problem Current industrial automation is strangled by Moravec’s Paradox. High-level reasoning is cheap; sensorimotor manipulation is brutally expensive, requiring up to $10^{18}$ FLOPS to automate reliably.
The Friction:
Legacy automation is entirely rigid. Hidden integration costs (safety, programming, spatial mapping) routinely run 4x to 6x the cost of the hardware itself.
The Macro Tailwind: The U.S. logistics and manufacturing sectors are facing structural, escalating labor shortages. Enterprise desperation for adaptable robotic labor has never been higher. Wage inflation provides a massive, immediate pricing umbrella for RaaS (Robotics-as-a-Software).
The TAM: Assuming a conservative base of 3.5M existing industrial robots and an emerging fleet of 100K humanoids by 2030, a software licensing fee of $10,000/node/year yields an immediate $36B Serviceable Obtainable Market (SOM). Because PI is hardware-agnostic, their TAM is a derivative of all kinetic hardware.
The Solution & Technical Supremacy PI is rendering legacy, single-purpose robotic coding obsolete via their $\pi_0$ (pi-zero) and $\pi_0.5$ models. By utilizing flow matching for continuous action outputs, they have achieved unprecedented step-function improvements in robotic dexterity and adaptability.
Shattering Benchmarks:
Across tested tasks, $\pi_0$ delivers an ~80% success rate. The nearest open-source baselines—Diffusion Policy (~35%), ACT (~45%), and OpenVLA (<10%)—are not even in the same structural tier.
Unmatched Dexterity: On complex tasks requiring high manual dexterity (e.g., stacking multi-sized bowls), $\pi_0$ achieved near 100% success against Diffusion Policy’s 55%.
Zero-Shot “In the Wild”: Using the $\pi_0$-FAST tokenizer for autoregressive discretization, PI achieved 42.3% average progress on over 300 completely unseen real-world kitchen tasks. This resilience to background human activity and variable environments proves their architecture scales.
Competition & The Strategic Wedge The embodied AI space is fiercely capitalized, but fragmented.
Vertically Integrated Players: Figure AI ($39B post-money) and Tesla Optimus are building massive data moats, but their software is heavily bound to their proprietary chassis.
Pure-Play Competitors: Skild AI recently closed a $1.4B Series C at a $14B–$15B valuation. Covariant has stagnated.
The Advantage: PI’s $5.6B valuation presents a highly attractive arbitrage opportunity compared to Skild’s $15B and Figure’s $39B markers. More importantly, PI’s hardware-agnostic approach allows them to become the default “brain” for the long tail of hardware manufacturers who cannot afford to train their own $1B foundation models. This shifts enterprise CapEx to highly scalable software OpEx.
Diligence Directives To validate our 10-year IRR and MOIC sensitivity models, we must immediately extract the following from the data room:
Data Velocity & Cost: What is the current run-rate for proprietary trajectory data collection (TB/month), and what is the fully-loaded cost per TB? Integration Friction: What is the average time (in days) and engineering hours required to integrate $\pi_0.5$ into a completely new, unseen hardware form factor from a cold start?
Compute Economics: What is the exact ratio of training compute cost to edge-inference compute cost? Specifically, how does the 4-5x higher inference cost of the FAST tokenizer impact hardware latency (ms) and success rates at the 99th percentile?
Commercial Conversion: Provide the conversion rates from unpaid/paid enterprise pilots to fleet-wide rollout contracts over the trailing 12 months. Correlate any churn directly to Mean Time Between Failures (MTBF). Conclusion
Physical Intelligence is building the TCP/IP of the physical economy. The technical risk is steep, but it is entirely held by the most capable team on the planet. The market risk is zero. I strongly recommend we lead the next tranche to secure maximum ownership.
Google Opal Template: LLM as a Judge
With AI enabled workflows enabling more output, it’s important to have systems in place to check the quality and accuracy of the work. Here’s the LLM as a judge template used to compare each Opal output to the Claude Opus memo for Physical Intelligence.
Here’s the prompt
# Task
Evaluate these investment memos
x/10 - Overall Memo
x/10 - Copy (is it specific, direct, easy to read)
x/10 - Data (data driven insights in each section)
x/10 - Unique take (interesting opinion)
# Memos
1. Claude Opus 4.6 Prompt (paste in)
New memo (add in)
We’re working on an AI tutor to help get feedback on prompts for agents and in platforms
We built an AI tutor that will help you dial in your prompts for each step of the workflow. It loads specific challenges, like for this Google Opal flow, and gives feedback at each step.
This is currently in private beta testing.
If you want to try the prompt tutor, you can explore an upcoming workshop with live testing.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.