RSS Amplifier

Deepnote's Substack · May 26, 2026

I/O drops Gemini Omni + 3.5 Flash, Karpathy joins napkin-profitable Anthropic, & everyone gets an FDE thanks to Wall Street

0
Sign in to vote or save

Data Deep Dives · Deepnote's Substack

Hi! You're receiving this newsletter because you're subscribed to Deepnote updates. Deepnote has been publishing bi-weekly AI news for the past year (check out prior editions here), and we are expanding this roundup to all of our subscribers, curated by our CEO.

Two weeks ago, I was at a dinner in Hayes Valley where someone argued that the cost of a token is a rounding error compared to the cost of caring about it. Then Peter Steinberger posted he’d burned $1.3M of OpenAI credits in 30 days, and the argument stopped being theoretical.

That’s the through-line: with the compute bottleneck relatively solved, LLM labs solve for distribution. Anthropic and OpenAI launched competing forward-deployment joint ventures the same day, both Wall Street-backed, with OpenAI guaranteeing PE LPs a 17.5% minimum return. Anthropic separately had a fortnight: $1.25B/month SpaceX deal for 220,000 GPUs, profitability on a $30B run rate, a $50B raise closing at $900B, Stainless acquired, Karpathy joining. This was somewhat soured by missing out on Pentagon deals and Microsoft quietly canceling internal Claude Code seats after the token bill broke the math.

Google I/O was rich with updates, but left investors unimpressed: 3.5 Flash now defaults Search at 1B+ users, Omni folds generative video into the reasoning loop, Antigravity is enterprise-ready. Elsewhere, Stripe now helps Agents checkout, Thinking Machines trains an interaction model on 200ms micro-turns so it hears you while it talks, Cerebras IPO’d at $5.5B and popped 108%, SAP put €1B into an 18-month-old German lab, and Exa raised $250M at $2.2B for agent-native search. Full breakdown below.

🏗️ The frontier bottleneck has moved from model capability to deployment capacity. Anthropic and OpenAI are launching parallel PE-backed joint ventures in the same week. Each targeting the same gap: not enough engineers who can integrate AI into real business operations. OpenAI sweetened its pitch to PE firms with a guaranteed minimum return of 17.5%

🌐 Google I/O confirmed Google is building an integrated AI stack, not a model portfolio. Gemini 3.5 Flash as Search default (1B+ monthly users), Gemini Omni for any-to-any generative creation, and Antigravity for agentic coding form a coherent vertical that is harder to replicate by assembling best-of-breed components from separate vendors.

🧫 Thinking Machines Lab’s interaction model is the clearest argument yet that turn-based design is a fundamental architectural constraint and that fixing it requires training a different kind of model from scratch.

🤖 Stripe solving agentic payments infrastructure: Link for Agents lets AI agents execute purchases within defined limits without ever handling raw financial credentials.

🧠 Karpathy joining Anthropic resets the research talent narrative. The most visible AI researcher of the last decade choosing Anthropic concentrates credibility at the lab in the run-up to its IPO.

Google’s I/O 2026 keynote was the most model-dense in the company’s history, centered on three releases with distinct strategic positions. Gemini 3.5 Flash launches as the new default for AI Mode Search (now surpassing 1 billion monthly users), outperforming Gemini 3.1 Pro on agentic coding benchmarks: Terminal-Bench 2.1 (76.2%), GDPval-AA (1,656 Elo), and MCP Atlas (83.6%). Gemini Omni (launching as Omni Flash) is the headline architectural bet: a unified generative world model that accepts text, images, audio, and video as input and produces physics-aware, conversationally editable video output, with full any-to-any output planned in subsequent releases; it is already live in the Gemini app, Google Flow, and YouTube Shorts with SynthID watermarking on every output frame. Google Antigravity, the agent-first development platform, now integrates with the Enterprise Agent Platform and rolls agentic coding into Workspace for organizational deployment at scale.

Key implications:

  • Gemini Omni is the first top-tier model to merge reasoning and generative media creation natively, a direct architectural challenge to the multi-model stacks competitors currently run

  • Gemini 3.5 Flash becoming the Search default at 1B+ monthly users means Google’s inference volume for a single model likely eclipses the entire usage of any competitor product

  • Antigravity signals Google’s answer to Claude Code and Codex: an agent platform embedded in existing enterprise Google Cloud deployments rather than a standalone tool

In the top-right quadrant of the Artificial Analysis index, 3.5 Flash delivers frontier-level intelligence at exceptional speed. (Source)

Google DeepMind released experimental demos showing a reimagined mouse pointer interface that allows users to direct Gemini using natural gestures, shorthand annotations, and voice alongside standard cursor input. The demos show real-time recipe navigation and document interaction using speech and motion gestures, without switching between a chat interface and the working application. This is early-stage research, not a shipping product, but it directly addresses the input-modality bottleneck that limits how fluidly people can direct AI agents through existing GUI surfaces.

Anthropic and OpenAI announced competing enterprise AI deployment ventures on the same day. It signals both labs have concluded the limiting factor for revenue growth is deployment capacity, not model capability.

Anthropic’s unnamed venture is capitalized at $1.5B with Blackstone, Hellman & Friedman, and Goldman Sachs as founding partners (each committing $300M, Goldman $150M), plus Apollo, General Atlantic, GIC, Leonard Green, and Sequoia. It will embed Anthropic engineers directly within portfolio companies.

OpenAI’s parallel entity (”The Deployment Company“) has 19 investors, including TPG, Bain Capital, Brookfield, and SoftBank, with access to 2,000+ mid-sized companies across PE portfolios, and is reportedly offering investors a guaranteed minimum return of 17.5%. The strategic logic is identical for both: private equity portfolios are the most efficient enterprise distribution channel available at scale, bypassing standard enterprise sales cycles entirely. The ventures represent a direct competitive threat to Big Three consulting firms, which currently charge substantially more for equivalent AI integration work.

Anthropic is paying SpaceX $1.25 billion per month through May 2029 for access to the Colossus and Colossus II clusters, covering 300MW and 220,000 NVIDIA GPUs. At full run rate, that’s $15 billion a year and roughly $45 billion over the contract’s three-year life. The immediate operational result: doubled rate limits for Claude Code Pro, Max, and Enterprise customers. The strategic one: Anthropic’s internal compute demand has outpaced what AWS and Google Cloud can physically build on the ground right now, and this deal gives the company independent capacity ahead of a likely IPO.

Anthropic’s SpaceX compute deal adds 300+ MW and 220,000+ NVIDIA GPUs, enabling immediate uplift to Claude Code and API usage limits. (Source)

The same week, Andrej Karpathy announced he’s joining Anthropic as a member of technical staff, framing it around returning to R&D at what he called an “especially formative” moment for frontier LLMs.

Most AI progress has gone into making models more autonomous. Thinking Machines Lab’s argument is that this has quietly created a different problem: humans are getting pushed out of the loop, not because the work doesn’t need them, but because the interface has no room for them.

Their response is an interaction model trained from scratch with a multi-stream, micro-turn design that processes 200ms input chunks while generating output concurrently. Unlike today’s models, which freeze perception while generating and are blind to what the user is doing mid-turn, this runs audio, video, and text continuously in parallel. Interruptions, interjections, and real-time steering are native to the model rather than bolted on. The demo is worth watching.

Turn-based models see an alternating token sequence. Time-aware interaction models see a continuous stream of micro-turns, so silence, overlap, and interruption remain part of the model’s context. (Source)

They’re now putting $100K grants (plus $25K in compute credits) behind the research infrastructure this approach still needs. Benchmarks for real-time multimodal interaction (none exist yet), safety for always-on systems, generative UI for agent outputs, and tools for steering agents mid-task.

The short version of Karpathy’s writeup: Agents started producing larger, more reliable chunks of work, and he found himself delegating whole tasks rather than writing lines of code. His framing for what’s happening: Software 1.0 was explicit code, 2.0 was neural networks trained on data, 3.0 is instructing LLMs through context, tools, and memory.

Boris Cherny made the practitioner version of the same point. He hasn’t written code himself in 2026, ships dozens of PRs a day from his phone, and says coding is effectively solved for the work he does. The key mechanism is loops, recurring tasks managed by Claude that handle code reviews, repair flaky CI tests, and analyze user feedback, running continuously even when his devices are offline. He also thinks Claude Code itself may be 100 lines of code a year from now, the tool compressing its own complexity as the models get better.

The Department of War’s CTO office announced agreements with SpaceX, OpenAI, Google, NVIDIA, Reflection, Microsoft, and AWS to deploy frontier AI capabilities on classified government networks. It’s the first time this many commercial AI providers have been formally authorized simultaneously for classified-tier infrastructure.

Notably absent is Anthropic, which the Pentagon blacklisted as a “supply chain risk” earlier this year after a contract dispute over whether the military’s use of Claude models would be subject to ethical guardrails.

Stripe launched Link for Agents: a wallet product that lets AI agents execute purchases on behalf of users, with payment credentials never exposed to the agent and user approval required per transaction, installed via a skill.md file at link.com/skill.md. This is the payment infrastructure layer that agentic commerce has been missing: a model where an agent can be authorized to spend within defined limits without being handed raw financial credentials.

Jiaqing Liang and 17 co-authors from the A3 Laboratory (Advantage AI Agent Lab) argue in the GenericAgent (GA) paper that long-horizon agent failures are primarily a context engineering problem: as interactions accumulate, tool schemas and memory retrievals progressively displace decision-relevant information, degrading reasoning even within a technically sufficient context window.

GA addresses this through four coordinated mechanisms:

  • a minimal atomic tool set

  • hierarchical on-demand memory

  • a self-evolution pipeline that compresses verified trajectories into reusable SOPs and executable code

  • context truncation layer that actively manages information density throughout execution

Crucially, the self-evolution mechanism converts prior task experience into compact structured procedures, so later runs begin from a denser, higher-signal context rather than raw logs. Across task completion, tool-use efficiency, memory effectiveness, and web browsing benchmarks, GA consistently outperforms leading agent frameworks while consuming significantly fewer tokens per task.

GA’s central design claim visualized: effective context is not a balance of three equal axes but a constrained optimization between completeness and conciseness. Left (verbose) examples preserve completeness at the cost of attention dilution; right (terse) examples improve conciseness at the cost of missing critical state. (Source)

Tsinghua/HIT researchers formalize a long-underappreciated problem: the orchestration layer wrapping an agent (its staging logic, failure taxonomy, role boundaries, artifact contracts, and stopping rules) determines performance as much as the base model. Nowadays, this “harness” is almost always buried in framework-specific controller code, making it untransferable, non-comparable, and scientifically opaque.

The Natural-Language Agent Harnesses (NLAHs), structured natural-language documents encoding harness policy, executed by a shared Intelligent Harness Runtime (IHR) that places an in-loop LLM to interpret harness logic against the current state and a runtime charter. Controlled evaluations across coding and computer-use benchmarks show that IHR-executed NLAHs match the task outcomes of code-coupled harness realizations while supporting clean module ablation (enabling practitioners to isolate the effect of individual control components (e.g., a verification gate or repair loop)). The practical implication is that harness design should be treated as a scientific and engineering discipline in its own right: teams should externalize their orchestration logic as versioned, inspectable artifacts rather than embedding it in runtime-specific scaffolding, enabling reproducibility, cross-system migration, and systematic ablation as agent complexity scales.

A canonical coding-agent harness expressed in controller code (left) versus as a Natural-Language Agent Harness executed by IHR (right). The NLAH version exposes the same control logic as an editable, portable artifact independent of any runtime. The IHR decomposes into an in-loop LLM interpreter, a backend providing tool and child-agent interfaces, and a runtime charter defining state and contract semantics. (Source)

Alibaba’s Qwen3.6-27B is the first major open-weight dense model to ship Multi-Token Prediction (MTP) heads as a first-class architectural feature rather than an add-on, enabling native speculative decoding where the model drafts multiple candidate tokens per forward pass and verifies them in parallel. The architecture combines a hybrid Gated DeltaNet (linear attention, 3 of every 4 sublayers) with standard Gated Attention layers using aggressive KV-head reduction (4 KV heads vs. 24 query heads), compressing KV cache memory significantly at serving time. On a consumer RTX 3090, enabling MTP via llama.cpp moves the same Qwen3.6-27B Q4_K_M from 38 to 65 tokens/sec.

This llmfan46 community variant preserves the native MTP checkpoint weights post-abliteration (safety-filter removal via representation engineering), allowing researchers to study or deploy the full MTP-accelerated stack locally. The broader implication: MTP is now a practical throughput lever for self-hosted models, and teams evaluating open-weight deployments should factor in native speculative decoding support as a first-order serving criterion alongside raw benchmark scores.

📊 CodexBar: free, open-source macOS menu bar app by Peter Steinberger (the OpenClaw/Clawdbot creator) that keeps AI coding-provider limits visible across 40+ providers (Codex, Claude Code, Cursor, Gemini, Copilot, and more) without requiring browser logins. It shows session and weekly limits with reset timers, uses OAuth/cookies/local CLI files to reuse existing sessions, and now ships a bundled codexbar CLI for scripts and CI.

🧪 superpowers-lab is a live staging ground for experimental Claude Code skills that haven’t made it into the main Superpowers framework yet: The current headline skill gives Claude Code direct access to a headless Windows VM via a thin tmux-based control layer (no GUI, no VNC), just scripted terminal interaction with interactive processes that normally resist automation. It installs as a single plugin line into claude.json and sits alongside the wider Superpowers ecosystem, which ships engineering culture (TDD, planning, verification gates) as a folder of markdown files that work identically across Claude Code, Cursor, Codex, Copilot CLI, and Gemini CLI.

Lucebox Hub rewrites LLM inference by hand, one consumer GPU at a time, and gets Apple Silicon efficiency numbers on a 2020 NVIDIA card. Two releases target the RTX 3090: a megakernel that fuses all 24 layers of Qwen3.5-0.8B into a single CUDA dispatch, hitting 1.87 tok/J, matching Apple M5 Max efficiency at 2× the throughput versus llama.cpp and a DFlash speculative decoder with DDTree for the 27B that reaches 207 tok/s (5.46× over autoregressive) while fitting 128K context in 24 GB VRAM via TQ3_0 KV cache quantization.

🎙️ GPT-Realtime-2 hits the OpenAI API and immediately gets demos with real-time transcription, low-latency voice, and session state. Within hours of release, developers were wiring it to voice chatbot UIs with live transcripts, Whisper-based transcription, and sub-second turn-taking. The demo in the linked tweet uses the marin voice with transcription-whisper and shows the full session handshake in real time.

📄 The Unreasonable Effectiveness of HTML: Thariq Shihipar from the Claude Code team argues Markdown is the wrong default output format for agent work, and HTML is better in almost every way. Markdown past ~100 lines stops being read by humans, which means plans and specs generated by agents effectively disappear from the review loop. HTML gives the same document color, diagrams, collapsible sections, and interactive elements, the kind of information density that makes a 500-line spec actually scannable. In Deepnote, the implications are direct: HTML artifacts produced by agents are shareable, renderable, and readable without a Markdown preview step.

🖥️ USB-C Headless Ghost Display Emulator: A $10 dummy plug that tricks a GPU into thinking a monitor is connected, which matters more than it sounds for agent-controlled machines. A quick hardware note: running Claude Code or any GUI-dependent agent on a headless server or mini PC without a display attached causes GPU drivers to disable hardware acceleration entirely, which tanks rendering performance and breaks screen-capture-based tools. This EDID dummy plug emulates a 4K@60Hz monitor over USB-C/Thunderbolt, keeping the display stack fully active with zero physical monitor required.

Running ~100 concurrent agents on every PR is what software development looks like when token cost approaches zero: Peter Steinberger’s breakdown of OpenClaw’s agentic infrastructure is one of the clearest pictures yet of what “AI-native” development actually means operationally. It is not a copilot that assists, but a mesh of specialized agents that review every commit for security regressions, deduplicate and cluster issues, auto-generate PRs from meeting discussions, and close six-month-old bugs when a relevant fix lands. The interesting design constraint isn’t capability, it’s that the whole system is designed to work lean: agents don’t replace headcount linearly, they compress it multiplicatively.

Steinberger’s CodexBar, which shows token spending on different AI coding tools. In 30 days, Steinberger had spent $1.3 million worth of tokens on OpenAI’s API. (Source)

GPT-Realtime-2 more than doubles its predecessor on instruction retention: Scale Labs’ Audio MultiChallenge S2S leaderboard places GPT-Realtime-2 at 48.45 (xHigh config) versus GPT-Realtime-1.5 at 34.73, with instruction retention rising from 36.7% to 70.8% APR. The ranking shows Gemini 3.1-flash-live-preview (Thinking) at 36.06, sitting just above the previous OpenAI model, which means voice model competition is compressing fast: last generation’s frontier is now mid-table. The real signal for practitioners is the instruction retention number: voice agents that can’t hold context across a conversation are toys; ones that hit 70%+ APR on a standardized benchmark are deployable infrastructure.

The Scale Labs Audio MultiChallenge S2S leaderboard chart shows the full competitive field and makes the gap between GPT-Realtime-2 configurations and everyone else immediately readable.

The chat interface is ending for power users, and the abstraction gap it leaves is the next product war: Peter Yang’s note captures a real bifurcation happening right now: practitioners have migrated to Claude Code and Codex for anything that requires actual work, while general chat interfaces remain the access point for everyone else. The unsolved problem is that the tools worth using require GitHub familiarity, worktrees, CLIs, and MCP setup, a barrier that realistically blocks the next billion users.

~$50B pre-IPO round (reported): Anthropic is in the final stages of closing its largest-ever raise at a ~$900B valuation, more than double its February 2026 valuation of $380B (driven by an annualized revenue run rate that crossed $30B + demand for compute to run Claude Mythos, its cybersecurity-focused model). The round, expected to be its last before an IPO as early as October 2026, would briefly make Anthropic the most valuable private AI company.

$13B project financing: Meta is arranging a $13B mostly-debt financing package via Morgan Stanley and JPMorgan for its El Paso, Texas data center campus, representing one of the largest single-site digital infrastructure financings on record.

$5.5B IPO, 108% first-day pop: Cerebras Systems priced at $185/share (above its $150–$160 range) on $510M in 2025 revenue, a $20B OpenAI contract, and a wafer-scale chip 57× larger than NVIDIA’s H100 that claims 32% lower inference cost per token than Blackwell; the stock closed above $300, opening 2026’s IPO season.

$3B fundraising target: Principal Asset Management is raising two private real estate equity funds ($2B US-focused, $1B European), targeting 18–20% net IRR over 8 years. This follows the firm’s February 2026 close of a $3.64B data center fund.

$2.1B Series B: Isomorphic Labs, Google DeepMind’s drug design spinout, raised $2.1B led by Thrive Capital with participation from Alphabet, GV, MGX, Temasink, CapitalG, and the UK Sovereign AI Fund to scale its IsoDDE AI drug design engine.

$1.16B investment over 4 years: SAP is acquiring German AI startup Prior Labs (an 18-month-old lab behind the TabPFN tabular foundation models) and investing €1B to build it into a sovereign European frontier AI lab focused on structured enterprise data.

$1B ARR milestone: Gusto reported $1B in actual trailing revenue (not ARR), putting the payroll and HR platform squarely in IPO territory.

$750M (reported, in talks): Ramp is in talks to raise $750M at a $40B+ pre-money valuation, just six months after its $32B November 2025 round. The 25%+ valuation step-up in half a year reflects how rapidly AI-powered financial automation is compressing the gap between fintech infrastructure and the accounting software incumbents it’s displacing.

$300M+ acquisition: Anthropic acquired Stainless, the SDK and MCP server generation platform that powered official developer libraries for OpenAI, Google, and Anthropic simultaneously, at a reported 2x+ premium to its $150M December 2024 Series A valuation.

$250M Series C: Exa, the web search API built for AI agents rather than humans, raised $250M at a $2.2B valuation led by a16z. The company already powers search for Cursor, Cognition, HubSpot, and over 400,000 developers, and the round reflects a growing bet that agent-native search infrastructure, optimized for machine consumption, not click-through.

$7M Seed: Altara raised $7M led by Greylock with Jeff Dean, and OpenAI and AMD leadership as angels, to build an AI intelligence layer that unifies fragmented experimental, sensor, and manufacturing data for semiconductor, battery, and advanced materials companies.

Acquisition (undisclosed): Palo Alto Networks is buying Portkey, an AI gateway that routes enterprise traffic between applications and model APIs processing trillions of tokens a month, and folding it into Prisma AIRS.

Q1 2026 revenue: $3.7M (+9,370% YoY): Quantum Computing Inc. (QUBT) posted a headline-grabbing revenue surge driven almost entirely by the acquisitions of Luminar Semiconductor and NuCrypt (organic revenue was $24K) while operating losses hit $20.6M on negative gross margins.

Acquisition (undisclosed, est. $10–20M): Check Point Software acquired Deepchecks, the AI testing, evaluation, and observability platform, as Check Point’s fourth Israeli acquisition of 2026, folding the team’s LLM evaluation expertise into its new Agentic Network Security Orchestration platform.

Acquisition (undisclosed): Medisolv acquired Health Elements AI to automate clinical data abstraction from medical records for its 1,800 healthcare organization customers.

No posts

Read the original on datadeepdives.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.