RSS Amplifier

Deepnote's Substack · Jul 3, 2026

Watermelon Sugar, Jalapeños, and Fable-ous return

0
Sign in to vote or save

Data Deep Dives · Deepnote's Substack

Sonnet 5 landed as the new default: Opus-level agentic vibes at a lower sticker price, yet somehow pricier per task in independent tests, slower than 4.6 under all the extra thinking, and roasted on Reddit and HN within hours as chattier, more adversarial, and a flat downgrade for a lot of people. Zhipu answered with GLM-5.2, fully open under MIT and basically matching Opus on agentic coding at a fifth of the cost. OpenAI and Broadcom taped out Jalapeño in nine months, Midjourney pivoted from catgirls to a 60-second full-body ultrasonic scanner plus a Spa opening in SF in 2027, and OpenAI floated handing the US government a 5% stake in the roughly $850B company to cool the regulatory drama.

The CEOs of Anthropic, DeepMind, OpenAI, and Mistral sat down for what ended up being a 2.5 hour lunch (no surprises there, as a Frenchman was in attendance). Welcome to the Bay Area, where the models are getting “safer” whether anyone asked, and a sandwich chain with Danny DeVito as a spokesperson mentions “AI” 22 times in its IPO filing.

🤖 Claude Sonnet 5 is now the default model for Free and Pro users, closing most of the performance gap with Opus 4.8 while costing well under half as much.

🔓 Fable 5 is back for all users as of today, after a two-and-a-half-week export-control freeze triggered by a jailbreak exploit that Anthropic’s own testing found wasn’t unique to Fable.

🐉 Z.ai GLM-5.2 released as a fully open, MIT-licensed model that almost matches Opus 4.8 at agentic coding and roughly a fifth of the price.

💰 AI funding stayed frenetic, with Baseten, Groq, General Intuition, and Upscale AI each closing raises of $190 million or more inside the same two weeks.

🌶️ OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom inference chip, with rival Anthropic already in early talks with Samsung on a chip of its own.

🍉 Meta is pouring 10x more compute into its upcoming Watermelon to catch up to GPT-5.5, while simultaneously launching a new cloud business to sell its excess AI capacity in xAI fashion.

Anthropic released Claude Sonnet 5, calling it its most agentic Sonnet model yet, with reasoning, tool use, and coding performance that closes much of the gap with Opus 4.8 while staying priced well below it. At higher “effort” settings, Sonnet 5 closes most of the gap with Opus 4.8 on tasks like the BrowseComp agentic search benchmark and OSWorld-Verified computer use, while offering a much wider range of cost-performance tradeoffs than its predecessor, Sonnet 4.6. It launches with introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026, rising to $3/$15 after that, versus $5/$25 for Opus 4.8, and it’s now the default model for Free and Pro users while remaining available on Max, Team, Enterprise, Claude Code, and the API.

On safety, Sonnet 5 shows a lower rate of misaligned behavior than Sonnet 4.6, though still higher than Opus 4.8 and Mythos Preview, and substantially weaker cyberattack capability, never producing a full working exploit in Anthropic’s Firefox vulnerability testing.

Claude Sonnet 5 benchmark scores versus Sonnet 4.6 and Opus 4.8. (Source)

Z.ai (formerly Zhipu AI) released GLM-5.2, a 744-billion-parameter Mixture-of-Experts model with 40 billion active parameters and a context window quadrupled to 1 million tokens, and gave away the weights under an unrestricted MIT license. On the Intelligence Index v4.1 it scores 51, ahead of Gemini 3.1 Pro Preview (46) and Gemini 3.5 Flash (50), and it lands within a percentage point of Anthropic’s Opus 4.8 on a key agentic coding benchmark at roughly a fifth of the price, about $1.40 per million input tokens and $4.40 per million output tokens versus $5/$30 for GPT-5.5 and $5/$25 for Claude Opus.

The timing is pointed: GLM-5.2 went live to paying customers on June 13, one day after a US export-control order forced Anthropic to pull Fable and Mythos globally. Z.ai‘s Hong Kong-listed shares jumped more than 30% on the news and are up over 800% since the company’s January debut.

Agentic coding performance by effort level. (Source)

Meta’s next major proprietary model, internally codenamed Watermelon, is currently in training and has reportedly caught up to OpenAI’s GPT-5.5 on key benchmarks, according to Superintelligence chief Alexandr Wang in a recent internal town hall. The model uses roughly 10x as much compute as its predecessor (Avocado / Muse Spark, released in April 2026) and represents Meta’s continued shift toward closed, high-compute frontier models.

This comes just days after reports that Meta is building out Meta Compute, a new cloud business to sell excess AI capacity and hosted models to external customers — the same strategy xAI/SpaceX has been aggressively pursuing with its Colossus clusters. In other words, while Meta pours unprecedented compute into Watermelon, it’s simultaneously trying to turn its massive infrastructure spend into a revenue stream by renting out the spare GPUs. Could it be because Meta is falling behind similar to xAI in its training ambitions?

The saga over Washington’s export restrictions on frontier models reached a resolution this week, though it exposed how ad hoc US oversight of AI releases has become.

Anthropic’s Fable 5 and Mythos 5 were pulled from public access on June 12 after the Commerce Department applied export controls following an Amazon-reported bypass that let Fable 5 walk through exploiting a software vulnerability. Mythos 5 was restored to approved organizations under an earlier partial exemption on June 26, ahead of the full export-control lift on June 30, with Fable 5 back for all global users July 1.

Anthropic’s own testing found the bypass wasn’t a unique Mythos-level risk, since weaker models including Opus 4.8, GPT-5.5, and Kimi K2.7 could reproduce the same behavior. Its new safety classifier now blocks the specific technique in over 99% of cases.

Anthropic, Amazon, Microsoft, Google, and other Project Glasswing partners are jointly drafting an industry framework for scoring jailbreak severity, a CVSS-style standard rating capability gain, breadth, ease of weaponization, and discoverability, meant to give labs and government a shared basis for triaging future bypass reports.

Separately, the White House pushed OpenAI to stagger release of its new GPT-5.6 models (Sol, Terra, Luna) to approved partners first. OpenAI’s June 26 system card rates all three as “High capability” for cybersecurity and bio/chem risk under its Preparedness Framework, backed by over 700,000 A100e GPU hours of red-teaming, its most intensive safety testing yet.

Anthropic’s framework for classifying jailbreak severity against its cybersecurity safety classifiers. (Source)

The redeployed model is already drawing skepticism; a post on r/ClaudeCode claims BridgeBench reruns show the returned Fable 5 scoring far below its pre-ban version (debugging dropping from 86.2 to 25.9), with the new guardrails allegedly triggering on benign tasks and falling back to Opus 4.8.

OpenAI and Broadcom unveiled Jalapeño, described as OpenAI’s first “Intelligence Processor,” an accelerator built from scratch for LLM inference rather than adapted from general-purpose AI chips. The two companies took it from design to tape-out in nine months, which they call the fastest ASIC development cycle ever achieved in advanced semiconductors, using OpenAI’s own models to help accelerate the design process. Engineering samples are already running production workloads including GPT-5.3-Codex-Spark, and OpenAI says early testing shows performance per watt substantially ahead of current state-of-the-art hardware, though full benchmarks haven’t been published yet. Deployment is planned at gigawatt scale with Microsoft and other data center partners starting later in 2026, with Celestica handling board and rack integration and Broadcom’s Tomahawk silicon handling networking.

OpenAI won’t have the custom-silicon lane to itself, though: a day after the announcement, Anthropic unveiled plans to partner with Samsung to manufacture its own custom AI chip as it looks to diversify beyond Google, Amazon, and Nvidia hardware.

Google spent the back half of June rebuilding its entire Gemini developer stack, and quietly revealed in the same window that the underlying compute is tighter than it looks.

  • A new default API. The Interactions API, in beta since December 2025, reached general availability as Google’s primary interface for Gemini models and agents. It adds Managed Agents that spin up a remote Linux sandbox from a single call, background asynchronous execution, tool calls that mix Search and Maps grounding with custom functions, and a Flex tier that cuts cost by 50%. Google says frontier agentic features will increasingly land here first, ahead of the legacy generateContent API.

  • Cheaper, faster generative media. Nano Banana 2 Lite, the fastest and cheapest image model in the family, generates a text-to-image output in about 4 seconds for $0.034 per 1K-resolution image, while Gemini Omni Flash opened to developers for video generation and multi-turn conversational editing at $0.10 per second of output, matching Veo 3.1 Fast.

    Performance benchmarks for Nano Banana 2 and 2 Lite compared to competitor AI image models, evaluating trade-offs between generation/editing quality (Elo scores), processing latency and cost per 1K-resolution image. (Source)
  • A capacity crunch behind the curtain. Per an FT report, Google told Meta around March it couldn’t supply the full Gemini API volume Meta wanted to buy, delaying some of Meta’s internal AI projects (Gemini powers Meta’s content moderation, ad chatbots, and customer service tools), and Meta has since told staff to conserve tokens. Other customers felt a lighter version of the same squeeze.

Midjourney, until now known purely for image generation, announced Midjourney Medical, a full-body ultrasonic imaging system built around a ring of roughly half a million tiny ultrasonic transducers that aims to produce MRI-comparable 3D body maps in about 60 seconds, which the company says is nearly 100 times faster than current MRI.

It’s pairing the scanner with the Midjourney Spa, a wellness venue where scanning happens incidentally during a normal spa visit, with the first location planned for San Francisco in 2027. The company, which says it has no investors and is funded entirely by its community, laid out a roadmap toward 50,000 scanners worldwide by 2031 with capacity for a billion scans a month, starting with body-composition mapping and adding FDA-cleared diagnostic capability over time.

Each slice continuously crossfades between the raw reconstruction and its AI segmentation — what the scan lets us identify inside the body. (Source)

Naveen Rao’s startup Unconventional AI released Un-0, an image-generation model built on an oscillator-based computing architecture instead of conventional chips, currently running as a software simulation but producing output comparable to Stable Diffusion or GPT Image 1. Rao, previously Databricks’ AI chief, says the architecture could eventually cut inference power use by as much as 1,000 times, with the roughly 50-person company now planning to release actual chip schematics and build a full inference stack around the technology.

NVIDIA released AutoModel as a PyTorch DTensor-native SPMD open-source training library targeting LLMs, VLMs, diffusion models, and retrieval models within the NeMo Framework. The core design bet is YAML-driven “recipes” with CLI-level overrides as the only interface, so the same configuration file runs on a single GPU or a multi-GPU, multi-node cluster without code changes. It integrates with Hugging Face for day-0 model support and deploys across local machines, SLURM, DGXC Lepton, and Kubernetes from a single run command. For teams maintaining separate experimental and production training stacks, AutoModel’s design collapses that boundary into one artifact, which reduces the translation cost between research iteration and deployment significantly.

DeepReinforce’s Ornith-1.0 introduces a self-improving RL loop in which the model jointly learns to solve coding tasks and generate the orchestration harnesses that guide its own problem-solving, rather than relying on a fixed, human-designed scaffold shared across task categories. At each RL step, the model first proposes a refined scaffold conditioned on the task and prior scaffold, then generates a solution rollout conditioned on that scaffold; reward from the rollout propagates back to both stages, so the model is optimized to author the orchestration that produces the best answers.

The 397B MoE flagship scores 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified, surpassing Claude Opus 4.7’s 70.3 and 80.8 on those same benchmarks, while the 9B dense model reaches 69.4 on SWE-Bench Verified, matching Gemma 4-31B at a third of the parameter count. Reward hacking is addressed through three layers: immutable environment and tool surface boundaries, a deterministic monitor that zero-rewards any attempt to read withheld test files or modify verification scripts, and a frozen LLM judge veto on top of the primary verifier.

Ornith-1.0-397B benchmark results across agentic coding evaluations. The model matches or exceeds Claude Opus 4.7 on Terminal-Bench 2.1 (77.5 vs. 70.3) and SWE-Bench Verified (82.4 vs. 80.8) while remaining fully open-source. (Source)

🧩 A four-layer pattern for surviving Claude outages treats a 529 error as a workflow bug rather than a vendor problem. Most AI coding setups keep all their state inside the chat session, so when Claude degrades, nobody can tell what the agent already touched, which tests passed, or whether resuming is even safe. The proposed fix: slice tasks small enough to checkpoint, write a receipt (commands run, files changed, tests passed) after every step, decide in advance which tasks may switch models and which must pause, and keep real state in repo-local files instead of chat history.

📊 Wix’s 250-run agent eval found that fixing your docs beats writing a skill, most of the time. The team tested the assumption that hand-curated “skills” outperform raw documentation by running identical coding tasks against baseline docs, optimized docs, and purpose-built skills, three times each, across 250 runs. Targeted doc fixes alone pushed CLI task completion from 67% to 87%, while skills-only runs trailed docs-optimized ones by seven points and lost their entire speed advantage the moment a skill had a stale field name or a mismatched scaffold.

Resetting your own Claude clock turns a rate limit into something you schedule around. Claude’s 5-hour usage window starts at your first message of the day, so a casual 7am question can leave you with no quota left by the time real work starts in the afternoon. The fix making rounds on r/ClaudeAI: fire a throwaway Haiku prompt on a timer early each morning, anchoring the reset to a time that suits you instead of whatever you happened to type first.

GLM-5.2 ties Opus-4.7 at pass@3 but loses by 6 points at pass@1, and the “2x more tokens” headline is misleading: Snowflake’s Coco team ran 103 dbt tasks against both models. GLM averages 99 turns and 860M billing tokens versus Opus’s 80 turns and 439M. That gap is driven by a small set of tasks where GLM spirals into excessive verification, not a consistent per-task pattern, on tasks both models solve, GLM only uses about 17% more calls. GLM’s real edge is dual-platform validation, it more reliably checks both DuckDB and Snowflake targets, which explains several of its wins despite the lower pass@1.

Aravind Srinivas says the model is no longer the product, and memory is the real bottleneck: Perplexity’s CEO discussed that chasing model quality or billion-user scale is the wrong game, the company he runs hit $20B in valuation and 45M users in three years with 400 people, and he thinks memory-chip makers like Micron could outvalue Meta within a year. It’s part fundraising pitch, part genuine thesis shift: compute and memory capacity, not benchmark wins, are what he’s betting determines who stays relevant in the phase of AI. Perplexity started as a search engine, shifted focus toward the hardware/compute layer, and is now positioning itself as a full-fledged AI lab, so Srinivas isn’t just describing the market; he’s describing where his own company has been moving.

Daytona leaving e2b as the last serious open-source sandbox standing: Daytona going proprietary is a signal for the whole agent-sandbox category, which is quietly following the same drift as the rest of AI infrastructure. He called out e2b’s founder specifically for staying open source “in the age of AI attackers,” framing it as a deliberate, harder choice rather than a default. Worth watching whether other infra providers in this space follow Daytona’s move or hold the line.

Coinbase kept AI spend nearly flat while token usage kept climbing, by changing defaults instead of adding friction: Brian Armstrong laid out the playbook: default engineers to cheaper open-weight models (GLM 5.2, Kimi 2.7) through an internal LLM gateway instead of capping usage, route prompts to the right model for planning versus execution, and push cache hit rates up, one internal tool went from 5% to 60%. Usage caps were barely binding anyway, 91% of employees never hit theirs, so Coinbase made spend visible rather than restrictive.

The chart he posted shows AI spend bars flattening while the token-usage line keeps climbing, which is the whole pitch in one image. (Source)

Chinese labs aren’t out-innovating Codex and Claude Code, they’re harvesting their outputs as training data: Burkov describes a fully scripted pipeline: have an LLM inject a subtle bug into a codebase, let Claude Code or Codex fix it, log every input and output, then use that transcript for supervised fine-tuning and the pass/fail result for reinforcement learning, no human in the loop. Tiezhen Wang, former Hugging Face APAC lead, makes the broader case in Rest of World: he calls distillation a neutral practice everyone in the field does, pointing to Musk’s own admission that xAI distilled OpenAI, and argues China’s open-source-by-default strategy wins on adoption speed precisely because cheap tokens let companies go “AI-native” internally, while Uber reportedly burned a year’s token budget in four months.

An AI paper reviewer just beat nine comparison systems in 90% of head-to-head matchups: Team at Refine ran 1,349 head-to-head matches across 150 economics preprints and won 90.4% of them, tying 5% and losing only 4.6%. That’s a lopsided result for a task long assumed to need domain judgment and taste rather than pattern-matching against prior literature. If review quality at that scale holds up under outside scrutiny, it’s an uncomfortable data point for journals and conferences about what peer review is actually rewarding.

Refine wins 90.4% of matches. (Source)

Cursor is building a GitHub competitor, and its own allies are joking the name collides with git’s most basic command: At its inaugural Compile event, Cursor unveiled Origin, an “agent-native” code hosting platform aimed squarely at GitHub for a world where AI agents write most commits. GitLab and Zed are pursuing similar rebuilds of version control for agent-first workflows, and Cursor’s fresh capital gives it a real shot at contesting GitHub’s incumbency.

The White House reportedly wants Anthropic to make Claude jailbreak-proof, security researchers say that bar doesn’t exist: Wired reports that Trump administration officials told the outlet any rerelease of “Fable 5” would need guardrails that can’t be circumvented, a requirement security experts quoted in the piece say no current alignment technique can satisfy. It’s a clean example of policymakers setting an engineering spec the field hasn’t solved for any model, from any lab, not just Anthropic’s.

$2.5 billion Fund VII: Valor Equity Partners, the growth-stage firm behind Musk-adjacent bets like SpaceX and Anduril, is targeting at least $2.5B for its latest fund, with a portion earmarked for further SpaceX investment.

$1.5 billion round (reported): Baseten is close to a $1.5B raise at a $13B valuation, co-led by Spark Capital, Sands Capital, Altimeter, and Wellington, a 160% valuation jump just five months after its last $300M round at $5B. The “inference gold rush” is now producing serial mega-rounds on a sub-annual cadence, a pace of re-pricing rarely seen outside of frontier labs themselves.

$650 million raise: Groq confirmed a $650M raise and restaffed its C-suite six months after Nvidia’s non-exclusive IP licensing deal poached its founder and CEO, pivoting the company toward its neocloud inference business. Investors are betting that inference infrastructure demand is strong enough to outlast even the loss of a company’s founding team and core IP exclusivity.

$320 million Fund VII: Seedcamp, the 18-year-old European early-stage investor, raised $320M, split between a $220M early-stage vehicle and a new $100M growth-stage fund to build out a US presence.

$320 million Series A at a $2.3 billion valuation: General Intuition, which trains AI agents on spatial-temporal reasoning using billions of first-person gaming clips, closed its round led by Khosla Ventures with General Catalyst, Jeff Bezos, Eric Schmidt, and Nico Rosberg, up from the ~$300M/$2B terms first reported and just eight months after a $134M seed. Proprietary interactive video data is emerging as a scarce, hotly contested asset for world-model training, with OpenAI reportedly among the suitors that tried to buy the underlying dataset outright.

$190 million Series A-1 extension at a $2 billion valuation: Upscale AI, which builds the hardware-and-software stack linking AI chips, memory, and storage for lower-latency training, raised the round led by Premji Invest with new backing from Nvidia, Salesforce Ventures, and Temasek, pushing total funding to $500M in under 18 months.

$110 million Series C: Taktile, which builds AI decisioning infrastructure for banks and fintechs across lending, onboarding, underwriting, fraud, and claims, raised $110M in a round led by Goldman Sachs, with participation from Balderton, Index Ventures, and other investors.

$100 million Series A: Scaled Cognition, founded by ex-Berkeley AI professor Dan Klein and Dan Roth, raised the round led by Khosla Ventures to build a hallucination-free enterprise model already deployed with Genesys and Fortune 500s in financial services and healthcare.

$50 million Series B: Patronus AI, which builds simulated “digital world” environments to stress-test AI agents before deployment, raised $50M led by Greenfield Partners with Notable Capital, Lightspeed, Datadog, and Samsung, on the back of 15x revenue growth.

$30 million extension: Runlayer, founded by ex-Zapier AI director Andrew Berman, raised $30M from Felicis and Khosla Ventures (bringing total funding to $42M) for a platform giving enterprises identity-aware permissions, observability, and runtime security over AI agents; customers include Instacart, Gusto, Opendoor, and dbt Labs.

$28 million Series A: Coval, a voice AI evaluation and simulation platform used by Zoom and Deepgram, raised $28M led by Norwest with Base10, Twilio Ventures, and Y Combinator. Voice agents are following the same trajectory as text agents: as soon as they’re deployed at scale, testing and monitoring tooling becomes a funded category in its own right.

$12.5 million Series A: Flagright, an AI-powered financial crime compliance startup, raised the round as financial institutions push to catch faster-moving fraud without sacrificing auditability. Compliance is becoming one of the clearest enterprise beachheads for AI, precisely because it demands the explainability that generic LLM deployments still struggle to provide.

$9 million seed: Probably, backed by Andreessen Horowitz, is building a validator “harness” that checks LLM outputs against deterministic systems to push accuracy toward 99.99% while running on smaller, cheaper models. As token costs rise, the market is rewarding startups that make weaker models reliable over those chasing frontier scale.

$4 million pre-seed: Fika Jobs, a Stockholm startup building a video-first hiring platform where AI agents conduct interviews and turn responses into shareable profiles, raised $4M led by Luminar Ventures. Recruiting is fragmenting into narrow AI wedges (sourcing, screening, now candidate-side video interviews) each attracting its own funding rather than consolidating under one player.

No posts

Read the original on datadeepdives.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.