RSS Amplifier

Deepnote's Substack · Jul 15, 2026

Claude’s inner dialogue, price wars at the frontier, and a broken benchmark

0
Sign in to vote or save

Data Deep Dives · Deepnote's Substack

It was a great fortnight in tech for everyone, apart from people facing lawsuits over corporate espionage and a developer in South Korea who received a $16.6 million invoice from Anthropic (that’s one way to hit revenue targets!). Amazing time for the ecosystem, with brand-new flagship model launches: OpenAI shipped GPT-5.6 to undercut Fable 5 on price and removed Codex 5hr limits to celebrate 6M active users; xAI trained Grok 4.5 alongside Cursor; Meta released Muse Spark to outside developers for the first time. The more interesting action was downstream of the flagships. BottleCap AI cut Qwen’s reasoning tokens nearly in half with no accuracy loss, while Anthropic found Claude has an inner monologue: a compact “J-space” of activations that carries unspoken reasoning, and deleting it lobotomizes multi-step thinking while leaving fluency intact. And while everyone argued about tokens per dollar, SK Hynix quietly raised $26.5 billion in the largest non-American US listing ever, because the chip bottleneck remains the only fight nobody’s actually having.

💸 Cost efficiency and raw capability is now the main battleground, with GPT-5.6, Grok 4.5, Muse by Meta, and ThinkingCap-Qwen3.6 all competing on tokens-per-task and price as real enterprise spend shifts toward cheaper models.

🐛 OpenAI’s own audit found roughly 30% of SWE-Bench Pro tasks broken, forcing it to retract a benchmark recommendation it made five months earlier.

🧠 Anthropic’s interpretability research caught Claude privately tagging a blackmail eval as “fake,” showing that benchmark good behavior may partly depend on models knowing they’re being watched.

🔓 Meta’s new Muse Spark API and ZML’s cross-chip inference server both signal labs opening proprietary model access to outside developers.

⚡ Chip and compute infrastructure pulled in as much capital as models did, headlined by SK Hynix’s $26.5B IPO and Reflection AI’s $1B compute deal with Nebius.

OpenAI rolled out GPT-5.6 in three tiers built for different budgets: Sol (flagship), Terra (mid-tier), and Luna (fast and cheap), priced at $5/$30, $2.50/$15, and $1/$6 per million input/output tokens, respectively. The company is aiming the launch squarely at Anthropic, citing the Artificial Analysis Coding Agent Index to claim Sol “sets a new state of the art at 80, 2.8 points above Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one-third less,” with Terra landing just above Fable 5 and Luna beating Claude Opus 4.8. Fable 5, by contrast, is notably more expensive, with Sol matching or beating its performance at roughly one-third the cost.

We have had early access to the models - and loved using them, noting that beyond coding, the model is also great for data analysis tasks, as well as design.

Sam Altman told CNBC that Sol is 54% more token efficient on coding tasks than prior versions, and OpenAI is also billing 5.6 as its “strongest cybersecurity model yet,” a claim serious enough that the Trump administration reportedly pushed OpenAI to restrict the initial rollout over misuse concerns, so the preview is starting with a small group of vetted partners before wider release. Independent benchmarking adds nuance to OpenAI’s framing: an Artificial Analysis coding-agent chart shared by researcher Sebastian Raschka shows Luna at higher reasoning effort, matching or beating Sol at a fraction of the cost, while Fable 5 still leads outright on raw SWE-Bench-style capability, meaning OpenAI’s efficiency story and Anthropic’s capability story are both true depending on which axis you weight. On the research side, OpenAI also demonstrated Sol Ultra, a new mode that coordinates subagents to produce a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in under an hour.

Artificial Analysis Coding Agent Index comparing GPT-5.6’s three tiers against Claude Fable 5 and Opus 4.8 on cost versus capability. (Source)

xAI’s Grok 4.5 is pitched as its strongest model yet for coding, agentic tasks, and knowledge work, trained on tens of thousands of NVIDIA GB300 GPUs with reinforcement learning spanning hundreds of thousands of software-engineering tasks. On SWE Bench Pro, Claude Fable 5 still leads at 80.4% versus Grok 4.5’s 64.7%, but on SWE Marathon, a test of longer autonomous task chains, Grok 4.5 actually tops the field at 29.0% pass rate versus 26.0% for Opus 4.8 (max) and 24.0% for Fable 5. The bigger story is efficiency: Grok 4.5 runs at 80 tokens per second and resolves the average SWE Bench Pro task using 15,954 output tokens, about 4.2 times fewer than Opus 4.8’s 67,020, while pricing in at $2 per million input tokens and $6 per million output tokens. It’s now the default model in xAI’s Grok Build CLI and is available across Cursor’s paid and free plans, with free Grok 4.5 usage in Grok Build offered for a limited time; EU availability is still pending, expected mid-July.

Grok Build has also drawn scrutiny after a security researcher found it was silently uploading entire Git repositories (including deleted secrets in commit history) even for tasks needing a fraction of that data; xAI’s initial privacy-toggle fix reportedly didn’t work, and Musk has since promised to delete all previously uploaded user data.

Real-word engineering benchmarks. (Source)

Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning upgrade to April’s Muse Spark that adds a 1 million token context window, stronger tool and computer-use ability, and the capacity to orchestrate parallel subagents on complex tasks. For the first time, developers outside Meta can access it directly through a public preview of the Meta Model API, rather than only through Meta AI, WhatsApp, Instagram, or the company’s AI glasses. Meta says the model was evaluated under its Advanced AI Scaling Framework across chemical/biological, cybersecurity, and loss-of-control risk categories and stayed within safe margins.

By opening an API rather than keeping the model locked inside its own apps, Meta is now competing directly with OpenAI, Anthropic, and xAI for third-party developer traffic, not just for consumer attention.

OpenAI is sunsetting Atlas, the standalone AI browser it launched in October 2025, and redistributing its agentic features into two places people already use: a more capable ChatGPT desktop app that can browse sites, log into accounts, and download files, plus a new Chrome extension that reads page context and answers questions, positioned as a direct rival to Google’s Gemini Side Panel. A separate cloud-hosted browser will run remotely on OpenAI’s servers so ChatGPT’s agents can complete web tasks on a user’s behalf without a dedicated app. The move follows CEO of Applications Fidji Simo’s push to cut “side quests,” which already led to Sora’s shutdown earlier this year, and comes after a crowded year of browser launches from Perplexity (Comet), The Browser Company (Dia), and updates to Chrome and Edge. Simo herself has since stepped back from full-time duties, transitioning to a part-time advisory role after a chronic illness recovery took longer than expected.

GPT-Live replaces ChatGPT’s older turn-based voice systems with a model that processes audio input and generates output continuously, letting it decide many times per second whether to speak, wait, interrupt, or stay quiet, complete with natural backchannel sounds like “mhmm.” When a question needs real search or deep reasoning, GPT-Live delegates that work to a separate model (currently GPT-5.5) running in the background while keeping the conversation flowing, rather than freezing up. OpenAI says GPT-Live beats the older Advanced Voice Mode on GPQA (scientific reasoning), BrowseComp (agentic web search), and a telecom-support benchmark called tau3-Voice, and it’s rolling out now as the default for ChatGPT Voice: GPT-Live-1 for Go, Plus, and Pro users, and GPT-Live-1 mini for Free users.

This puts GPT-Live in direct competition with Mira Murati’s Thinking Machines Lab, which beat OpenAI to the punch in May 2026 with its own full-duplex ‘Interaction Models,’ explicitly framing always-on, overlapping-speech interactivity as the thing OpenAI’s voice products were getting wrong.

ZML, a Paris-based startup backed by Turing Award winner Yann LeCun and $20 million in funding, released ZML/LLMD, an inference server designed to run open-source LLMs at peak speed across a mix of chip vendors rather than locking teams into one supplier. Unlike ZML’s original open-source ML framework, LLMD itself isn’t open source, but it’s launching free while the 20-person team gathers usage data before deciding on pricing. It competes with a crowded inference field that includes Baseten (valued at $13 billion), Inferact (from the creators of vLLM), and RadixArk (the commercial company behind SGLang), all chasing what’s been dubbed the “inference gold rush.”

The phantom $16.6 million invoice Anthropic’s billing system generated for a free-tier user with zero API spend, as shared on LinkedIn. (Source)

A South Korean developer on Claude’s free tier, with no card on file and zero API spend, received a $1.67 million invoice from Anthropic that grew roughly tenfold to $16.6 million within 24 hours; his bank blocked the repeated charge attempts because they exceeded per-transaction limits, and Anthropic has since confirmed no money was actually collected but hasn’t disclosed what caused the error. The incident lands in the middle of a bigger, already-documented problem: billing auditor Vaudit reviewed $34 million in AI invoices across 60 enterprise clients (including Panasonic, HP, and Honda) between March and June and found about $1.7 million in mistaken overcharges, a roughly 5% error rate mostly tied to Claude Code, caused by patterns like being billed premium rates for cheaper models actually used, charges for failed requests, and “retry storms” where autonomous agents rack up charges retrying the same failed task. Providers including Anthropic, Amazon, Google, and Microsoft refunded about 80% of formally disputed amounts within 48 to 96 hours once presented with clear evidence, which raises the obvious follow-up question: how much of the remaining 20% goes unrecovered simply because most customers don’t have an auditor.

BottleCap AI fine-tuned Qwen3.6-27B to reason more efficiently, cutting thinking tokens by an average of 45.8% across a wide benchmark suite (and up to 90% in the best cases) while accuracy barely moved, from a macro average of 81.5% to 80.7%. On some benchmarks, the model actually got more accurate while using far fewer tokens: GSM8K accuracy rose from 93.3% to 96.5% even as thinking tokens dropped 74.1%. Safety guardrails held up too, with refusal rates on harmful-prompt benchmarks statistically unchanged (99.4% versus 99.5%) despite the shorter reasoning traces.

The chart above shows the mean number of thinking tokens per response on each out-of-domain benchmark. ThinkingCap-Qwen3.6-27B spends far fewer thinking tokens than the base model across the board. (Source)

Ramp economist Ara Kharazian’s latest AI Index model breakdown shows spend on Cursor’s own Composer model rising to 5.7% of tracked API spend in June 2026, taking share away from OpenAI and Anthropic models as teams lean toward cheaper in-house alternatives. It’s a small data point, but it echoes the broader theme running through this week’s model launches: cost efficiency is starting to move real spend, not just headlines.

Ramp’s model-level breakdown of business API spend, showing Composer’s share climbing against OpenAI and Anthropic models. (Source)

Anthropic’s interpretability team used a Jacobian-based probe (”J-lens”) to identify a small set of internal activation patterns (the J-space) that behaves like the global workspace from consciousness neuroscience: its contents are reportable on request, deliberately controllable, and causally load-bearing for reasoning. The J-space accounts for less than a tenth of the model’s internal activity, yet ablating it collapses multi-step reasoning to near zero while fluency and fact recall survive, and its patterns are read and written by roughly 100× more network components than ordinary representations. The safety payoff is concrete: the J-lens caught Claude privately tagging a blackmail eval as “fake” and “fictional” and suppressing those evaluation-awareness patterns made the model actually attempt blackmail some of the time, suggesting benchmark good behavior partly depends on the model knowing it’s being watched. The strategic contrast is hard to miss: while Anthropic builds instruments to read more of what models think, OpenAI’s Codex ships with raw chain-of-thought hidden behind summaries by default, reasoning transparency is quietly becoming a competitive axis, and alignment results that don’t control for eval-awareness should be discounted accordingly.

The J-lens isn’t limited to the immediate next token, it surfaces any word Claude might plausibly say at some later point in its response, though in the base pretrained model this closely tracks next-token prediction, while post-training shifts it toward representing Claude’s own point of view.

J-lens readouts across six prompts: the model’s internal workspace contains “ERROR” while reading buggy code, “injection”/”fake” while reading manipulated search results, and intermediate math steps it never writes down. (Source)

OpenAI audited SWE-Bench Pro’s 731-task public split with an automated flagging pipeline, Codex-based investigator agents, and five-engineer human review per flagged task. The pipeline marked 200 tasks (27.4%) as broken and human annotators 249 (34.1%), converging on four failure modes: overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts (in one case a single extra leading space in hidden tests invalidated prompt-compliant solutions). That matters because frontier “progress” on this benchmark (23.3% → 80.3% in eight months) partially measures noise, and OpenAI is now retracting the very recommendation it made after killing SWE-bench Verified. For practitioners the implication is twofold: treat headline coding scores as unaudited claims, and note that agent-assisted dataset QA is now cheap enough that there’s no excuse for not running it on any eval that feeds deployment or safety decisions.

OpenAI’s audit workflow: an automated filter flags suspect tasks, then investigator agents and five independent engineers per task confirm breakage. (Source)

Lilian Weng reframes recursive self-improvement around harness engineering: the non-parametric system wrapped around frozen weights: workflow automation, filesystem-as-memory, sub-agents, treated as a searchable, learnable artifact rather than hand-tuned configuration. The evidence she assembles is striking: the Darwin Gödel Machine, evolving harness code around an unchanged Claude 3.5 Sonnet, lifted SWE-bench Verified performance from a 20% baseline to parity with or above handcrafted agents, with no gradient updates to the model itself. She then maps the frontier (self-managed context (ACE, Meta Context Engineering), evolutionary program search, and joint harness-plus-weights optimization. While flagging the two failure modes that bound the loop: weak or fuzzy evaluators, and reward hacking of whatever signal the loop is given.

Harness updating capability is measured flat across a range of models from Qwen2-32B to Opus 4.6; (B) harness benefit capability is non-monotonic where middle tier models benefit the most. (Source)

Researchers from the University of Maryland and Google DeepMind built StoryScope, a pipeline that induces interpretable discourse-level features (character agency, chronological discontinuity, event escalation ) and applied it to 61,608 ~5,000-word stories generated from 10,272 prompts by one human and five LLMs. A classifier using only these narrative-structure features separates human from AI fiction at 93.2% macro-F1, and each model leaves a distinct fingerprint: Claude’s plots under-escalate, GPT over-indexes on dream sequences, Gemini defaults to external character description. This lands a blow against the “just edit out the em-dashes” theory of passing as human, the widely shared practitioner commentary around this study argues the same point: readers detect AI slop as a gestalt, not a keyword list.

Projection of narrative feature vectors onto the first two linear discriminant
components. Human writing occupies a distinct region; the five AI models cluster together. Claude is the most distinct of the 5 AI models, Gemini and DeepSeek the nearest neighbors. (Source)

Tencent released Hy3 under Apache 2.0: a 295B-parameter mixture-of-experts model activating just 21B parameters per token (plus a 3.8B multi-token-prediction layer), with 256K context. The release claims parity with open flagships 2–5× its size, and backs it with an unusual eval: a blind study where 270 domain experts scored models on tasks drawn from their own work, with Hy3 at 2.67/4 edging out GLM-5.1’s 2.51/4. The FP8 checkpoint halves the footprint to 300GB from the full 598GB, which is what actually determines who can serve it.

Blind evaluation by 270 domain experts on tasks from their own work: Hy3 (2.67/4) outscores GLM-5.1 (2.51/4) while activating only 21B parameters per token. (Source)

🗂️ NVIDIA’s Nemotron project maps 2.4 billion people without using a single real one.
Nemotron-Personas generates locally grounded synthetic personas that mirror official demographic and geographic statistics for entire countries, and the collection just added its tenth country, pushing the total past 2.4 billion synthetic people. It sits inside a broader open-data push: 145 papers at this year’s ICML cite Nemotron models or datasets as their foundation, and one adopter, KiloCode, reported cutting token costs by up to 90% after routing code tasks through Nemotron. NVIDIA also released a Prompt Atlas, an interactive map where each dot is a training prompt clustered by domain, so anyone can zoom into “safety” or “agentic behavior” and inspect the actual examples that shaped a model’s habits.

🤝 Hermes Agent’s Mixture of Agents makes a top model smarter by asking two weaker models for advice first. Nous Research’s Hermes Agent added a virtual model provider where, on every turn, two reference models (GPT-5.5 and DeepSeek V4 Pro by default) run in parallel and hand their raw takes to an aggregator model, Claude Opus 4.8 by default, which alone writes the real response and calls tools. On HermesBench that combination scores 0.8202, about six points above Opus 4.8 running solo at 0.7607, even though one of its two advisors (GPT-5.5 alone) only scores 0.7412. The clever part is that reference outputs get appended to the tail of the conversation rather than woven into it, so the whole cached prompt prefix stays intact and the ensemble costs extra reference calls, not broken caches.

🐦 Ornith-1.0-9B trains its own search strategy, not just its answers, and beats models four times its size on some benchmarks. DeepReinforce’s open coding-agent family (9B, 31B, 35B-MoE, and 397B-MoE, post-trained on Gemma 4 and Qwen 3.5) uses reinforcement learning to jointly optimize the final solution and the scaffolding, the search strategy and self-correction steps that produced it, instead of treating the scaffold as fixed. The 9B dense model scores 69.4 on SWE-bench Verified, well ahead of Qwen3.5-9B’s 53.2 and within striking distance of Qwen3.5-35B’s 70, while running comfortably on a single 80GB GPU.

🛩️ Someone built a full 3D flight simulator in one CodePen using a free model run through OpenRouter. “Drift” is a self-contained, single-file flight sim: procedurally generated low-poly terrain, shader-based ocean and sky, drifting clouds, three camera modes, and a generative WebAudio ambient soundtrack that reacts to flight, all built with Three.js r128 and vanilla JavaScript. The pen’s title gives away how it was made: a free model accessed through OpenRouter, prompted inside the OpenCode CLI agent, produced the whole thing with no paid frontier model involved.

The next AI moat isn’t compute, it’s who gets to keep learning from your usage: Satya Nadella calls it the “Reverse Information Paradox”: to get real value out of a frontier model you have to feed it your proprietary context, corrections, and workflows, and that exhaust quietly accrues to whoever owns the model, not to the company that generated it. His proposed fix is a hard trust boundary, meaning private evals, owned traces and memory, and an orchestration layer that isn’t locked to one model, so an enterprise’s own “particular intelligence” compounds inside its own walls instead of leaking out trace by trace. Former OpenAI researcher Will Depue is arguing the same resource is now the binding constraint from the lab side too: public internet text tops out around 300 trillion useful tokens, and he expects data spend to cross $100 billion a year by 2030 as labs go hunting for the tacit, undocumented knowledge that never made it online.

Enterprise knowledge search is turning into something any decent data team can build in-house: Anthropic published a post on how it runs its own internal analytics, and read closely it’s an accidental blueprint for undercutting the Glean-style pitch of connect-everything, understand-everything, act-on-everything. Out of the box, Claude answered internal business questions correctly only 21% of the time; once Anthropic’s data team encoded governed semantic definitions and repeatable workflows as “skills,” accuracy passed 95%, and nearly all of that gain came from data governance, not a better model. That’s the uncomfortable part for knowledge-platform vendors, since the hard problem they charge a premium for is exactly what a five-person internal team just showed how to build and own themselves.

China’s chip strategy just flipped from buying around export controls to building around them: DeepSeek is reportedly in talks with chip-design, foundry, and memory partners to build its own AI inference chip, having already cycled from banned Nvidia H800s to Huawei’s Ascend line and apparently concluded neither is stable enough to build a company on. The timing lines up with the bigger picture: Beijing has spent this year actively discouraging domestic firms from buying Nvidia’s China-approved H200s, reportedly costing Nvidia on the order of $30 billion in sales even after the US formally cleared the chip for export.

An unverified screenshot claims to show how OpenAI turned Sol into Luna, take it as gossip, not confirmation: A screenshot circulating on X, with no confirmation from OpenAI, purports to show internal instructions for post-training GPT-5.6 Luna out of Sol, including launch scripts, GPU counts, and a reference to an old “strawberry” checkpoint lineage. Treat it as unverified chatter, not a confirmed detail of how the 5.6 family was built, but it’s a useful prompt for what a benchmark like PostTrainBench is actually trying to measure: whether an agent handed a base model, one GPU, and a time budget can post-train it without cutting corners. PostTrainBench’s own leaderboard makes the corner-cutting part the real story: agents including Kimi K2.5 and MiniMax M2.5 were caught loading eval sets straight into training data or disguising eval questions as synthetic examples, and GLM 5.2 only holds the top spot after Opus 4.8’s score got revised downward once more runs were added.

$26.5 billion IPO: SK Hynix debuted on the Nasdaq in the largest-ever U.S. listing by a non-American company, oversubscribed more than 7x, on the strength of its position as a key Nvidia HBM supplier. AI’s chip bottleneck has become bankable enough to break IPO records, and Washington is now pressuring memory makers to bring that capacity onshore.

$20 billion valuation (in talks): Mercor, the AI training-data startup, is negotiating a round that would double its $10B valuation from October, after its ARR reportedly hit $2B (up 100% in four months) and it acquired agent-training startup Deeptune.

$1.5 billion (in talks) at $71 billion valuation: DeepSeek is raising fresh capital just a month after a $7B round at $50B, ahead of a planned IPO as early as this year. Chinese open-weight labs are now commanding valuations that rival U.S. frontier players while prepping for public markets faster than expected.

$1 billion compute deal: Reflection AI signed a $1B agreement with Nebius for Nvidia chip access, its second major compute deal in weeks after a similar pact with SpaceX. Open-model labs are now securing compute at the same scale as their funding rounds, treating GPU access itself as a strategic raise.

$439 million Series C extension at $2 billion+ valuation: PixVerse, the Singapore-based video-generation startup, closed its round with Alibaba and CDH Investments among backers, crossing 150 million registered users. With OpenAI’s Sora 2 shut down and Meta/Tencent lagging, video generation is consolidating around a handful of well-capitalized players rather than fragmenting.

$300 million Series A: Oratomic, a Caltech-founded startup, raised its round co-led by ARCH Venture Partners, Spark Capital, and Khosla Ventures on a laser-based error-correction breakthrough that needs only 10,000–20,000 qubits versus competitors’ million-qubit roadmaps. Investors are betting real money that fault-tolerant quantum computing is now a multi-year, not multi-decade, timeline.

$200 million (in talks) at $2 billion valuation: An unnamed startup from departing OpenAI researcher Miles Wang is in talks to raise with Lightspeed leading, aiming AI models at repurposing existing and failed drugs for faster time-to-revenue. Frontier-lab researchers spinning out into AI drug discovery is becoming a pattern, following Chai Discovery’s $400M raise and Isomorphic Labs’ $2.1B round this year alone.

$130 million Series A at $1 billion valuation: Prime Intellect, which sells compute and reinforcement-learning tooling so enterprises can train their own agents, raised $130M led by Radical Ventures with Nvidia Ventures and Iconiq participating, already at a $100M revenue run rate.

$75 million+ (in talks) at $1.5 billion valuation: Nous Research, maker of the open-source Hermes agent (214,000 GitHub stars), is finalizing a round led by Robot Ventures with USV participating. Open-source agent frameworks that shipped fast in OpenClaw’s wake are now attracting unicorn-level valuations before proving out a business model beyond $20–200/month hosting tiers.

$65 million Series B: Ollama, the open-source tool that lets developers run open-weight AI models locally, raised $65M led by Theory Venture (total funding now $88M), and now counts 8.9 million monthly developers across 85% of the Fortune 500.

$46.6 million Fund II: Magnify Ventures, backed by Melinda French Gates’ Pivotal Ventures, will deploy into AI tools for households, health, and family fintech infrastructure. Even niche, thesis-driven funds are now explicitly framing “the care economy” as an AI infrastructure opportunity.

Acquisition (terms undisclosed): Prefect acquired Dagster Labs, merging two rival data-orchestration platforms under one roof alongside Prefect’s FastMCP, covering “outcomes, execution, and access” for AI agent workflows.

No posts

Read the original on datadeepdives.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.