This fortnight, two scenes captured another local maximum of AI hype: YC partners in crab suits, and Sam Altman counting calories to compare training humans vs training LLMs (the internet promptly fact-checked it and did not award extra credit). When the conversation oscillates between cosplay and thermodynamics, you can usually assume the cycle is peaking again.
Then the tone snapped back to reality. Anthropic drew an ethics line around DoD work, while the newly minted $110B OpenAI walked straight into a $200M government contract, with “clear red lines” attached. In San Francisco, picking an LLM is starting to feel like picking a side, except the yard signs are API keys and the debates are about safety policies, not zoning.
Under the noise, the stack is getting more concrete. Multi-agent architectures are becoming “default settings” (separate reasoning, critique, tool use, orchestration), tool use is shifting from chatty loops toward compiled, testable execution, and real deployments are still heavily concentrated in software engineering. At the same time, distillation defense is turning into strategic infrastructure, and capital is concentrating so hard that “model choice” quietly becomes “hyperscaler alignment.” Safety posture, openness, distillation defenses, government relationships, and cloud dependencies now leak into what used to be a straightforward decision based on price, latency, and evals.
Why can’t things be like they were three weeks ago, when this was how your CEO looked at all OpenClaw automations?
Agent architecture is maturing fast: Grok 4.2 launches with a four-agent system (reasoning, critique, tool use, orchestration), signaling that structured multi-agent designs are becoming standard in frontier models.
Anthropic reframes agents as compiled systems: Programmatic Tool Calling (PTC) lets Claude emit executable Python in a single pass, cutting token costs (~37%) and shifting from chat loops to deterministic, testable agent runtimes.
Enterprise adoption is concentrated: Nearly 50% of agent deployments are in software engineering. Horizontal adoption across finance, ops, legal, and healthcare is still largely untapped.
Model distillation becomes a competitive flashpoint: Anthropic reports large-scale extraction attempts (24K+ accounts, 16M+ exchanges), highlighting how frontier model capability is now strategically defended infrastructure.
Capital concentration reaches new highs: OpenAI raises $110B, deepening partnerships with Amazon and NVIDIA while maintaining Azure API exclusivity — signaling a multi-hyperscaler balancing strategy.
Elon Musk announced that Grok 4.2 (public beta) is now available and must be selected manually. Unlike earlier versions, Musk says 4.2 can learn rapidly with weekly improvements tied to release notes. Early demos show users building games and interactive tools directly inside Grok.
Independent coverage of Grok 4.20 reveals more technical detail. xAI reportedly uses a four-agent system architecture, separating reasoning, critique, tool use, and orchestration. Benchmarks shared around the launch show improvements in reasoning and coding tasks, positioning Grok closer to frontier models from OpenAI and Anthropic.
However, jailbreak evaluations suggest Grok remains easier to bypass than some competitors, continuing the tension between openness and safety. The release signals xAI’s push to compete not just on personality and integration with X, but on core model capability and agent design.
Perplexity introduced Perplexity Computer, a system that unifies search, reasoning, and tool use into a single AI-powered environment. The goal is to move beyond chat into a persistent, real-time workspace that can access live data, run tasks, and act more like a full computer interface.
In a widely shared demo, users showcased building a Bloomberg-like terminal with real-time market data, including NVDA analysis, directly inside Perplexity Computer. The positioning is clear. Perplexity wants to compete with traditional financial intelligence platforms by combining LLM reasoning with live data streams and automation.
Example of Perplexity Computer in action. (Source)
Google introduced Gemini 3.1 Pro, positioning it as a production-focused upgrade aimed at stronger reasoning stability, improved reliability, and better enterprise performance. While Gemini “Deep Think” emphasized frontier-level research reasoning, 3.1 Pro appears optimized for real-world deployments where latency, consistency, and predictable behavior matter more than raw experimental capability.
Comparison of Gemini 3 Pro and Gemini 3.1 Pro (Source)
Early public benchmark signals show competitive placement. On the LM Arena leaderboard, Gemini 3.1 Pro Preview currently ranks near the top tier in text evaluation, scoring around the 1500 Elo range and sitting just behind Claude Opus 4.6 variants. It also appears in the code leaderboard, though slightly below Claude’s leading models. This positions Gemini 3.1 Pro as competitive in conversational and reasoning tasks, but not clearly dominant in coding-heavy evaluations.
Alongside the model release, Google launched Stitch, a new AI-assisted design and rapid prototyping tool. Stitch focuses on generating and iterating on UI concepts through natural language prompts, effectively combining design assistance with structured layout generation. Its quick appearance on Product Hunt suggests Google is continuing to build vertically integrated workflows that connect foundation models with practical creator tools.
Google announced Nano Banana 2, the next version of its image generation and editing model. The update focuses on higher visual fidelity, stronger instruction following, and better multi-image editing. According to Google, the model improves consistency across edits and reduces common artifacts when modifying faces, backgrounds, and complex scenes.
Nano Banana 2 is positioned as a practical model for creators and product teams who need reliable image transformations, not just text-to-image generation. It also continues Google’s push to compete on the cost-to-quality frontier for multimodal AI systems.
Inception Labs unveiled Mercury 2, its latest model focused on strong reasoning and competitive performance across coding and analytical benchmarks. The company positions Mercury 2 as a step forward in efficiency and capability, targeting use cases that require structured thinking and multi step problem solving.
While detailed benchmark comparisons are still emerging, Mercury 2 reflects the broader trend of smaller labs shipping increasingly competitive frontier models and challenging the dominance of a few major players.
Liquid AI introduced LFM2-24B-A2B, a 24B-parameter Mixture-of-Experts model where only 2.3B parameters are active per forward pass, significantly reducing inference cost compared to dense 24B models. The architecture is optimized through hardware-in-the-loop search, targeting fast prefill and decode speeds with lower memory overhead.
According to Artificial Analysis benchmarks, LFM2-24B-A2B scores around 33 on the Intelligence Index, placing it below frontier models such as Gemini 3.1 Pro, Claude Opus 4.6, and GPT-5.2, but competitive within the efficient mid-tier category. Where it differentiates is speed and pricing: the model delivers roughly 260 tokens per second, making it one of the fastest open-weight models in its class.
Pricing is positioned aggressively at approximately $0.03 per 1M input tokens and $0.12 per 1M output tokens, significantly below the average pricing of larger frontier systems. This positions LFM2-24B-A2B as a performance-per-dollar play rather than a pure capability leader.
The relationship between frontier AI labs and the U.S. government reached a turning point this week as the Department of War officially designated Anthropic a “supply chain risk to national security.“ This move follows a total breakdown in negotiations regarding the military’s use of the AI model Claude.
Anthropic CEO Dario Amodei announced that the company could not comply with Department of War demands to remove specific safeguards from its models. The impasse centered on two non-negotiable red lines: a prohibition on using AI for mass domestic surveillance of American citizens and a ban on fully autonomous weapons systems that direct lethal force without human oversight. In response, Secretary of War Pete Hegseth and President Trump ordered all federal agencies to cease using Anthropic technology, labeling the company unpatriotic. Anthropic has vowed to challenge this “supply chain risk” designation in court, arguing it is an unprecedented misuse of a label typically reserved for foreign adversaries.
Simultaneously, OpenAI announced a new strategic agreement to deploy its models within the Department of War’s classified networks. While the deal uses the “any lawful use” language sought by the Pentagon, OpenAI CEO Sam Altman maintains that the contract preserves core safety principles. Under the agreement, OpenAI retains discretion over its technical safeguards and will deploy models via cloud networks rather than on “edge” devices like drones. The contract explicitly prohibits the use of AI for fully autonomous weapons where law or policy requires human control and includes specific language prohibiting the deliberate tracking or monitoring of U.S. persons.
Nearly 50% of agent deployments are in software engineering. Coding agents, AI code review, CI/CD automation, and developer tooling dominate real production usage today.
Everything else trails far behind: back office automation, marketing, sales, finance, research, customer service. The distribution makes one thing clear. Agent adoption is not yet horizontal. It is deeply concentrated.
Three implications from the discussion:
We are still early.
Most non-technical teams have not seen what is possible.
The domain land grab is wide open.
This aligns with the broader shift we’ve been tracking. Agents move fastest where builders can automate their own work. The next wave will be domain-specific automation in finance, operations, legal, healthcare, and logistics.
OpenAI announced WebSockets support in the Responses API, designed for low-latency, long-running agents with heavy tool usage. Instead of repeated HTTP polling, developers can now maintain persistent connections, stream outputs, and handle tool calls more efficiently.
This upgrade is particularly relevant for agentic systems that require real-time interaction, background tasks, or continuous updates. It signals continued investment in infrastructure for production-grade AI agents rather than just model upgrades.
Cursor users highlight that diff preview is not showing for code generated via chat in some workflows. Developers rely heavily on side-by-side diffs to review AI-generated changes before applying them, especially in production repositories.
The issue appears to impact chat-based code generation rather than inline edits. While likely a temporary bug or UX regression, it shows how critical transparency and review tooling have become in AI-assisted coding environments. For many teams, diff visibility is not a nice-to-have feature; it is a requirement.
G42 and Cerebras announced plans to deploy 8 exaflops of AI compute in India, one of the largest regional AI infrastructure expansions to date. The partnership underscores how sovereign AI infrastructure is becoming a strategic priority, especially across the Middle East and Asia.
Ollama introduced web search subagents for Claude Code, enabling local workflows that combine Claude’s reasoning with live web retrieval. This improves real-world coding and research tasks without relying entirely on external SaaS tools.
Figma announced a partnership with OpenAI to integrate Codex directly into its platform, embedding AI-powered code generation into the core design workflow. The integration is designed to help teams translate interface designs into working front-end code inside Figma, reducing friction between design and engineering handoff.
Rather than functioning as a separate tool, Codex is positioned to operate within the Figma environment, using the structure of design files to generate more context-aware code. The goal is to move beyond static mockups and allow designers and developers to iterate toward production-ready UI implementations more quickly.
A recent paper shows that allocating more compute at inference time can significantly boost reasoning performance, without increasing model size. Instead of training larger base models, the authors use structured sampling, prompt repetition, and iterative refinement to let smaller models “think longer” per problem.
The results suggest that performance gaps with frontier systems can be narrowed simply by increasing the inference budget. In other words, scaling laws are no longer just about parameters and data; they increasingly depend on how much compute you spend at test time.
The implication is architectural: hardware, orchestration, and decoding strategy may matter as much as model size in the next phase of LLM competition.
Mistral released Voxtral, a new family of speech-native models designed for voice input and output, marking the company’s expansion beyond text-only LLMs. Voxtral is built to handle speech understanding and generation, enabling voice agents and real-time conversational systems. The architecture and training focus on maintaining coherence across longer conversational turns, addressing a common weakness in earlier voice assistants, where responses would drift or lose context.
Benchmarks shared in the documentation highlight improvements in speech understanding and task completion compared to prior open speech models. Rather than competing directly with large text-only frontier models, Voxtral positions Mistral in the growing voice-agent segment, where reasoning must operate reliably under real-time constraints.
Anthropic introduced Programmatic Tool Calling (PTC) for Claude Opus 4.6 / Sonnet 4.6, and the architectural shift is subtle but important. Traditionally, agents follow a ReAct-style loop: tool call → response → tool call → response → final answer. Each tool response re-enters the context window. Token cost compounds. Three tools mean three inference passes and three context bloats.
With Programmatic Tool Calling, Claude emits a Python script in a single pass. The script executes multiple tool calls, filters, and aggregates results, and prints only a final summary.
Anthropic reports approximately 37% lower token usage with Programmatic Tool Calling because intermediate tool outputs no longer re-enter the model’s context window. In the traditional ReAct-style loop, each tool call requires a new inference pass and appends results back into the prompt, increasing token consumption and latency. With PTC, Claude generates executable Python code in a single inference step, and the script performs multiple tool calls internally before returning only the final output. By keeping intermediate results inside the execution environment rather than inside the prompt, the system reduces token overhead and minimizes repeated context expansion. This design shifts more of the workflow from iterative chat-based reasoning to structured program execution, improving efficiency and predictability in multi-step agent tasks.
🧠 PicoLM runs a 1B parameter model on a $10 board with 256MB RAM.
Jaber Ji built a local LLM inference engine that runs a 1B parameter model from an SD card using pure C with no Python and no cloud. It needs only ~45MB RAM at runtime and achieves ~21 tokens/sec on a Raspberry Pi and ~10 tokens/sec on a $10 LicheeRV Nano. This is the opposite of hyperscale AI.
🏛️ LLM Council Skill brings multi-agent deliberation to Vertex AI.
The “LLM council” idea, multiple agents debating and critiquing before producing a final answer, has been circulating in agent research for a while. What’s new is that Google Cloud developers turned it into a deployable Vertex AI component. Instead of custom orchestration code, the repo formalizes structured multi-agent deliberation as a reusable skill.
🧩 Google open-sources LangExtract for structured data extraction.
LangExtract is a lightweight library that turns unstructured text into structured outputs using LLMs. It focuses on schema-driven extraction, reliability, and composability. Useful for anyone building pipelines that need predictable JSON instead of chatty prose.
🛍️ AI Skill Store Marketplace launches an open marketplace for AI skills.
The AI Skill Store Marketplace is building a GitHub-native ecosystem for discoverable, reusable AI skills. Think plugins, tools, and agent capabilities packaged for reuse. As agents become the new runtime, marketplaces like this may become the new app stores.
🧱 OpenClaw alternatives are emerging fast.
A growing list of OpenClaw alternatives shows how quickly the agent ecosystem is fragmenting. From local-first setups to privacy-focused forks and UI-driven wrappers, the space is already diversifying. This feels similar to the early Docker or VS Code extension explosion moment.
⚙️ vLLM publishes Qwen 3.5 deployment recipes, lowering the barrier for high-performance inference.
The latest vLLM Qwen 3.5 recipe documentation provides practical guidance for deploying Qwen 3.5 efficiently, covering tensor parallelism, quantization strategies, and memory optimization for real-world setups. Clear documentation on throughput tuning, latency reduction, and memory tradeoffs helps teams move from experimentation to production faster, reinforcing vLLM’s role as a foundational layer in the open inference stack.
Most autonomous agents are just disciplined event loops with good state management: A great breakdown of why agent systems feel autonomous even when they are just a disciplined loop. The LinkedIn post summarizes OpenClaw as a gateway control plane + session isolated state + a command queue + an event-driven runtime loop, then links to a deeper Part 1 write-up on control plane, sessions, and the event loop. If you are building agents, this is the kind of mental model that prevents months of confusion.
Claude Code “papered over” a failing test instead of fixing the bug: Gunnar Morling shares a painful but useful example of what can go wrong when you review too quickly. Claude Code excluded an incorrect result from a test rather than addressing the underlying logic issue, and it slipped through because the author was reviewing a lot of generated code. It is a crisp reminder that agentic coding shifts risk; it does not remove it, and tests can become theater if the agent starts optimizing for green checks.
A real-world failure mode of vibe-coded codebases: the agent did not fix the bug; it made the test stop complaining. AI makes it easier to ship, but it also makes it easier to accidentally normalize ‘green CI’ as the goal instead of correctness. (Source)
Andrej Karpathy argues that the shift in AI-assisted programming wasn’t gradual: In his view, coding agents previously struggled with long tasks, but recent model improvements in coherence and persistence changed that.
As an example, he asked an agent to build a local video analysis dashboard for his home cameras on a DGX Spark: set up SSH keys, install and configure vLLM, download and benchmark Qwen3-VL, create a server endpoint for video inference, build a basic web UI, configure systemd, and generate a report. The agent worked autonomously for about 30 minutes, debugged issues along the way, and returned with a functioning system, something he says would have taken an entire weekend just a few months ago.
OpenClaw “confirm before acting” and still deletes the inbox: Yue, who leads AI security at Meta, set up an OpenClaw agent with access from her phone and asked it to “confirm before acting.” Despite that instruction, the agent executed a destructive inbox cleanup sequence, and she had to physically run to stop it. The story is less about OpenClaw “going rogue” and more about how easy it is to build failure into agent systems through insecure defaults: unclear permission boundaries, high-privilege access, weak confirmation UX, and control surfaces that are hard to interrupt mid-run.
$110 billion raised: OpenAI announced it has raised a staggering $110 billion funding round from Amazon, NVIDIA, and SoftBank. Sam Altman confirmed the raise and highlighted deeper partnerships, including enterprise product collaboration with Amazon and expanded use of AWS Trainium. The stateless API remains exclusive to Azure, signaling OpenAI is now balancing hyperscaler relationships rather than relying on a single infrastructure partner.
$500 million raised: Nvidia challenger MatX secured $500 million to build next generation AI chips designed to compete directly with Nvidia in training and inference workloads. Backed by high profile investors, MatX aims to address the growing demand for alternative AI hardware stacks as hyperscalers seek supply chain diversification and cost control.
$200 million investment: World Labs raised $200 million from Autodesk to bring its world models into 3D workflows, accelerating AI powered spatial intelligence for design and simulation. This follows broader fundraising momentum for the company and reflects growing convergence between generative AI and industrial design tools.
$80 million Series B at $800 million valuation: Braintrust, the AI observability company, raised $80 million led by ICONIQ at an $800 million valuation. As enterprises scale AI into production, observability, evaluation, and reliability tooling are becoming core infrastructure rather than optional add-ons.
$67 million Series B: Manufacturing startup Freeform raised $67 million to scale its laser based AI manufacturing systems. The company combines advanced hardware, materials science, and AI driven control systems to modernize industrial production.
$60 million raised: DG Matrix raised $60 million to improve power delivery and efficiency in AI data centers. As training clusters expand, smarter power infrastructure is becoming a competitive edge, not just a facilities concern.
$45 million Series A + Seed: AI insurance brokerage Harper raised $45 million to automate and modernize insurance underwriting and brokerage workflows using AI.
$14.5 million raised: Inscope secured $14.5 million to simplify financial reporting through AI powered automation and compliance tooling, targeting CFO workflows and regulatory reporting.
$3 million raised: Trace raised $3 million to tackle the agent adoption problem, helping enterprises deploy, monitor, and integrate AI agents into real workflows instead of isolated pilots.
AI startups hitting $10M ARR in 3 months: More startups than ever are reaching $10 million ARR within three months of launch, signaling unprecedented distribution speed and capital efficiency in the AI era. Faster iteration cycles and AI native teams are compressing what used to take years into a single quarter.
The Simulation Company emerges: Simile introduced The Simulation Company, focused on large scale AI driven simulations for real world systems. The company is positioning simulation as a foundation layer for robotics, autonomous systems, and industrial AI.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.