AI agents are experiencing explosive growth, but building them still depends almost entirely on human expertise. Every major ADK today requires developers to manually design agent topology, tooling systems, and memory architecture from scratch. This human-centric paradigm does not scale, cannot dynamically adapt across tasks, and may not be optimal for AI reasoning — mirroring early ML’s reliance on hand-crafted features, right before end-to-end learning took over.
New research from teams at UC Santa Barbara, UC Berkeley, Columbia, Duke, Google DeepMind, and UCLA—including Berkeley RDI co-director, Professor Dawn Song—argues that agent development is now at the same inflection point: instead of manually designing agent structures and capabilities, we should move toward an AI-centric paradigm, where a base “agent scaffold” is provided and the AI itself learns how to organize topology, tools, and memory from experience and feedback.
OpenSage is an Agent Development Kit built around three capabilities that no existing ADK provides:
Self-generating agent topology: AI dynamically creates and manages sub-agents at runtime, forming vertical (sequential) or horizontal (parallel) topology automatically.
Dynamic tool synthesis: Agents write and register their own tools (Python modules, Bash scripts, and more), execute them in isolated dependency-aware sandboxes, and run tools asynchronously in the background while continuing to reason. Agents can also create reusable skills beyond tools.
Hierarchical graph-based memory: Short- and long-term memory are stored as typed Neo4j graphs supporting both graph-based traversal and semantic retrieval, managed by a dedicated memory agent.
Evaluated against four diverse and popular benchmarks, SageAgent (built on OpenSage) consistently ranks first under the same backbone model:
60.2% on CyberGym vs. 39.4% for OpenHands.
78.4% on Terminal-Bench 2.0.
59.0% on SWE-Bench Pro Python vs. 40.2% for SWE-agent.
46.8% on DevOps-Gym, the only agent completing end-to-end tasks.
On a 300-instance CyberGym subset, removing vertical or horizontal topology drops performance to 60.3% and 62.3% respectively; disabling all features collapses it to 33.7%.
Tooling System Powers Complex Tasks
The tooling system is equally critical. Without the dynamic tooling system, agents fall back to a raw terminal interface and performance collapses.
On SWE-Bench Pro, full hierarchical memory achieves 59.0% vs. 56.4% for a baseline graph memory and 56.2% for none.
Emergent Capabilities and What’s Next
Across experiments, agents exhibit novel self-programming behaviors: spawning debugging sub-agents, constructing syntax-aware fuzzers, and selectively persisting high-signal memory. Current models don’t yet exploit these capabilities optimally — invocation frequency remains low in some cases, and models occasionally forget to reuse existing agents or memory, hallucinate tools, or create sub-agents with mismatched toolsets. To help build stronger models, OpenSage now includes integrated RL training support (AReaL, slime) and sandbox infrastructure for large-scale parallel rollouts.
The longer-term goal is for OpenSage to be not only an agent construction framework, but also a training scaffold for next-generation reasoning models, where AI can design, coordinate, and refine agents through interaction and feedback, unifying agent construction and problem solving within a single AI-centric loop.
Read the Paper | Read the Blog | Visit the Website | View the Code | Share Professor Dawn Song’s Tweet and LinkedIn Post |
Phase 2, Sprint 2 of the AgentX–AgentBeats competition started this week, and we’re extremely excited to see what you build over the next few weeks!
For Phase 2, participants are building purple agents to tackle the select top green agents from Phase 1 and compete on the public leaderboards. Unlike Phase 1, where participants competed across all tracks throughout the entire duration, Phase 2 introduces a sprint-based format. The competition is organized into four rotating sprints.
📋 March 23 – April 12, 2026
Three tracks and associated benchmarks/green agents are live for the second sprint:
Research Agent Track
FieldWorkArena (GitHub, Leaderboard)
MLE-Bench (GitHub, Leaderboard)
Multi-Agent Evaluation Track
MAizeBargAIn (GitHub, Leaderboard)
τ²-Bench Track
τ²-Bench (GitHub, Leaderboard)
Computer Use & Web Agent Track
CAR-bench (GitHub, Leaderboard)
OSWorld-Verified (GitHub, Leaderboard)
We’ve opened up the Sprint 2 submission form, which you can access by clicking the button below!
Sprint 3 (4/13 – 5/3): Agent Safety, Coding Agent, Cybersecurity Agent
Sprint 4 (5/4-5/24): General Purpose Agents, the grand finale of AgentBeats Phase 2, where everything culminates.
AgentX–AgentBeats is the first competition to explicitly spotlight general-purpose agents, testing broad capability, adaptability, and robustness across diverse tasks rather than a single domain. While earlier sprints emphasize depth, this final sprint showcases breadth and real-world readiness.
Participants are encouraged to compete in multiple tracks across multiple sprints during Phase 2. Teams and team members who submit purple agents in any sprint will also be eligible to enter a raffle for free tickets to the Agentic AI Summit later this year.
For more details on each sprint and how to compete in Phase 2, please refer to the AgentX–AgentBeats website!
The Berkeley Xcelerator, a non-dilutive accelerator program designed to support pre-seed and seed-stage startups building at the forefront of Agentic AI, is now open for applications!
The Xcelerator is built in partnership with Berkeley RDI’s research community and ecosystem partners, offering selected teams the support, resources, and guidance to take their startup to the next level! In addition, the Xcelerator is open to everyone; you do not need to be affiliated with UC Berkeley to apply!
Why apply to the Xcelerator?
Unparalleled access to frontier research and expertise through close collaboration with Berkeley RDI’s community across agentic AI, AI safety and security, and the broader AI landscape.
Practical enablement through industry partnerships, including cloud, GPU, and API credits provided by industry partners such as Google Cloud, Google DeepMind, OpenAI, and Nebius, with more to be announced!
Visibility and network effects through the Berkeley ecosystem and Berkeley RDI’s global community of 56,000+ developers and builders, including the rapidly growing Agentic AI MOOC community
A culminating Demo Day at the Agentic AI Summit (August 1–2, 2026), bringing together 5,000+ in-person attendees and placing your startup directly in front of top-tier VCs, leading AI researchers, industry executives, and strategic partners.
We’re looking for AI and Agentic AI startups at the pre-seed or seed stage. If you think that you or your team are a good fit, we encourage you to learn more and apply via the Xcelerator website and form below!
📅 As a reminder, we’ve extended the deadline to apply to Tuesday, March 31, 2026! We can’t wait to review your application!
Our sincerest thanks to all of our sponsors and partners:
Save the date! The Agentic AI Summit returns to Berkeley on August 1–2, 2026, welcoming 5,000+ expected in-person attendees for two days of insights and innovation. Building on last year’s sold-out success—with 2,000+ in‑person attendees and 40,000+ global livestream participants—the summit will bring together researchers, builders, industry leaders, and the global agentic AI community for keynotes, technical talks and panels, hands-on workshops, live demos, and more!
In addition, we are excited to introduce our speakers for the Summit! We are honored to have such a great group of academics, founders, executives, and investors participate in this year’s event, and more will be announced soon!
🎟️ Early‑Bird Pricing (Limited Capacity)
A limited number of early‑bird tickets are still available:
Student Early-Bird: $149
Standard Early-Bird: $299
If you’re looking to secure the best ticket price and be part of the conversation shaping the future of Agentic AI, we encourage you to register early. We look forward to welcoming you to Berkeley this August.
We’re also thrilled to share that the Call for Speaking Proposals (CFP) for the Agentic AI Summit 2026 is now open!
If you’re interested in sharing your work through a technical talk, panel discussion, workshop, or tutorial, or poster presentation—and helping advance the frontiers of the Agentic AI—we warmly invite you and/or your team to apply and be part of the conversation at the Summit.
Please complete the form below to submit your proposal. The program committee will review submissions on a rolling basis. The application deadline is 4/15—we can’t wait to hear from you.
Partner with us to shape the future of Agentic AI. If you’re interested in sponsoring the summit, please complete the sponsorship application form. Sponsorship opportunities are limited and reviewed/allocated on a rolling basis, so we encourage you to apply early.
At GTC, Nvidia CEO Jensen Huang said the company now projects “at least $1 trillion” in AI infrastructure demand through 2027, up from roughly $500 billion projected for Blackwell and Rubin systems through 2026. In addition, Nvidia also announced NemoClaw, a new stack for deploying autonomous AI agents on OpenClaw with built-in privacy, security, and hybrid local/cloud model support via the OpenShell runtime. Alongside it, the company officially unveiled the Vera Rubin platform, integrating seven chips across compute, networking, and storage into a unified system designed to support training, post-training, and real-time agentic inference.
OpenAI released GPT-5.4 mini and nano, extending many of GPT-5.4’s capabilities to faster, lower-cost models optimized for high-volume workloads. GPT-5.4 mini improves on GPT-5 mini across coding, reasoning, multimodal understanding, and tool use, while running more than 2× faster and approaching GPT-5.4 performance on benchmarks like SWE-Bench Pro and OSWorld-Verified. OpenAI highlighted the models’ potential in compositional systems, where larger models handle planning while GPT-5.4 mini operates as a fast subagent, alongside strong multimodal performance in computer-use tasks.
Anthropic released what it describes as the largest qualitative study of attitudes toward AI to date. It uses a “Claude Interviewer” system to conduct open-ended conversations with 81,000 users across 159 countries in 70 languages. The top reported perceived benefit of AI was professional excellence, with users citing time savings, financial independence, and improved life management, while the most common concern was unreliability—27% worried about AI hallucinations and instruction-following errors. Concerns about jobs and the economy (22%) and loss of human agency (22%) followed closely, with job anxiety emerging as the strongest predictor of overall sentiment. Regional sentiment varied, with India and South America above average, while the U.S., Europe, Japan, and South Korea were neutral or below. Around 11% of respondents reported no concerns, often describing AI as a neutral tool similar to electricity or the internet.
In a recent interview, OpenAI Chief Scientist Jakub Pachocki said the company is prioritizing a fully automated “AI researcher” as its next major research goal, positioning it as a “North Star” that combines advances in reasoning models, agents, and interpretability. The system is intended to operate as an agent-based framework capable of independently tackling complex problems across domains like math, physics, biology, and business. First, OpenAI plans to release an “autonomous AI research intern” by September, designed to complete discrete research tasks that would take humans several days, followed by a multi-agent research system targeted for 2028. Pachocki stated, “We are getting close to a point where we’ll have models capable of working indefinitely in a coherent way… You kind of have a whole research lab in a data center.”
Don’t miss the developments shaping Agentic AI. Subscribe for weekly coverage of groundbreaking research, emerging trends, and critical insights across Agentic AI and the broader AI landscape.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.