RSS Amplifier

Agentic AI Weekly · Jul 8, 2026

Agentic AI Weekly | Berkeley RDI | July 8, 2026

0
Sign in to vote or save

This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.

Researching Highlight: SageCTF | Agentic AI Summit Agenda Now Live and Tickets are Running Out!

SageCTF: The Most Capable AI Agent for The Most Challenging CTF

DEF CON CTF is the hardest competition in cybersecurity. Its challenges, typically involving reverse-engineering binaries and crafting working exploits, have historically demanded teams of 20, 50, even hundreds of elite hackers working nonstop across a three-day window.

This year, one solo competitor changed the math. It wasn’t human.

Meet SageCTF

Built by teams at UC Santa Barbara and UC Berkeley on a next-generation agent scaffold called OpenSage, SageCTF entered the DEF CON CTF 2026 qualifiers as a solo player. The results were striking:

  • 8 flags recovered across 7 of the hardest challenges in the contest

  • 1,743 points, landing in the top 5% of 686 scored teams

  • Outscored every team that self-reported using “No AI” or “Low AI”

To respect the competition’s rules and the culture of CTF, the team solved the challenges but never submitted flags. The scoring is a measure of what SageCTF could have achieved.

These weren’t easy problems

DEF CON challenges sit far beyond routine CTFs. The single easiest one had a solve rate of just 12.8%. Six of SageCTF’s seven solves had solve rates below 10%. Its hardest full solve, gitvfs, was cracked by only 30 of 686 teams.

And these solves took time. SageCTF averaged five hours per challenge with zero human intervention, even with one run for 12 hours straight. Throughout, it inspected artifacts, built environments, tested hypotheses, debugged scripts, abandoned dead ends, and preserved partial progress, all on its own.

How it works

SageCTF’s edge isn’t just a strong underlying model, it’s architecture. Rather than following a fixed workflow, OpenSage lets the agent build its own team for each problem. A root agent reads the challenge, then spins up specialized subagents for reverse engineering, exploit testing, and integration. Subagents can even create their own subagents.

A real SageCTF agent topology from the final `Birdhouse` solve.

Other ingredients in the mix:

  • Inter-agent communication: a message bus broadcasts breakthroughs so agents stay synced without sharing one giant context

  • Multi-model orchestration: across seven solves, SageCTF made nearly 24,000 model calls and processed 2.4 billion tokens, routing work between GPT-5.5, Claude-Opus-4.6, and DeepSeek-V4-Pro based on each model’s strengths

  • Hierarchical memory for managing knowledge across marathon solves

In a head-to-head benchmark of 50 hard challenges, SageCTF solved 39 to Claude Code’s 13, and every Claude Code solve was also a SageCTF solve.

SageCTF vs Claude Code on CTF benchmarks

The honest caveat

SageCTF still missed plenty. On one challenge, it rebuilt nearly the entire exploit pipeline but flubbed the final byte order, landing a hair away from the answer with no scoreboard feedback to correct course.

The takeaway from the team is clear-eyed: the best human teams remain stronger, and CTF is still a craft built on taste, creativity, and intuition. But the threshold has shifted. Hard CTF autonomy is now an agentic problem and SageCTF just proved how much ground these systems have gained.

Read the Blog

Read on OpenSage

Agentic AI Summit 2026: Tickets are Running Out!

Save the date! The Agentic AI Summit returns to Berkeley on August 1–2, 2026, welcoming 5,000+ expected in-person attendees for two days of insights and innovation. Building on last year’s sold-out success—with 2,000+ in‑person attendees and 40,000+ global livestream participants—the summit will bring together researchers, builders, industry leaders, and the global agentic AI community for keynotes, technical talks and panels, hands-on workshops, live demos, and more!

Full Summit Agenda Now Available!

We are thrilled to announce that the full agenda for the Summit is now available on our website! The Summit is set to feature 200+ speakers across 4 stages, plus 200+ poster presentations during the event! If you’ve been waiting for the schedule to come out before grabbing your tickets, now is the time to take a look and secure your spot!

View the Agenda

In addition, we are excited to showcase our expanded list of speakers for the Summit! We are honored to have such a great group of academics, founders, executives, and investors participate in this year’s event, and more will be announced soon!

🎟️ Standard Tickets (Limited Capacity)
A limited number of Standard Tickets are still available:

  • Standard: $499

Get your tickets


Join Us Virtually

Can't attend the Summit in person? Join us virtually—register for the livestream to access all sessions from anywhere in the world.

Register for Livestream


Apply to the Startup Spotlight

Building the future of Agentic AI? We’re looking for you.

We are excited to announce that applications are now open for the Startup Spotlight at the Agentic AI Summit! Selected startups will have the opportunity to showcase their products and innovations to researchers, founders, engineers, investors, industry leaders, and policymakers from across the AI ecosystem. Apply to showcase your product and innovation!

Applications will be reviewed on a rolling basis, so we encourage interested teams to apply as soon as possible. The deadline to apply is this Friday, July 10, at 11:59 PM PT.

Apply Now

You can also learn more about the Summit, the full agenda, and our event sponsors by reading Professor Dawn Song’s recent LinkedIn and Twitter/X updates:

LinkedIn

Twitter/X

Trends This Week

  • Anthropic announced Claude Sonnet 5, the latest version of its mid-tier Sonnet model, focused on stronger agentic performance across planning, coding, tool use, and computer-use tasks. Anthropic says Sonnet 5 narrows the gap with Opus 4.8 while running at a lower cost, positioning it as a more practical default model for enterprise agent workflows. In addition, Anthropic announced Claude Science, a beta AI workbench for researchers that integrates literature analysis, code execution, scientific databases, reproducible artifacts, and compute management across local, remote, and HPC environments. Lastly, the company redeployed Claude Fable 5, its higher-capability Claude model designed for broad use with stronger safeguards than Mythos 5, after U.S. export controls on both models were lifted. Anthropic said the redeployment includes updated cybersecurity safeguards, a proposed industry framework for scoring jailbreak severity, and deeper collaboration with government partners on pre-release testing and AI security.

  • OpenAI introduced GeneBench-Pro, a benchmark for testing AI agents on research-level computational biology tasks. The benchmark includes 129 synthetic problems across 10 domains, requiring models to analyze messy datasets, choose appropriate methods, identify quality-control issues, and produce final scientific judgments. OpenAI says GPT-5.6 Sol reached a 28.7% pass rate at the highest reasoning level, or 31.5% with Pro mode enabled, while typical problems were estimated to take human experts 20–40 hours. OpenAI says that the results point to progress in agentic scientific reasoning, while showing that models still struggle to reliably complete complex research workflows independently.

  • This past week, Microsoft announced Microsoft Frontier Company, a new venture focused on enterprise AI deployment and transformation, backed by a $2.5 billion investment and 6,000 industry and engineering experts. The initiative will embed AI engineers directly with customers to co-design, deploy, and continuously improve AI systems using Microsoft’s existing AI platform. Microsoft positions the organization as broader than traditional forward-deployed engineering, embedding engineering experts directly into environments to co-design, deploy, and continuously improve AI systems at scale. The launch adds to a growing wave of AI deployment initiatives, following Amazon Web Services’ $1 billion commitment to its own effort just two days earlier, while OpenAI and Anthropic have also introduced similar joint ventures backed by private equity capital.

Don’t miss the developments shaping Agentic AI. Subscribe for weekly coverage of groundbreaking research, emerging trends, and critical insights across Agentic AI and the broader AI landscape.

Get cutting-edge AI news sent to you every week by subscribing below!

Read on berkeleyrdi.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.