RSSAmplifier

VentureBeat · Aug 16, 2026

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

0
Sign in to vote or save

This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.

DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail,…

Read on venturebeat.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.