VentureBeat · Aug 16, 2026
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
0Sign in to vote or save
This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail,…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.