A Very Big Video Reasoning Suite

We bet on a future that video reasoning is the next fundamental intelligence paradigm, after language reasoning, where spatiotemporal embodied world experiences could be more naturally captured.

Data Engines

View All
clock
GitHub
Knowledge in-domain testset
The clock shows 6:53. Show what the clock will look like after 2 hours.
First Frame
Last Frame
Abstraction training set
The scene shows a tree structure with nodes connected by edges, a green starting node at the top, a red ending node at the bottom, and a yellow smiley agent positioned at the green starting node. The agent can move along edges, visiting one node per step. Move the yellow smiley agent from the green starting node to the red ending node along the unique tree path, showing the complete movement process step by step.
First Frame
Last Frame
undirected_graph_navigation
GitHub
Spatiality training set
The scene shows a network of nodes connected by undirected edges (edges without arrows) with a green starting node, a red ending node, and a purple triangular agent positioned at the green starting node. The agent can move along any edge, traversing it from one end to the other in either direction, moving from one node to an adjacent node each step. Move the purple triangular agent from the green starting node to the red ending node along the path with the minimum number of steps.
First Frame
Last Frame
handle_object_reappearance
GitHub
Transformation training set
There is an object in the center. Move the object right off-screen, then return along the same path to the center.
First Frame
Last Frame
color_addition
GitHub
Perception in-domain testset
Two balls of different colors are placed at separate locations. Show them moving toward each other at identical speeds. When overlapping, display the additive mixture of their colors. Continue until completely merged at the midpoint.
First Frame
Last Frame

Inference Results

View Full Bench
Communicating Vessels - Samples
00
01
02
03
04
Task Domains 1/5
Communicating Vessels
Knowledge in-domain testset
Symmetry Completion
Abstraction out-of-domain testset
Key Door Matching
Spatiality in-domain testset
Grid Shift
Transformation in-domain testset
Multi Object Placement
Perception in-domain testset
Prompt
Loading...
Ground Truth
First
First Frame
Final
Final Frame
Model Outputs
1/
VBVR-Wan2.2
VBVR-Wan2.2
CogVideoX 1.5
Kling 2.6
LTX-2
Runway Gen-4
Sora 2
Veo 3
Wan 2.2 I2V
Hunyuan I2V
Seedance 2.0

Leaderboard

Modality
Split
Type
Category
2026-04-28