Pinned Loading [NeurIPS 2022 Spotlight] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training Python 1.8k 170 A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response. Python 397 8 [CVPR 2022 Oral & TPAMI 2024] MixFormer: End-to-End Tracking with Iterative Mixed Attention Python 540 75 [CVPR 2025] Multiple Object Tracking as ID Prediction Python 553 53 [NeurIPS 2025 Spotlight] StreamForest: Efficient Online Video Understanding with Persistent Event Memory Python 129 5 [CVPR 2026] DDT: Decoupled Diffusion Transformer Python 416 22