GitHub

Pinned Loading

  1. [NeurIPS 2022 Spotlight] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

    Python 1.8k 170

  2. A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.

    Python 397 8

  3. [CVPR 2022 Oral & TPAMI 2024] MixFormer: End-to-End Tracking with Iterative Mixed Attention

    Python 540 75

  4. [CVPR 2025] Multiple Object Tracking as ID Prediction

    Python 553 53

  5. [NeurIPS 2025 Spotlight] StreamForest: Efficient Online Video Understanding with Persistent Event Memory

    Python 129 5

  6. [CVPR 2026] DDT: Decoupled Diffusion Transformer

    Python 416 22

Read the original on github.com ↗