CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
Meta AI Research, GenAI; University of Oxford, VGG
Nikita Karaev, Iurii Makarov, Jianyuan Wang, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, Christian Rupprecht
Project Page | Paper #1 | Paper #2 | X Thread | BibTeX
CoTracker is a fast transformer-based model that can track any point in a video. It brings to tracking some of the benefits of Optical Flow.
CoTracker can track:
- Any pixel in a video
- A quasi-dense set of pixels together
- Points can be manually selected or sampled on a grid in any video frame
Try these tracking modes for yourself with our Colab demo or in the Hugging Face Space 🤗.
Updates:
-
[January 21, 2025] 📦 Kubric Dataset used for CoTracker3 now available! This dataset contains 6,000 high-resolution sequences (512×512px, 120 frames) with slight camera motion, rendered using the Kubric engine. Check it out on Hugging Face Dataset.
-
[October 15, 2024] 📣 We're releasing CoTracker3! State-of-the-art point tracking with a lightweight architecture trained with 1000x less data than previous top-performing models. Code for baseline models and the pseudo-labeling pipeline are available in the repo, as well as model checkpoints. Check out our paper for more details.
-
[September 25, 2024] CoTracker2.1 is now available! This model has better performance on TAP-Vid benchmarks and follows the architecture of the original CoTracker. Try it out!
-
[June 14, 2024] We have released the code for VGGSfM, a model for recovering camera poses and 3D structure from any image sequences based on point tracking! VGGSfM is the first fully differentiable SfM framework that unlocks scalability and outperforms conventional SfM methods on standard benchmarks.
-
[December 27, 2023] CoTracker2 is now available! It can now track many more (up to 265*265!) points jointly and it has a cleaner and more memory-efficient implementation. It also supports online processing. See the updated paper for more details. The old version remains available here.
-
[September 5, 2023] You can now run our Gradio demo locally.
Quick start
The easiest way to use CoTracker is to load a pretrained model from torch.hub:
Offline mode:
pip install imageio[ffmpeg], then:
import torch # Download the video url = 'https://github.com/facebookresearch/co-tracker/raw/refs/heads/main/assets/apple.mp4' import imageio.v3 as iio frames = iio.imread(url, plugin="FFMPEG") # plugin="pyav" device = 'cuda' grid_size = 10 video = torch.tensor(frames).permute(0, 3, 1, 2)[None].float().to(device) # B T C H W # Run Offline CoTracker: cotracker = torch.hub.load("facebookresearch/co-tracker", "cotracker3_offline").to(device) pred_tracks, pred_visibility = cotracker(video, grid_size=grid_size) # B T N 2, B T N 1
Online mode:
cotracker = torch.hub.load("facebookresearch/co-tracker", "cotracker3_online").to(device) # Run Online CoTracker, the same model with a different API: # Initialize online processing cotracker(video_chunk=video, is_first_step=True, grid_size=grid_size) # Process the video for ind in range(0, video.shape[1] - cotracker.step, cotracker.step): pred_tracks, pred_visibility = cotracker( video_chunk=video[:, ind : ind + cotracker.step * 2] ) # B T N 2, B T N 1
Online processing is more memory-efficient and allows for the processing of longer videos. However, in the example provided above, the video length is known! See the online demo for an example of tracking from an online stream with an unknown video length.
Visualize predicted tracks:
After installing CoTracker, you can visualize tracks with:
from cotracker.utils.visualizer import Visualizer vis = Visualizer(save_dir="./saved_videos", pad_value=120, linewidth=3) vis.visualize(video, pred_tracks, pred_visibility)
We offer a number of other ways to interact with CoTracker:
- Interactive Gradio demo:
- A demo is available in the
facebook/cotrackerHugging Face Space 🤗. - You can use the gradio demo locally by running
python -m gradio_demo.appafter installing the required packages:pip install -r gradio_demo/requirements.txt.
- A demo is available in the
- Jupyter notebook:
- You can run the notebook in Google Colab.
- Or explore the notebook located at
notebooks/demo.ipynb.
- You can install CoTracker locally and then:
-
Run an offline demo with 10 ⨉ 10 points sampled on a grid on the first frame of a video (results will be saved to
./saved_videos/demo.mp4)):python demo.py --grid_size 10
-
Run an online demo:
python online_demo.py
-
A GPU is strongly recommended for using CoTracker locally.
Installation Instructions
You can use a Pretrained Model via PyTorch Hub, as described above, or install CoTracker from this GitHub repo. This is the best way if you need to run our local demo or evaluate/train CoTracker.
Ensure you have both PyTorch and TorchVision installed on your system. Follow the instructions here for the installation. We strongly recommend installing both PyTorch and TorchVision with CUDA support, although for small tasks CoTracker can be run on CPU.
Install a Development Version
git clone https://github.com/facebookresearch/co-tracker cd co-tracker pip install -e . pip install matplotlib flow_vis tqdm tensorboard
You can manually download all CoTracker3 checkpoints (baseline and scaled models, as well as single and sliding window architectures) from the links below and place them in the checkpoints folder as follows:
mkdir -p checkpoints cd checkpoints # download the online (multi window) model wget https://huggingface.co/facebook/cotracker3/resolve/main/scaled_online.pth # download the offline (single window) model wget https://huggingface.co/facebook/cotracker3/resolve/main/scaled_offline.pth cd ..
You can also download CoTracker3 checkpoints trained only on Kubric:
# download the online (sliding window) model wget https://huggingface.co/facebook/cotracker3/resolve/main/baseline_online.pth # download the offline (single window) model wget https://huggingface.co/facebook/cotracker3/resolve/main/baseline_offline.pth
For old checkpoints, see this section.
Evaluation
To reproduce the results presented in the paper, download the following datasets:
And install the necessary dependencies:
pip install hydra-core==1.1.0 mediapy
Then, execute the following command to evaluate the online model on TAP-Vid DAVIS:
python ./cotracker/evaluation/evaluate.py --config-name eval_tapvid_davis_first exp_dir=./eval_outputs dataset_root=your/tapvid/path
And the offline model:
python ./cotracker/evaluation/evaluate.py --config-name eval_tapvid_davis_first exp_dir=./eval_outputs dataset_root=/fsx-repligen/shared/datasets/tapvid offline_model=True window_len=60 checkpoint=./checkpoints/scaled_offline.pth
We run evaluations jointly on all the target points at a time for faster inference. With such evaluations, the numbers are similar to those presented in the paper. If you want to reproduce the exact numbers from the paper, add the flag single_point=True.
These are the numbers that you should be able to reproduce using the released checkpoint and the current version of the codebase:
| Kinetics, |
|---|