GitHub

Pinned Loading

  1. AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.

    Python 301 24

  2. Scaling Vision Pre-Training to 4K Resolution

    Python 226 12

  3. When do we not need larger vision models?

    Python 419 16

  4. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.

    Python 3.9k 330

  5. Official code for "TOAST: Transfer Learning via Attention Steering"

    Python 187 10

  6. Official code for "Top-Down Visual Attention from Analysis by Synthesis" (CVPR 2023 highlight)

    Jupyter Notebook 165 14

Read the original on github.com ↗