Pinned Loading AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x. Python 301 24 Scaling Vision Pre-Training to 4K Resolution Python 226 12 When do we not need larger vision models? Python 419 16 VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud. Python 3.9k 330 Official code for "TOAST: Transfer Learning via Attention Steering" Python 187 10 Official code for "Top-Down Visual Attention from Analysis by Synthesis" (CVPR 2023 highlight) Jupyter Notebook 165 14