X (formerly Twitter)

🎯 Introducing Sparse Attention Vectors (SAVs): A breakthrough method for extracting powerful multimodal features from Large Multimodal Models (LMMs). SAVs enable SOTA performance on discriminative vision-language tasks (classification, safety alignment, etc.)! Links in replies! 🔎Using just ~20 attention heads & only few-shot examples, SAVs: - Outperform both LoRA and few-shot baselines - Work with image, text, & interleaved inputs - Extract features without finetuning - ready to go at test time! This project was a cross-collaborative effort between researchers from UC Berkeley, Carnegie Mellon University, and MIT-IBM Research (

@berkeley_ai

,

@CMU_Robotics

,

@MITIBMLab

). Many thanks to all of the collaborators and co-authors on this work: Brandon Huang, Tianning (Ray) Chai,

@ZhiqiuLin

,

@ArbelleAssaf

,

@RogerioFeris

,

@leokarlin

,

@trevordarrell

,

@RamananDeva

,

@roeiherzig

Read the original on x.com ↗