🎯 Introducing Sparse Attention Vectors (SAVs): A breakthrough method for extracting powerful multimodal features from Large Multimodal Models (LMMs). SAVs enable SOTA performance on discriminative vision-language tasks (classification, safety alignment, etc.)! Links in replies! 🔎Using just ~20 attention heads & only few-shot examples, SAVs: - Outperform both LoRA and few-shot baselines - Work with image, text, & interleaved inputs - Extract features without finetuning - ready to go at test time! This project was a cross-collaborative effort between researchers from UC Berkeley, Carnegie Mellon University, and MIT-IBM Research (
@berkeley_ai,
@CMU_Robotics,
@MITIBMLab). Many thanks to all of the collaborators and co-authors on this work: Brandon Huang, Tianning (Ray) Chai,
@ZhiqiuLin,
@ArbelleAssaf,
@RogerioFeris,
@leokarlin,
@trevordarrell,
@RamananDeva,
@roeiherzig