X (formerly Twitter)

If you want a vision encoder for dexterous manipulation, what should be the most important part to model? πŸ€” Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control. But collecting annotated robotic trajectories at scale is SUPER expensive and largely unrealistic. So, how do we bridge this gap? We introduce CAIP (Contrastive Action-Image Pre-training) ⬇️ πŸ”Έ Action-centric upstream: we align visual observations with action chunks through a contrastive objective. πŸ”Έ Human video as a proxy: we represent 3D human hand poses analogously to robotic end-effector actions, tapping into a massive source of human demonstrations. πŸ”Έ Massive scale: pre-trained on over 32,000 hours of manipulation video, driving both sample efficiency and robust generalization. πŸ”Έ Hardware proven: achieves a 76% average success rate on a real-world Dexmate Vega bimanual

@DexmateAI

manipulator with dual 22-DoF Sharpa Wave hands

@SharpaRobotics

. πŸ”Έ State-of-the-art: significantly outperforms strong baselines like DINOv2, SigLIP, MVP, and Qwen3.5 ViT across complex tasks, even under unexpected lighting changes and visual distractors. 🌐 Project: caip-encoder.github.io πŸ“ Blog: x.com/roeiherzig/sta… πŸ“„ Paper: arxiv.org/abs/2606.17256 πŸ’» Model: huggingface.co/yuvansharma/ca…

Read the original on x.com β†—