sites.google.com

September 9, 2026

Malmö, Sweden

OVERVIEW

Real-world deployment introduces challenges that benchmarks often overlook, including limited compute, strict latency requirements, domain shifts, and noisy or incomplete data. This workshop addresses these gaps by focusing on Real-World Video Representation Learning: transitioning from evaluation-driven development to model representations designed for real-world deployment.

By shifting the focus from leaderboard-driven improvements to reliability, efficiency, adaptability, and real-world generalization, the workshop aims to bridge cutting-edge research in video representation learning with the operational demands of real-world AI systems. Overall, it promotes a research agenda that prioritizes impact, deployability, and sustained performance beyond controlled benchmarks.

INVITED SPEAKERS & PANELISTS

University of Leiden, NL

University of Tokyo, JP

Google DeepMind, US

AGENDA

09:00 - 09:10: Introduction 

09:10 - 09:40: Speaker 1: Hazel Doughty 

09:40 - 10:10: Speaker 2: Viorica Patraucean

10:10 - 10:30: Coffee Break (Malmo Massan Exhibit Hall) 

10:30 - 11:00: Oral Session 

11:00 - 11:30: Speaker 3: Gül Varol

11:30 - 12:00: Panel Discussion [Hazel Doughty, Yoichi Sato, Gül Varol, Ye Xia]

12:00 - 12:05: Closing Remarks

12:05 - 12:30: Break 

12:30 - 13:30: Poster Session  

ACCEPTED POSTERS

  1. Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study, Hao Dong

  2. CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling, Sayan Deb Sarkar, Rémi Pautrat, Ondrej Miksik, Marc Pollefeys, Iro Armeni, Mahdi Rad, Mihai Dusmanu [Oral]

  3. Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition, Lars Doorenbos, Duc Manh Vu, Serday Ozsoy, Juergen Gall

  4. Toward Real-World Neural Video Representation: Compact, Low-Complexity, and Scalable, Ho Man Kwan, Tianhao Peng, Fan Zhang, David Bull

  5. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement, Seohyun Lee, Seoung Choi, Dohwan Ko, Jongha Kim, Hyunwoo Kim [Oral]

  6. TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning, Soumya Jahagirdar, Edson Araujo, Anna Kukleva, Jehanzeb Mirza, Saurabhchand Bhati, Samuel Thomas, Brian Kingsbury, Rogerio Feris, James Glass, Hilde Kuehne [Oral]

  7. What Moves? Context-Aware Localized Latent Actions, Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Björn Ommer

  8. When Conditional Sequence Matching Does Not Transfer to Global Video Retrieval, Arjang Talattof

  9. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model, Narges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero, Tommie Kerssies, Bastian Leibe, Gijs Dubbelman, Daan de Geus

  10. I Have a Stream: Making Self-Supervised Learning Work on Continuous Video, Ivan Martinović, Lukas Knobel, Yuki M. Asano

  11. Adapting MLLMs for Nuanced Video Retrieval, Piyush Bagad, Andrew Zisserman


POSTER GUIDELINES

Please prepare the posters following the official ECCV 2026 poster guidelines.


ORGANIZING TEAM

PROGRAM COMMITTEE

Ana Manzano - UvA / Amsterdam UMC

Melissa Tijink - University of Twente

Mohammadreza Salehi - University of Amsterdam

Niccolò Cavagnero - Eindhoven University of Technology

Rakshith Srinivasa Murthy - Adobe

Riccardo Santambrogio - Politecnico di Milano

Simone Alberto Peirone - Politecnico di Torino

Zhiqi Miao - University of Groningen

Read the original on sites.google.com ↗