Research Scientist, Google DeepMind
I am a Research Scientist at Google DeepMind in Paris. I completed my PhD in 2024 in computer vision at the Visual Geometry Group, at the University of Oxford, collaborating with Prof. Andrew Zisserman. I am interested in Video understanding, Vision and Language, Multi-modal learning, Learning with weak supervision, Sign Language Recognition and Translation.
04 / 2025
10 / 2024
06 / 2024
07 / 2023
I have started a 6-month internship at Meta in New York, working with Nicolas Carion.
07 / 2023
09 / 2022
09 / 2022
08 / 2022
07 / 2022
07 / 2022
I have started a 6-month internship at Google in Grenoble, working with Cordelia Schmid.
10 / 2021
10 / 2021
02 / 2021
02 / 2021
09 / 2020
08 / 2020
07 / 2020
09 / 2019
I have joined VGG as a PhD student under the supervision of Profressor Andrew Zisserman.
04 / 2024
04 / 2024
Research
Verbs in Action: Improving verb understanding in video-language models
L. Momeni*, M. Caron, A. Nagrani, A. Zisserman, C. Schmid
ICCV, 2023.
@INPROCEEDINGS{Momeni23,
title = {Verbs in Action: Improving verb understanding in video-language models},
author = {Momeni, Liliane and Caron, Mathilde and Nagrani, Arsha and Zisserman, Andrew and Schmid, Cordelia},
year = {2023},
booktitle = {ICCV},
}
Understanding verbs is crucial to modelling how people and objects interact with each other and the environment through space and time. Recently, state-of-the-art video-language models based on CLIP have been shown to have limited verb understanding and to rely extensively on nouns, restricting their performance in real-world video applications that require action and temporal understanding. In this work, we improve verb understanding for CLIP-based video-language models by proposing a new Verb-Focused Contrastive (VFC) framework. This consists of two main components: (1) leveraging pretrained large language models (LLMs) to create hard negatives for cross-modal contrastive learning, together with a calibration strategy to balance the occurrence of concepts in positive and negative pairs; and (2) enforcing a fine-grained, verb phrase alignment loss. Our method achieves state-of-the-art results for zero-shot performance on three downstream tasks that focus on verb understanding: video-text matching, video question-answering and video classification. To the best of our knowledge, this is the first work which proposes a method to alleviate the verb understanding problem, and does not simply highlight it.
Weakly-supervised Fingerspelling Recognition in British Sign Language
Prajwal K R*, H. Bull*,
, S. Albanie, G. Varol, A. Zisserman
BMVC, 2022.
@INPROCEEDINGS{Prajwal22,
title = {Weakly-supervised Fingerspelling Recognition in British Sign Language Videos},
author = {Prajwal, K R and Bull, Hannah and Momeni, Liliane and Albanie, Samuel and Varol, G{\"u}l and Zisserman, Andrew},
year = {2022},
booktitle = {BMVC},
}
The goal of this work is to detect and recognize sequences of letters signed using fingerspelling in British Sign Language (BSL). Previous fingerspelling recognition methods have not focused on BSL, which has a very different signing alphabet (e.g., two-handed instead of one-handed) to American Sign Language (ASL). They also use manual annotations for training. In contrast to previous methods, our method only uses weak annotations from subtitles for training. We localize potential instances of fingerspelling using a simple feature similarity method, then automatically annotate these instances by querying subtitle words and searching for corresponding mouthing cues from the signer. We propose a Transformer architecture adapted to this task, with a multiple-hypothesis CTC loss function to learn from alternative annotation possibilities. We employ a multi-stage training approach, where we make use of an initial version of our trained model to extend and enhance our training data before re-training again to achieve better performance. Through extensive evaluations, we verify our method for automatic annotation and our model architecture. Moreover, we provide a human expert annotated test set of 5K video clips for evaluating BSL fingerspelling recognition methods to support sign language research.
Automatic dense annotation of large-vocabulary sign language videos
L. Momeni*, H. Bull*, Prajwal K R*, S. Albanie, G. Varol, A. Zisserman
ECCV, 2022.
@INPROCEEDINGS{Momeni22,
title = {Automatic dense annotation of large-vocabulary sign language videos},
author = {Momeni, Liliane and Bull, Hannah and Prajwal, K R and Albanie, Samuel and Varol, G{\"u}l and Zisserman, Andrew},
year = {2022},
booktitle = {ECCV},
}
Recently, sign language researchers have turned to sign language interpreted TV broadcasts, comprising (i) a video of continuous signing and (ii) subtitles corresponding to the audio content, as a readily available and large-scale source of training data. One key challenge in the usability of such data is the lack of sign annotations. Previous work exploiting such weakly-aligned data only found sparse correspondences between keywords in the subtitle and individual signs. In this work, we propose a simple, scalable framework to vastly increase the density of automatic annotations. Our contributions are the following: (1) we significantly improve previous annotation methods by making use of synonyms and subtitle-signing alignment; (2) we show the value of pseudo-labelling from a sign recognition model as a way of sign spotting; (3) we propose a novel approach for increasing our annotations of known and unknown classes based on in-domain exemplars; (4) on the BOBSL BSL sign language corpus, we increase the number of confident automatic annotations from 670K to 5M. We will make these annotations publicly available to support the sign language research community.
BOBSL: BBC-Oxford British Sign Language Dataset
S. Albanie*, G. Varol*,
, H. Bull*, T. Afouras, Himel Chowdhury, Neil Fox, Bencie Woll, Rob Cooper, Andrew McParland, A. Zisserman
ArXiv, 2021.
@InProceedings{Albanie2021bobsl,
author = {Samuel Albanie and G{\"u}l Varol and Liliane Momeni and Triantafyllos Afouras and Hannah Bull and Himel Chowdhury and Neil Fox and Bencie Woll and Rob Cooper and Andrew McParland and Andrew Zisserman},
title = {{BOBSL}: {BBC}-{O}xford {B}ritish {S}ign {L}anguage {D}ataset},
howpublished = {\url{https://www.robots.ox.ac.uk/~vgg/data/bobsl}},
year = {2021}
}
In this work, we introduce the BBC-Oxford British Sign Language (BOBSL) dataset, a large-scale video collection of British Sign Language (BSL). BOBSL is an extended and publicly released dataset based on the BSL-1K dataset introduced in previous work. We describe the motivation for the dataset, together with statistics and available annotations. We conduct experiments to provide baselines for the tasks of sign recognition, sign language alignment, and sign language translation. Finally, we describe several strengths and limitations of the data from the perspectives of machine learning and linguistics, note sources of bias present in the dataset, and discuss potential applications of BOBSL in the context of sign language technology.
Visual Keyword Spotting with Attention
KR Prajwal*,
, T. Afouras, A. Zisserman
BMVC, 2021.
@InProceedings{Prajwal2021,
author = {Prajwal, KR and Momeni, Liliane and Afouras, Triantafyllos and Zisserman, Andrew},
title = {Visual Keyword Spotting with Attention},
booktitle = {BMVC},
year = {2021}
}
In this paper, we consider the task of spotting spoken keywords in silent video sequences -- also known as visual keyword spotting. To this end, we investigate Transformer-based models that ingest two streams, a visual encoding of the video and a phonetic encoding of the keyword, and output the temporal location of the keyword if present. Our contributions are as follows: (1) We propose a novel architecture, the Transpotter, that uses full cross-modal attention between the visual and phonetic streams; (2) We show through extensive evaluations that our model outperforms the prior state-of-the-art visual keyword spotting and lip reading methods on the challenging LRW, LRS2, LRS3 datasets by a large margin; (3) We demonstrate the ability of our model to spot words under the extreme conditions of isolated mouthings in sign language videos.
Aligning Subtitles in Sign Language Videos
H. Bull*, T. Afouras*, G. Varol, S. Albanie,
@InProceedings{Bull2021,
author = {Bull, Hannah and Afouras, Triantafyllos and Varol, G{\"u}l and Albanie, Samuel and Momeni, Liliane and Zisserman, Andrew},
title = {Aligning Subtitles in Sign Language Videos},
booktitle = {ICCV},
year = {2021}
}
The goal of this work is to temporally align asynchronous subtitles in sign language videos. In particular, we focus on sign-language interpreted TV broadcast data comprising (i) a video of continuous signing, and (ii) subtitles corresponding to the audio content. Previous work exploiting such weakly-aligned data only considered finding keyword-sign correspondences, whereas we aim to localise a complete subtitle text in continuous signing. We propose a Transformer architecture tailored for this task, which we train on manually annotated alignments covering over 15K subtitles that span 17.7 hours of video. We use BERT subtitle embeddings and CNN video representations learned for sign recognition to encode the two signals, which interact through a series of attention layers. Our model outputs frame-level predictions, i.e., for each video frame, whether it belongs to the queried subtitle or not. Through extensive evaluations, we show substantial improvements over existing alignment baselines that do not make use of subtitle text embeddings for learning. Our automatic alignment model opens up possibilities for advancing machine translation of sign languages via providing continuously synchronized video-text data.
Read and Attend: Temporal Localisation in Sign Language Videos
G. Varol*,
*, S. Albanie*, T. Afouras*, A. Zisserman
CVPR, 2021.
@InProceedings{Varol21,
author = {Varol, G{\"u}l and Momeni, Liliane and Albanie, Samuel and Afouras, Triantafyllos and Zisserman, Andrew},
title = {Read and Attend: Temporal Localisation in Sign Language Videos},
booktitle = {CVPR},
year = {2021}
}
The objective of this work is to annotate sign instances across a broad vocabulary in continuous sign language. We train a Transformer model to ingest a continuous signing stream and output a sequence of written tokens on a large-scale collection of signing footage with weakly-aligned subtitles. We show that through this training it acquires the ability to attend to a large vocabulary of sign instances in the input sequence, enabling their localisation. Our contributions are as follows: (1) we demonstrate the ability to leverage large quantities of continuous signing videos with weakly-aligned subtitles to localise signs in continuous sign language; (2) we employ the learned attention to automatically generate hundreds of thousands of annotations for a large sign vocabulary; (3) we collect a set of 37K manually verified sign instances across a vocabulary of 950 sign classes to provide a more robust sign language benchmark; (4) by training on the newly annotated data from our method, we outperform the prior state of the art on the BSL-1K sign language recognition benchmark.
Signer Diarisation in The Wild
S. Albanie*, G. Varol*,
, T. Afouras, A. Brown, C. Zhang, E. Coto, N. C. Camgöz, B. Saunders, A. Dutta, N. Fox, R. Bowden, B. Woll, A. Zisserman
Technical Report, 2021.
@InProceedings{Albanie21a,
author = {Samuel Albanie and G{\"u}l Varol and Liliane Momeni and Triantafyllos Afouras and Andrew Brown and Chuhan Zhang and Ernesto Coto and Necati Cihan Camg{\"o}z and Ben Saunders and Abhishek Dutta and Neil Fox and Richard Bowden and Bencie Woll and Andrew Zisserman},
title = {SeeHear: Signer Diarisation and a New Dataset},
booktitle = {International Conference on Acoustics, Speech, and Signal Processing},
year = {2021},
keywords = {Signer Diarisation, Sign Language Datasets}
}
In this work, we propose a framework that enables collection of large-scale, diverse sign language datasets that can be used to train automatic sign language recognition models. The first contribution of this work is SDTRACK, a generic method for signer tracking and diarisation in the wild. Our second contribution is to show how SDTRACK can be used to automatically annotate 90 hours of British Sign Language (BSL) content featuring a wide range of signers, and including interviews, monologues and debates. Using SDTRACK, this data is annotated with 35K active signing tracks, with corresponding video-level signer identifiers and subtitles, and 40K automatically localised sign labels.
Watch, Read and Lookup: Learning to Spot Signs from Multiple Supervisors *
L. Momeni*, G. Varol*, S. Albanie*, T. Afouras, A. Zisserman
ACCV, 2020.
@InProceedings{Momeni2020bsldict,
author = {Liliane Momeni and G{\"u}l Varol and Samuel Albanie and Triantafyllos Afouras and Andrew Zisserman},
title = {Watch, Read and Lookup: Learning to Spot Signs from Multiple Supervisors},
booktitle = {Asian Conference on Computer Vision},
year = {2020}
}
The focus of this work is sign spotting: for a given sign corresponding to a keyword, the task is to identify whether and where it has been signed in a continuous, co-articulated sequence in a sign language video. To achieve this, we train a model using multiple types of available supervision by: (1) watching existing sparsely labelled footage; (2) reading associated subtitles, which provide weak supervision; (3) looking up words in visual sign language dictionaries to enable spotting of novel signs. This multi-source approach enables robust and scalable training for real-world sign spotting applications.
Seeing Wake Words: Audio-Visual Keyword Spotting
L. Momeni, T. Afouras, T. Stafylakis, S. Albanie, A. Zisserman
BMVC, 2020.
@InProceedings{Momeni2020seeing,
title = {Seeing wake words: Audio-visual Keyword Spotting},
author = {Liliane Momeni and Triantafyllos Afouras and Themos Stafylakis and Samuel Albanie and Andrew Zisserman},
booktitle = {British Machine Vision Conference},
year = {2020}
}
In this work, we consider the task of audio-visual keyword spotting, where the goal is to detect occurrences of a pre-specified keyword in unconstrained video content. We introduce KWS-Net, a model that fuses audio and visual streams using a novel spatiotemporal attention mechanism. The model exploits complementary cues from visual speech to improve robustness in noisy or silent environments. Our contributions include a new dataset tailored for AV-KWS, architectural innovations for effective multimodal fusion, and an extensive evaluation demonstrating state-of-the-art performance across challenging scenarios. This work highlights the importance of visual information in improving keyword spotting accuracy, particularly in real-world applications such as voice assistants and transcription tools.
BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues
S. Albanie*, G. Varol*,
, T. Afouras, J.S. Chung, N. Fox, A. Zisserman
ECCV, 2020.
@InProceedings{Albanie2020bsl1k,
author = {Samuel Albanie and G{\"u}l Varol and Liliane Momeni and Triantafyllos Afouras and Joon Son Chung and Neil Fox and Andrew Zisserman},
title = {{BSL-1K}: Scaling up co-articulated sign language recognition using mouthing cues},
booktitle = {European Conference on Computer Vision},
year = {2020}
}
Recent progress in fine-grained gesture and action classification, and machine translation, point to the possibility of automated sign language recognition becoming a reality. A key stumbling block in making progress towards this goal is a lack of appropriate training data, stemming from the high complexity of sign annotation and a limited supply of qualified annotators. In this work, we introduce a new scalable approach to data collection for sign recognition in continuous videos. We make use of weakly-aligned subtitles for broadcast footage together with a keyword spotting method to automatically localise sign-instances for a vocabulary of 1,000 signs in 1,000 hours of video. We make the following contributions: (1) We show how to use mouthing cues from signers to obtain high-quality annotations from video data - the result is the BSL-1K dataset, a collection of British Sign Language (BSL) signs of unprecedented scale; (2) We show that we can use BSL-1K to train strong sign recognition models for co-articulated signs in BSL and that these models additionally form excellent pretraining for other sign languages and benchmarks - we exceed the state of the art on both the MSASL and WLASL benchmarks. Finally, (3) we propose new large-scale evaluation sets for the tasks of sign recognition and sign spotting and provide baselines which we hope will serve to stimulate research in this area.