Focus: The focus of this workshop is the task of searching for content within a corpus of videos using natural language queries.
Background: Convolutional neural networks have yielded unprecedented progress on a wide range of image-centric benchmarks, driven through a combination of well-annotated datasets and end-to-end training. However, naively extending this approach from images to higher-level video understanding tasks quickly becomes prohibitive with respect to the computation and data annotation required to jointly train multi-modal high-capacity models. An attractive alternative is to repurpose collections of existing pretrained models as "experts", offering representations which have been specialised for semantically relevant machine perception tasks. In addition to efficacy, this approach offers a second key advantage---it encourages researchers without access to industrial computing clusters to contribute towards questions of fundamental importance to video understanding: How should temporal information be used to maximum effect? How best to exploit complementary and redundant signals across different modalities? How can models be designed that function robustly across different video domains? To stimulate research into these questions, we are hosting a challenge that focuses on learning from videos and language with experts: making available a diverse collection of carefully curated visual and audio pre-extracted features across a set of five influential video datasets as part of a "pentathlon" of video understanding.
Schedule
A video of the event can be found here.
| 14:00 - 14:10 | Welcome (slides) | ||
| 14:10 - 14:40 | Invited Keynote (slides) | Chen Sun (Google) |
![]() |
| 14:40 - 14:48 | (Invited oral): Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning (paper, slides) | Shizhe Chen | |
| 14:49 - 14:57 | (Invited oral): ActBERT: Learning Global-Local Video-Text Representations (paper, slides) | Linchao Zhu | |
| 14:58 - 15:06 | (Invited oral): Visual Grounding in Video for Unsupervised Word Translation (paper, slides) | Jean-Baptiste Alayrac | |
| 15:06 - 15:15 | Virtual coffee break | ||
| 15:15 - 15:45 | Invited Keynote (slides) | Carl Vondrick
(Columbia University) |
![]() |
| 15:45 - 15:50 | Challenge results (slides) | 15:50 - 15:56 | (1st place): MMT (technical report, slides) | Valentin Gabeur | 15:56 - 16:02 | (2nd place): cszhe (technical report, slides) | Shizhe Chen |
| 16:02 - 16:08 | (3rd place): LEgGOdt (technical report, slides) | Kaixu Cui | |
| 16:10 - 16:40 | Virtual coffee break | ||
| 16:40 - 16:48 | (Invited oral): Action Modifiers: Learning from Adverbs in Instructional Videos (paper, slides) | Hazel Doughty | |
| 16:49 - 16:57 | (Invited oral): Condensed Movies: Story Based Retrieval with Contextual Embeddings (paper, slides) | Max Bain | |
| 16:57 - 17:10 | Virtual coffee break | ||
| 17:10 - 17:40 | Invited Keynote (slides) | Anna Rohrbach (UC
Berkeley) |
![]() |
| 17:40 - 17:45 | Closing remarks (slides) |
Challenge: The Video Pentathlon
| Ranking | Team name | Peantathlon score | |
| 1st place | MMT | 2511.43 | technical report |
| 2nd place | cszhe | 2448.56 | technical report |
| 3nd place | LEgGOdt | 1895.01 | technical report |
| 4th place | haoxiaoshuai | 1496.98 | technical report |
Dates
Challenge track
- Challenge Opens 9th April
- Test set Release 9th May
-
Challenge Closes
2nd June
23:59 UTC 4th June - Report Deadline 23:59 UTC 9th June
- Workshop Day 15th June
Invited Speakers
Organisers

Ernesto Coto
Oxford
Ivan Laptev
INRIA
Rahul Sukthankar
Google/CMU
Bernard Ghanem
KAUST
Andrew Zisserman
Oxford
Contact
Contact email for any queries relating to the workshop: albanie[AT]robots.ox.ac.uk






