Focus: The focus of this workshop is the task of searching for content within a corpus of videos using natural language queries.

Background: Convolutional neural networks have yielded unprecedented progress on a wide range of image-centric benchmarks, driven through a combination of well-annotated datasets and end-to-end training. However, naively extending this approach from images to higher-level video understanding tasks quickly becomes prohibitive with respect to the computation and data annotation required to jointly train multi-modal high-capacity models. An attractive alternative is to repurpose collections of existing pretrained models as "experts", offering representations which have been specialised for semantically relevant machine perception tasks. In addition to efficacy, this approach offers a second key advantage---it encourages researchers without access to industrial computing clusters to contribute towards questions of fundamental importance to video understanding: How should temporal information be used to maximum effect? How best to exploit complementary and redundant signals across different modalities? How can models be designed that function robustly across different video domains? To stimulate research into these questions, we are hosting a challenge that focuses on learning from videos and language with experts: making available a diverse collection of carefully curated visual and audio pre-extracted features across a set of five influential video datasets as part of a "pentathlon" of video understanding.

Schedule

This was a virtual half-day workshop that took place as part of CVPR 2020 on Monday 15th June 2020.
A video of the event can be found here.
14:00 - 14:10 Welcome (slides)
14:10 - 14:40 Invited Keynote (slides) Chen Sun (Google)

14:40 - 14:48 (Invited oral): Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning (paper, slides) Shizhe Chen
14:49 - 14:57 (Invited oral): ActBERT: Learning Global-Local Video-Text Representations (paper, slides) Linchao Zhu
14:58 - 15:06 (Invited oral): Visual Grounding in Video for Unsupervised Word Translation (paper, slides) Jean-Baptiste Alayrac
15:06 - 15:15 Virtual coffee break
15:15 - 15:45 Invited Keynote (slides) Carl Vondrick (Columbia University)
15:45 - 15:50 Challenge results (slides)
15:50 - 15:56 (1st place): MMT (technical report, slides) Valentin Gabeur 15:56 - 16:02 (2nd place): cszhe (technical report, slides) Shizhe Chen
16:02 - 16:08 (3rd place): LEgGOdt (technical report, slides) Kaixu Cui
16:10 - 16:40 Virtual coffee break
16:40 - 16:48 (Invited oral): Action Modifiers: Learning from Adverbs in Instructional Videos (paper, slides) Hazel Doughty
16:49 - 16:57 (Invited oral): Condensed Movies: Story Based Retrieval with Contextual Embeddings (paper, slides) Max Bain
16:57 - 17:10 Virtual coffee break
17:10 - 17:40 Invited Keynote (slides) Anna Rohrbach (UC Berkeley)
17:40 - 17:45 Closing remarks (slides)

Challenge: The Video Pentathlon

Congratulations to everyone who took part in the challenge. The top ranked results and technical reports are provided below. For a detailed breakdown of the scores, see here.
Ranking Team name Peantathlon score
1st place MMT 2511.43 technical report
2nd place cszhe 2448.56 technical report
3nd place LEgGOdt 1895.01 technical report
4th place haoxiaoshuai 1496.98 technical report
Please see challenge page for details about how the challenge was run.

Dates

Challenge track

  • Challenge Opens 9th April
  • Test set Release 9th May
  • Challenge Closes 2nd June
    23:59 UTC 4th June
  • Report Deadline 23:59 UTC 9th June
  • Workshop Day 15th June

Invited Speakers

Organisers

The organisers would like to thank Erika Lu and Robert McCraith for their assistance in helping to ensure a smooth runnning the workshop.

Contact

Contact email for any queries relating to the workshop: albanie[AT]robots.ox.ac.uk