AVATAR submission to the Ego4D AV Transcription Challenge

11/18/2022
by   Paul Hongsuck Seo, et al.
6

In this report, we describe our submission to the Ego4D AudioVisual (AV) Speech Transcription Challenge 2022. Our pipeline is based on AVATAR, a state of the art encoder-decoder model for AV-ASR that performs early fusion of spectrograms and RGB images. We describe the datasets, experimental settings and ablations. Our final method achieves a WER of 68.40 on the challenge test set, outperforming the baseline by 43.7

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset