MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation (original) (raw)

View PDF

Abstract:We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages. It is fully transcribed and covers 6 English-to-X translation as well as 6 X-to-English translation directions. To the best of our knowledge, this is the first open benchmark for audio-visual speech-to-text translation and the largest open benchmark for multilingual audio-visual speech recognition. Our baseline results show that MuAViC is effective for building noise-robust speech recognition and translation models. We make the corpus available at this https URL.

Submission history

From: Changhan Wang [view email]
[v1] Wed, 1 Mar 2023 16:31:01 UTC (43 KB)
[v2] Tue, 7 Mar 2023 16:41:01 UTC (43 KB)