Science Cast

Enhancing expressivity transfer in textless speech-to-speech translation

Jarod DuretOctober 12, 2023 5:54am

Views (394)
Comments (0)

Export Citation

Voice is AI-generated

Connected to paperThis paper is a preprint and has not been certified by peer review

Enhancing expressivity transfer in textless speech-to-speech translation

arXivPDFOctober 11, 2023 12:00am

Authors

Jarod Duret LIA, Benjamin O'Brien LIA, Yannick Estève LIA, Titouan Parcollet CAM

Abstract

Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and transferring expressivity accurately across different languages. Expressivity plays a vital role in conveying emotions, nuances, and cultural subtleties, thereby enhancing communication across diverse languages. To address this issue this study presents a novel method that operates at the discrete speech unit level and leverages multilingual emotion embeddings to capture language-agnostic information. Specifically, we demonstrate how these embeddings can be used to effectively predict the pitch and duration of speech units in the target language. Through objective and subjective experiments conducted on a French-to-English translation task, our findings highlight the superior expressivity transfer achieved by our approach compared to current state-of-the-art systems.

TwitterandLinkedIn

0 comments

Add comment

Enhancing expressivity transfer in textless speech-to-speech translation

Enhancing expressivity transfer in textless speech-to-speech translation

AI-powered Paper ChatBeta

AI-powered Paper ChatBeta

0 comments