LIFI: Towards Linguistically Informed Frame Interpolation

التفاصيل البيبلوغرافية
العنوان: LIFI: Towards Linguistically Informed Frame Interpolation
المؤلفون: Rajiv Ratn Shah, Changyou Chen, Roger Zimmermann, Aradhya Neeraj Mathur, Devansh Batra, Yaman Kumar Singla
المصدر: ICASSP
بيانات النشر: IEEE, 2021.
سنة النشر: 2021
مصطلحات موضوعية: Signal processing, business.industry, Computer science, Deep learning, Speech recognition, ComputingMethodologies_IMAGEPROCESSINGANDCOMPUTERVISION, Viseme, Entertainment, Web traffic, Artificial intelligence, Motion interpolation, business, Set (psychology), Interpolation
الوصف: Here we explore the problem of speech video interpolation. With close to 70% of web traffic, such content today forms the primary form of online communication and entertainment. Despite high performance on conventional metrics like MSE, PSNR, and SSIM, we find that the state-of-the-art frame interpolation models fail to produce faithful speech interpolation. For instance, we observe the lips stay static while the person is still speaking for most interpolated frames. With this motivation, using the information of words, sub-words, and visemes, we provide a new set of linguistically informed metrics targeted explicitly to the problem of speech video interpolation. We release several datasets to test video interpolation models of their speech understanding. We also design linguistically informed deep learning video interpolation algorithms to generate the missing frames.
URL الوصول: https://explore.openaire.eu/search/publication?articleId=doi_________::e6ca9b5bf00ea03126c5444537b21250
https://doi.org/10.1109/icassp39728.2021.9413998
حقوق: OPEN
رقم الأكسشن: edsair.doi...........e6ca9b5bf00ea03126c5444537b21250
قاعدة البيانات: OpenAIRE