Pegasus-v1 Technical Report

التفاصيل البيبلوغرافية
العنوان: Pegasus-v1 Technical Report
المؤلفون: Jung, Raehyuk, Go, Hyojun, Yi, Jaehyuk, Jang, Jiho, Kim, Daniel, Suh, Jay, Lee, Aiden, Han, Cooper, Lee, Jae, Kim, Jeff, Kim, Jin-Young, Kim, Junwan, Park, Kyle, Lee, Lucas, Ha, Mars, Seo, Minjoon, Jo, Abraham, Park, Ed, Kianinejad, Hassan, Kim, SJ, Moon, Tony, Jeong, Wade, Popescu, Andrei, Kim, Esther, Yoon, EK, Heo, Genie, Choi, Henry, Kang, Jenna, Han, Kevin, Seo, Noah, Nguyen, Sunny, Won, Ryan, Park, Yeonhoo, Giuliani, Anthony, Chung, Dave, Yoon, Hans, Le, James, Ahn, Jenny, Lee, June, Saini, Maninder, Sanders, Meredith, Lee, Soyoung, Kim, Sue, Couture, Travis
سنة النشر: 2024
المجموعة: Computer Science
مصطلحات موضوعية: Computer Science - Multimedia, Computer Science - Artificial Intelligence, Computer Science - Computation and Language, Computer Science - Computer Vision and Pattern Recognition
الوصف: This technical report introduces Pegasus-1, a multimodal language model specialized in video content understanding and interaction through natural language. Pegasus-1 is designed to address the unique challenges posed by video data, such as interpreting spatiotemporal information, to offer nuanced video content comprehension across various lengths. This technical report overviews Pegasus-1's architecture, training strategies, and its performance in benchmarks on video conversation, zero-shot video question answering, and video summarization. We also explore qualitative characteristics of Pegasus-1 , demonstrating its capabilities as well as its limitations, in order to provide readers a balanced view of its current state and its future direction.
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2404.14687
رقم الأكسشن: edsarx.2404.14687
قاعدة البيانات: arXiv