Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models

التفاصيل البيبلوغرافية
العنوان: Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
المؤلفون: Lu, Tianyi, Zhang, Xing, Gu, Jiaxi, Xu, Hang, Pei, Renjing, Xu, Songcen, Wu, Zuxuan
سنة النشر: 2023
المجموعة: Computer Science
مصطلحات موضوعية: Computer Science - Computer Vision and Pattern Recognition, Computer Science - Artificial Intelligence
الوصف: Latent Diffusion Models (LDMs) are renowned for their powerful capabilities in image and video synthesis. Yet, video editing methods suffer from insufficient pre-training data or video-by-video re-training cost. In addressing this gap, we propose FLDM (Fused Latent Diffusion Model), a training-free framework to achieve text-guided video editing by applying off-the-shelf image editing methods in video LDMs. Specifically, FLDM fuses latents from an image LDM and an video LDM during the denoising process. In this way, temporal consistency can be kept with video LDM while high-fidelity from the image LDM can also be exploited. Meanwhile, FLDM possesses high flexibility since both image LDM and video LDM can be replaced so advanced image editing methods such as InstructPix2Pix and ControlNet can be exploited. To the best of our knowledge, FLDM is the first method to adapt off-the-shelf image editing methods into video LDMs for video editing. Extensive quantitative and qualitative experiments demonstrate that FLDM can improve the textual alignment and temporal consistency of edited videos.
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2310.16400
رقم الأكسشن: edsarx.2310.16400
قاعدة البيانات: arXiv