Recent advances in Generative Artificial Intelligence have significantly improved automatic multimedia content creation. This paper presents an AI-based Text-to-Video Generation framework that converts natural language prompts into short animated video sequences using diffusion and deep learning techniques. The proposed framework employs the Stable Diffusion v1.5 model for generating high-quality image frames and AnimateDiff for introducing smooth temporal motion while preserving scene consistency.
The DDIM scheduler is utilized to accelerate the denoising process and improve inference efficiency without compromising output quality. The generated frames are sequentially processed and assembled into MP4 videos using FFmpeg. The proposed system is implemented using Python, PyTorch, Hugging Face Diffusers, and Google Colab, providing a lightweight and cost-effective platform for video generation. Experimental evaluation demonstrates that the proposed framework generates visually realistic and semantically meaningful videos with improved motion continuity and reduced computational complexity compared to conventional frame-by-frame generation approaches. The proposed framework has potential applications in digital media, education, entertainment, advertising, animation, and content creation.
Introduction
This research presents an AI-Based Text-to-Video Generation System that converts natural language descriptions into short, coherent video clips using advanced diffusion models. With the growing demand for automated video creation in social media, education, marketing, and entertainment, the proposed system aims to reduce the time, cost, and expertise required for traditional video production.
The framework integrates Stable Diffusion v1.5, AnimateDiff, and the DDIM Scheduler. Stable Diffusion generates high-quality visual frames from text prompts, AnimateDiff ensures smooth motion and temporal consistency across frames, and the DDIM Scheduler speeds up the diffusion process while maintaining visual quality. Finally, FFmpeg combines the generated frames into an MP4 video. The system is implemented using Python, PyTorch, Hugging Face Diffusers, and Google Colab, making it accessible through a cloud-based environment.
The literature survey reviews major advancements in diffusion-based video generation, including Latent Diffusion Models, AnimateDiff, Video Diffusion Models, Imagen Video, Make-A-Video, and CogVideo. While these methods achieve impressive visual quality, they often suffer from high computational costs, temporal inconsistency, and difficulties generating long, coherent videos. The proposed framework addresses these challenges by combining efficient pre-trained models to improve motion continuity and computational efficiency.
The methodology follows an experimental approach, where users provide positive and negative text prompts that guide the generation process. The prompts are converted into text embeddings, latent noise is refined using diffusion-based denoising, and motion-aware modules generate temporally consistent video frames. The resulting frames are encoded into an MP4 video using FFmpeg. The system is evaluated using both qualitative measures (visual quality, prompt relevance, motion smoothness, and object consistency) and quantitative metrics, including Prompt Alignment Score, Temporal Consistency Score, Motion Smoothness Score, Frame Coherence Score, Video Generation Success Rate, and Average Video Generation Time.
Experimental results demonstrate that the proposed system successfully generates coherent videos from a variety of prompts, including human activities, animals, and animated scenes. A Gradio-based user interface enables interactive video generation with configurable parameters. The generated videos accurately match the input descriptions while maintaining stable object appearance, realistic backgrounds, and smooth motion, confirming the effectiveness of the proposed AI-based text-to-video generation framework.
Conclusion
This paper presented an AI-Based Text-to-Video Generation framework using Stable Diffusion v1.5, AnimateDiff, the Motion Adapter, and the DDIM Scheduler to generate short videos from natural language prompts. The proposed system successfully produced visually coherent videos with satisfactory semantic alignment, motion continuity, and temporal consistency. Experimental results demonstrated strong performance in prompt relevance, frame coherence, and motion smoothness while enabling efficient video generation on a cloud-based platform. Overall, the proposed framework provides an effective solution for automated text-to-video generation and has potential applications in education, digital content creation, storytelling, and multimedia production.
References
[1] T. Lee, S. Kwon, and T. Kim, “Grid Diffusion Models for Text-to-Video Generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 8734–8743.
[2] P. Esser, R. Rombach, and B. Ommer, “Diffusion Models for Video Generation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition Workshops (CVPRW), 2023.
[3] R. Rombach, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 10684–10695.
[4] J. An, S. Zhang, H. Yang, S. Gupta, J. Huang, J. Luo, and X. Yin, “Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation,” arXiv preprint arXiv:2304.08477, 2023.
[5] M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,” in Proc. IEEE/CVF Int. Conf. Computer Vision (ICCV), 2021, pp. 1728–1738.
[6] Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreiss, M. Aittala, T. Aila, S. Laine, and B. Catanzaro, “eDiffi: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers,” arXiv preprint arXiv:2211.01324, 2022.
[7] A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models,” arXiv preprint arXiv:2304.08818, 2023.
[8] T. Brooks, A. Holynski, and A. A. Efros, “InstructPix2Pix: Learning to Follow Image Editing Instructions,” Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2023.
[9] D. Ceylan, C.-H. P. Huang, and N. J. Mitra, “Pix2Video: Video Editing Using Image Diffusion,” arXiv preprint arXiv:2303.12688, 2023.
[10] H. Chen, M. Xia, Y. He, Y. Zhang, X. Cun, S. Yang, J. Xing, Y. Liu, Q. Chen, X. Wang, C. Weng, and Y. Shan, “VideoCrafter1: Open Diffusion Models for High-Quality Video Generation,” arXiv preprint, 2023.