A variety of degradations, including noise, blur, low resolution, and environmental aberrations, frequently impair digital photos taken in real world settings. Both humans and automated systems find it challenging to understand images due to these problems, which also lower image quality. However, it takes a lot of time and is not feasible to hand create descriptions for big image collections. This research introduces an integrated system that integrates automated image understanding and picture improvement to address these issues. To enhance the visual quality of deteriorated images, the proposed method makes use of a diffusion-based image restoration technique. To eliminate noise, lessen blur, and restore image details, a pretrained Stable Diffusion x4 Upscaler is used, resulting in outputs that are clearer and of higher quality. A feature extraction module built on the VGG16 deep learning architecture processes the improved image after image restoration. A Convolutional Neural Network Long Short Term Memory (CNN-LSTM) caption generation model receives the extracted visual data and uses them to automatically produce relevant textual descriptions of the image content. Python and deep learning frameworks are used in the proposed system implementation, and it is deployed via an interactive web application built on Streamlit. Results from experiments show that combining diffusion-based restoration with image captioning enhances caption production efficiency and image quality.
Introduction
The proposed project, DenoiseCap, integrates diffusion-based image restoration with deep learning-based image caption generation to improve the understanding of degraded images. Images captured in real-world conditions often suffer from noise, blur, low resolution, and poor lighting, making them difficult for both humans and AI systems to interpret. Traditional restoration techniques have limitations in preserving important visual details, while modern diffusion models can effectively reconstruct high-quality images. Similarly, image captioning uses computer vision and natural language processing to automatically generate meaningful descriptions of images, supporting applications such as accessibility, surveillance, multimedia management, and image retrieval.
The proposed system first restores degraded images using a Stable Diffusion x4 Upscaler, then extracts visual features through VGG16, and finally generates captions using a CNN-LSTM model. The entire pipeline is deployed as a Streamlit web application, allowing users to upload degraded images, restore them, and obtain descriptive captions. Image restoration performance is evaluated using PSNR and SSIM metrics.
The literature review highlights recent advances in diffusion-based image restoration and deep learning-based image captioning, emphasizing the effectiveness of CNNs, BLIP, and LSTM models in improving image quality and caption accuracy.
Experimental results demonstrate that the proposed system effectively removes various image degradations, including Gaussian noise, Gaussian blur, salt-and-pepper noise, and motion blur. The restored images achieve improved PSNR and SSIM values and enable more accurate caption generation. Sample outputs show successful restoration of blurred images and the generation of meaningful captions, confirming that the integrated framework enhances both image quality and semantic understanding for real-world applications.
Conclusion
The proposed DenoiseCap architecture effectively combines deep learning-based image caption creation with diffusion-based image restoration into a single system for improved visual comprehension. By lowering noise and blur while maintaining crucial structural features, the image restoration module makes use of the Stable Diffusion x4 Upscaler. The VGG16 feature extraction model is then used to process the restored images, and a CNN-LSTM architecture produces context-aware and relevant captions. Compared to methods that directly analyze deteriorated images, the framework improves feature extraction and generates more accurate and descriptive captions by undertaking image restoration before caption generation. Through an easy-to-use and interactive user interface, the Streamlit based deployment further illustrates the practical usefulness of the suggested system.
Experimental evaluation confirms the effectiveness of the proposed approach in both image restoration and semantic image understanding. While the caption generating module created logical and pertinent descriptions of the image material, the restoration module produced visually improved images with increased structural quality. The system\'s overall accuracy of 89.74% shows that it can reliably improve images and automatically generate captions in a single pipeline. By integrating image restoration and caption generation into a cohesive solution, the suggested DenoiseCap framework advances computer vision and is appropriate for use in assistive technologies, surveillance, medical imaging, digital asset management, and intelligent multimedia systems.
References
[1] H. Wu, M. Zhang, L. He, J.-Z. Zhou, and J. Lv, “Enhancing Diffusion Model Stability for Image Restoration via Gradient Management,” arXiv preprint arXiv:2507.06656, 2025.
[2] Q. Liang, X. Zheng, K. Huang, Y. Zhang, J. Chen, and Y. Tian, “Event-Diffusion: Event-Based Image Reconstruction and Restoration with Diffusion Models,” in Proceedings of the ACM Multimedia Conference (MM), 2023.
[3] J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 6840–6851, 2020.
[4] V. Tejesh and M. Supriya, “Enhanced Image Captioning Using CNN and BLIP Models,” in Proceedings of the International Conference on Multi-Agent Systems for Collaborative Intelligence (ICMSCI), IEEE, 2025.
[5] S. Bhatt, A. Rai, P. Chauhan, A. Dua, A. Kumar, and R. Sharma, “Deep Fusion: A CNN-LSTM Image Caption Generator for Enhanced Visual Understanding,” 2023.
[6] O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and Tell: A Neural Image Caption Generator,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 2015, pp. 3156–3164.
[7] M. Hodosh, P. Young, and J. Hockenmaier, “Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics,” Journal of Artificial Intelligence Research, vol. 47, pp. 853–899, 2013. (Flickr8K Dataset)
[8] E. Agustsson and R. Timofte, “NTIRE 2017 Challenge on Single Image Super-Resolution: Dataset and Study,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017. (DIV2K Dataset)
[9] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis With Latent Diffusion Models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10684–10695.
[10] F. Chollet et al., “Keras,” 2015. [Online]. Available: https://keras.io
[11] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019.
[12] T. Streamlit Inc., “Streamlit Documentation,” 2025. [Online]. Available: https://streamlit.io