Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Mallikarjunarao Sunke, Srikanth Gudi, Sriharsha Gudi
DOI Link: https://doi.org/10.22214/ijraset.2026.84791
Certificate: View Certificate
Thanks to the development of basic models and the high quality of the data, the emergence of AI-generated content has accelerated. Despite its incredible success, there are still challenges that are yet to be addressed in AI content generation, such as the processing of long-trail information, maintaining up-to-date knowledge, addressing high inference and training expenses, and addressing data leakage. The shift to address those challenges has been called Retrieval Augmented Generation (RAG). RAG has brought the process of gathering information, which improves the process of data generation by recovering relevant information from available data sources, resulting in robustness and accuracy. RAG has become the foundation of Natural Language Processing (NLP) to effectively fill the gap between factual accuracy of knowledge and fluency of “Large Language Models (LLMs)”. This study traces the root of RAG from its beginnings as a framework for knowledge-based tasks to its current state as an agentic, modular and complex runtime of knowledge. This study is an in-depth analysis of the evolution of Naïve RAG to Modular and Advanced RAG models, and the introduction of new innovations, such as self-reflection, dense vector recovery, and the use of different models. They are then examined to provide detailed feedback on how to make RAG truly dynamic and usable as a verifier when applying them to organizations.
The text provides a comprehensive overview of Retrieval-Augmented Generation (RAG) as an important development for improving the capabilities and reliability of Large Language Models (LLMs). LLMs such as GPT, PaLM, and LLaMA can generate highly human-like text, but they depend heavily on static training data. This creates problems such as outdated information, misinformation, hallucinations, lack of real-time knowledge, and difficulty maintaining factual consistency.
RAG addresses these limitations by combining an LLM with an external retrieval system. Instead of relying only on information stored in the model's parameters, RAG retrieves relevant information from databases, search engines, documents, knowledge bases, or knowledge graphs and provides it to the LLM during generation. This can make responses more current, relevant, factual, transparent, and verifiable, without requiring the entire model to be retrained whenever new information becomes available.
RAG has two fundamental processes:
Advanced techniques such as hybrid search, query rewriting, re-ranking, knowledge distillation, contrastive learning, reinforcement learning, and attention-based fusion can further improve retrieval quality.
The text identifies four major stages in the evolution of RAG:
RAG has applications across many academic and professional domains, particularly where information changes frequently or accuracy is important. These include:
Although RAG improves LLM performance, it also introduces several challenges:
RAG has been bred to be bigger, smarter and more realistic. RAG has solved the biggest problem with LLM reliability – that of pipelines. RAG can be used in conjunction with real-time streams of data and long-context windows, and this is more likely to yield an intelligence and active memory, capable of navigating large swathes of complexity and intellectualism with unexpected accuracy. From Naïve RAG to smart Multi-agent RAG, the RAG has been matured significantly by 2025. These approaches are more than just concepts of research. They play a key role in the production-level AI programmes. When creating a virtual assistant, chatbot or knowledge engine, it\'s crucial to select the right RAG framework to address the needs of the organization. This study provides comprehensive knowledge about RAG framework when transforming AI models, with special emphasis on improvements and applications. This study has summarized and organized the basic changes in RAG, and provided insight into interaction among the generators and retrievers. This study offers possible practical tips on RAG for various tasks and modalities. Lastly, existing RAG benchmarks are also introduced with some limitations of RAG and highlighted with some promising areas of research.
[1] Husain, H., Wu, H. H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436. [2] Berant, J., Chou, A., Frostig, R., & Liang, P. (2013, October). Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 conference on empirical methods in natural language processing (pp. 1533-1544). [3] Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., ... & Fung, P. (2023, November). A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. In Proceedings of the 13th international joint conference on natural language processing and the 3rd conference of the asia-pacific chapter of the association for computational linguistics (volume 1: Long papers) (pp. 675-718). [4] Wang, Y., Lipka, N., Rossi, R. A., Siu, A., Zhang, R., & Derr, T. (2024, March). Knowledge graph prompting for multi-document question answering. In Proceedings of the AAAI conference on artificial intelligence (Vol. 38, No. 17, pp. 19206-19214). [5] Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., ... & Kaplan, J. (2022). Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073. [6] Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., ... & Yih, W. T. (2020, November). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) (pp. 6769-6781). [7] Saha, S., Junaed, J. A., Saleki, M., Sharma, A. S., Rifat, M. R., Rahouti, M., ... & Amin, M. R. (2023, December). Vio-lens: A novel dataset of annotated social network posts leading to different forms of communal violence and its evaluation. In Proceedings of the First Workshop on Bangla Language Processing (BLP-2023) (pp. 72-84). [8] Levesque, H. J. (1984, August). A logic of implicit and explicit belief. In AAAI (Vol. 84, pp. 198-202). [9] Saha, A., Pahuja, V., Khapra, M., Sankaranarayanan, K., & Chandar, S. (2018, April). Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32, No. 1). [10] NVIDIA (2025). Spectrum-X: End-to-End Networking for AI and High-Performance Computing. Available at https://www.nvidia.com/en-us/networking/spectrumx/. [11] Ma, L., Zhang, R., Han, Y., Yu, S., Wang, Z., Ning, Z., ... & Lu, C. T. (2023). A comprehensive survey on vector database: Storage and retrieval technique, challenge. arXiv preprint arXiv:2310.11703. [12] Neha, F., Bhati, D., Shukla, D. K., Guercio, A., & Ward, B. (2025, January). Exploring AI text generation, retrieval-augmented generation, and detection technologies: A comprehensive overview. In 2025 IEEE 15th Annual Computing and Communication Workshop and Conference (CCWC) (pp. 00633-00639). IEEE. [13] Han, B., Susnjak, T., & Mathrani, A. (2024). Automating systematic literature reviews with retrieval-augmented generation: A comprehensive overview. Applied Sciences, 14(19), 9103. [14] Parekh, K. V., Saxena, N., & Ansari, M. A. (2025, June). A Comparative Study of Retrieval-Augmented Generation (RAG) Chatbots. In 2025 International Conference Automatics, Robotics and Artificial Intelligence (ICARAI) (pp. 1-6). IEEE. [15] Singh, A., Ehtesham, A., Kumar, S., Khoei, T. T., & Vasilakos, A. V. (2025). Agentic retrieval-augmented generation: A survey on agentic rag. arXiv preprint arXiv:2501.09136. [16] Ramdurai, B. (2025). Large language models (LLMs), retrieval-augmented generation (RAG) systems, and convolutional neural networks (CNNs) in application systems. International Journal of Marketing and Technology, 15(01), 2249-1058. [17] FEZARI, M., & Al-Dahoud, A. (2025). The Evolution of Retrieval-Augmented Generation (RAG) in AI. [18] Kalotra, S. (2025). Trends in Active Retrieval Augmented Generation. Signity. Available at https://www.signitysolutions.com/blog/trends-in-active-retrieval-augmented-generation [19] Dogan, E. (2024). RAG, or Retrieval Augmented Generation: Revolutionizing AI in 2025. Glean. Available at https://www.glean.com/blog/rag-retrieval-augmented-generation [20] Naresh (2025). Beyond Vanilla RAG: The 7 Modern RAG Architectures Every AI Engineer Must Know. Available at https://dev.to/naresh_007/beyond-vanilla-rag-the-7-modern-rag-architectures-every-ai-engineer-must-know-4l0c [21] Karlsson, D. (2025). Evaluating modular RAG with reasoning models. Available at https://www.kapa.ai/blog/evaluating-modular-rag-with-reasoning-models [22] Naminas, K. (2025). RAG Evaluation: Metrics and Benchmarks for Enterprise AI Systems. Label Your Data. Available at https://labelyourdata.com/articles/llm-fine-tuning/rag-evaluation [23] Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., ... & Clark, P. (2023). Self-refine: Iterative refinement with self-feedback. Advances in neural information processing systems, 36, 46534-46594. [24] Huang, J., & Chang, K. C. C. (2023, July). Towards reasoning in large language models: A survey. In Findings of the association for computational linguistics: ACL 2023 (pp. 1049-1065). [25] Saad-Falcon, J., Khattab, O., Potts, C., & Zaharia, M. (2024, June). Ares: An automated evaluation framework for retrieval-augmented generation systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 338-354). [26] Zhang, J. (2023). Graph-toolformer: To empower llms with graph reasoning ability via prompt augmented by chatgpt. arXiv preprint arXiv:2304.11116. [27] Jiang, X., Zhang, R., Xu, Y., Qiu, R., Fang, Y., Wang, Z., ... & Wang, Y. (2025, July). HyKGE: A hypothesis knowledge graph enhanced RAG framework for accurate and reliable medical LLMs responses. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 11836-11856). [28] Ye, D., Lin, Y., Du, J., Liu, Z., Li, P., Sun, M., & Liu, Z. (2020, November). Coreferential reasoning learning for language representation. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) (pp. 7170-7186). [29] Zeng, H. (2023). Measuring massive multitask chinese understanding. arXiv preprint arXiv:2304.12986. [30] Jiang, H., Wu, Q., Lin, C. Y., Yang, Y., & Qiu, L. (2023, December). Llmlingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 13358-13376). [31] Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., ... & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997. [32] Yang, A., Nagrani, A., Seo, P. H., Miech, A., Pont-Tuset, J., Laptev, I., ... & Schmid, C. (2023). Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 10714-10726). [33] Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., & Hajishirzi, H. (2023, July). When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) (pp. 9802-9822). [34] Xia, M., Huang, G., Liu, L., & Shi, S. (2019, July). Graph based translation memory for neural machine translation. In Proceedings of the AAAI conference on artificial intelligence (Vol. 33, No. 01, pp. 7297-7304). [35] Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., ... & Wang, T. (2016). MS MARCO: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268. [36] Seo, M., Baek, J., Thorne, J., & Hwang, S. J. (2024). Retrieval-augmented data augmentation for low-resource domain tasks. arXiv preprint arXiv:2402.13482. [37] Gao, Y., Xiong, Y., Wang, M., & Wang, H. (2024). Modular rag: Transforming rag systems into lego-like reconfigurable frameworks. arXiv preprint arXiv:2407.21059. [38] Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., ... & Grave, E. (2023). Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research, 24(251), 1-43. [39] Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., ... & Sifre, L. (2022, June). Improving language models by retrieving from trillions of tokens. In International conference on machine learning (pp. 2206-2240). PMLR. [40] Wang, C., Long, Q., Xiao, M., Cai, X., Wu, C., Meng, Z., ... & Zhou, Y. (2024). Biorag: A rag-llm framework for biological question reasoning. arXiv preprint arXiv:2408.01107. [41] KHAPRA, C. (2024). A survey on end-to-end speech recognition systems. INTERNATIONAL JOURNAL OF COMMUNICATION, 5(2), 122-127. [42] Shukla, S., Mishra, S., Singh, J., Mishra, K., Khapra, C., & Morey, P. (2024, August). Deep Belief Networks for Unsupervised Feature Learning and Improved Image Recognition. In International Conference on Artificial Intelligence on Textile and Apparel (pp. 133-145). Singapore: Springer Nature Singapore. [43] Khapra, P. (2022). Evolution of cybercrime and deepfakes–Exploring intervention strategies of international organizations against AI threats. NeuroQuantology, 20(22), 1425-1434. [44] Barnett, S., Kurniawan, S., Thudumu, S., Brannelly, Z., & Abdelrazek, M. (2024, April). Seven failure points when engineering a retrieval augmented generation system. In Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering-Software Engineering for AI (pp. 194-199). [45] Cuconasu, F., Trappolini, G., Siciliano, F., Filice, S., Campagnano, C., Maarek, Y., ... & Silvestri, F. (2024, July). The power of noise: Redefining retrieval for rag systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 719-729). [46] Qiu, L., Shaw, P., Pasupat, P., Shi, T., Herzig, J., Pitler, E., ... & Toutanova, K. (2022, December). Evaluating the impact of model scale for compositional generalization in semantic parsing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 9157-9179). [47] Zhang, H., Zhao, P., Miao, X., Shao, Y., Liu, Z., Yang, T., & Cui, B. (2023). Experimental analysis of large-scale learnable vector storage compression. arXiv preprint arXiv:2311.15578. [48] Aksitov, R., Chang, C. C., Reitter, D., Shakeri, S., & Sung, Y. (2023). Characterizing attribution and fluency tradeoffs for retrieval-augmented large language models. arXiv preprint arXiv:2302.05578. [49] Liu, J. (2022). LlamaIndex [Computer software]. https://doi.org/10.5281/zenodo.1234 [50] Jiang, W., Zhang, S., Han, B., Wang, J., Wang, B., & Kraska, T. (2024). Piperag: Fast retrieval-augmented generation via algorithm-system co-design. arXiv preprint arXiv:2403.05676. [51] Hu, Z., Murthy, V., Pan, Z., Li, W., Fang, X., Ding, Y., & Wang, Y. (2025, October). Hedrarag: Co-optimizing generation and retrieval for heterogeneous rag workflows. In Proceedings of the ACM SIGOPS 31st symposium on operating systems principles (pp. 623-638). [52] Krishnan, N. (2025). Ai agents: Evolution, architecture, and real-world applications. arXiv preprint arXiv:2503.12687. [53] Wu, P., Zhang, M., Wan, K., Zhao, W., He, K., Du, X., & Chen, Z. (2025). Hiprag: hierarchical process rewards for efficient agentic retrieval augmented generation. arXiv preprint arXiv:2510.07794. [54] Fu, Y., Peng, H., Sabharwal, A., Clark, P., & Khot, T. (2022). Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
Copyright © 2026 Mallikarjunarao Sunke, Srikanth Gudi, Sriharsha Gudi. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84791
Publish Date : 2026-09-02
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
