Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Neeraja R A, Sherin Shaji, Devende M R , Ajmalsha S R , Mrs. Viji C
DOI Link: https://doi.org/10.22214/ijraset.2026.85064
Certificate: View Certificate
Legal artificial intelligence (AI) systems increasingly draw on retrieval-augmented generation (RAG), transformer-based embeddings, multi-stage structured reasoning, and ensemble learning to support tasks such as legal question answering, legal judgment prediction, and multilingual outcome prediction. Hallucination, limited domain grounding, and uneven performance across languages and legal systems remain open problems. This survey synthesizes four recent peer-reviewed studies used as core sources: a recursive-feedback retrieval-augmented generation framework for legal question answering (LQ-RAG); a four-stage large language model (LLM) framework for legal judgment prediction combining knowledge infusion, case-law retrieval, multi-hop reasoning, and generative judgment synthesis (LegalReasoner); a retrieval-augmented generation system evaluated across six embedding-LLM pairings on a small corpus of Indian statutory documents; and a stacked-ensemble system for multilingual legal outcome prediction from case summaries in English, Telugu, Tamil, and Kannada. Rather than treating these four studies as an exhaustive literature base, we use them as core studies and interpret their reported findings critically against their own stated methodology, sample sizes, evaluation metrics, and self-identified limitations, supplementing the discussion with the related work each study itself cites. Across the reviewed studies, we find scope-limited evidence that retrieval grounding, fine-tuned or domain-adapted embeddings, and stacked ensembles improved reported performance relative to the corresponding baselines evaluated within those studies; these improvements are dataset-, language-, and metric-specific and should not be read as a general ranking across architectures or studies. We identify recurring gaps, including small or proof-of-concept datasets, reliance on a proprietary LLM as an automated evaluator, absence of legal-domain-expert human evaluation, narrow language and jurisdiction coverage, and the broader challenge of temporal drift in legal knowledge, and we outline evidence-based directions for future work. This review is presented as a focused technical survey rather than a systematic literature review, and its conclusions are bounded accordingly.
The text presents a technical survey of AI and NLP methods for legal decision-making, focusing on how modern systems can improve legal information retrieval, reasoning, and prediction while reducing hallucinations and reliability problems.
Legal AI must accurately retrieve statutes and precedents, apply legal reasoning, and provide transparent justifications. However, large language models (LLMs) can generate hallucinated or fabricated legal content, including nonexistent cases and incorrect citations. Such errors are particularly serious in legal practice because they can undermine the validity of legal arguments and decisions.
The survey examines four major technical approaches for improving legal AI:
Retrieval-Augmented Generation (RAG): Retrieves relevant statutes, regulations, or case law before generating an answer, helping ground responses in external evidence.
Multi-stage reasoning: Breaks legal prediction into stages such as legal knowledge integration, case retrieval, multi-hop reasoning, and final judgment synthesis.
Domain-specific embeddings: Uses transformer models adapted to legal terminology, jurisdictions, or languages to improve representation of legal documents.
Ensemble and stacked models: Combines multiple predictive models to improve robustness and reduce overfitting.
The survey emphasizes that none of these methods completely eliminates hallucination or guarantees legal correctness.
The work is a focused technical survey rather than a systematic literature review. It analyzes four peer-reviewed core studies covering:
Retrieval-augmented legal question answering.
Multi-stage LLM-based legal judgment prediction.
RAG for reducing hallucinations in legal guidance.
Multilingual legal outcome prediction using transformer embeddings and stacked models.
Additional papers are used mainly as contextual references from the four core studies. Because the studies use different datasets, jurisdictions, languages, tasks, and evaluation metrics, their results cannot be directly ranked or treated as interchangeable.
Legal-domain LLMs
Models such as Legal-BERT, LawGPT, Lawyer LLaMA, SaulLM, and other domain-adapted systems attempt to improve understanding of legal terminology and reasoning. However, domain adaptation does not fully solve problems such as hallucination, bias, incomplete training data, and knowledge cutoffs.
Retrieval-augmented legal systems
RAG systems retrieve relevant legal documents and provide them to the LLM as supporting evidence. The reviewed studies show that retrieval can improve grounding, but performance depends heavily on the quality, completeness, relevance, and authority of the retrieved documents.
Legal judgment prediction
Legal judgment prediction has evolved from conventional machine-learning models to neural networks, attention mechanisms, and transformer architectures. Modern approaches attempt to represent facts, statutes, precedCaseLawBERT, mBERT, XLM-R, and IndicBERT are designed to represent legal or multilingual text. Their effectiveness varies according to the language, legal domain explores structured reasoning techniques. However, automated metrics such as BLEU, ROUGE, and BERTScore do not necessarilyents, and relationships between cases more effectively.
Domain-specific and multilingual embeddings
Models such as Legal-BERT, InLegalBERT, InCaseLawBERT, mBERT, XLM-R, and IndicBERT are designed to represent legal or multilingual text. Their effectiveness varies according to the language, legal domain, dataset, task, and evaluation method.
Reasoning and evaluation
Legal reasoning requires connecting facts, statutes, precedents, and multiple pieces of evidence. The literature therefore increasingly explores structured reasoning techniques. However, automated metrics such as BLEU, ROUGE, and BERTScore do not necessarily demonstrate actual legal correctness. Human evaluation by legal experts remains important.
Two principal RAG approaches are discussed:
A relatively lightweight system evaluated on 17 Indian ordinances and regulations (~297 pages) using different embedding–LLM combinations. One reported configuration, Llama 2 with all-mpnet-base-v2, achieved a BERTScore F1 of 56.87%.
Legal Query RAG (LQ-RAG), which uses hybrid retrieval, dense and lexical search, reranking, query processing, and recursive evaluation-feedback mechanisms. It reports improvements in retrieval metrics such as Hit Rate and MRR, but introduces greater computational and response-time costs.
Importantly, the studies use different experimental settings, so their results cannot be directly compared.
The survey discusses LegalReasoner, a four-stage representations, structured reasoning, and multiple predictive models. RAG improves access to authoritative information, domain-specific embeddings improve legal text representation, and legal judgment prediction framework based on LLaMA-2 (7B). Its stages include:
Legal knowledge infusion through contrastive learning.
Case-law retrieval using a Graph Attention Network.
Multi-hop reasoning using a transformer with hierarchical attention.
Judgment synthesis to produce the predicted legal outcome.
This architecture attempts to make legal reasoning more structured by separating retrieval, reasoning, and decision-making rather than relying on a single generation step.
The survey is organized around five questions:
RQ1: How effectively does retrieval reduce hallucination and improve grounding?
RQ2: How do multi-stage reasoning and knowledge integration improve legal judgment prediction?
RQ3: How do domain-specific and multilingual transformer representations affect legal prediction?
RQ4: What is the specific contribution of ensemble/stacked models?
RQ5: What limitations remain concerning datasets, multilinguality, evaluation, computational cost, fairness, and changes in law over time?
This survey has examined four recent, peer-reviewed studies spanning retrieval-augmented legal question answering [1], [3], multi-stage legal judgment prediction [2], and multilingual, ensemble-based legal outcome prediction [4], using them as core evidence rather than as an exhaustive account of legal AI research. Addressing the review questions posed in Section I: retrieval-based approaches are reported to improve grounding and relevance relative to un-grounded or less elaborate baselines, but residual failures in context relevance and groundedness persist even in the most sophisticated system reviewed, indicating that retrieval helps provide external grounding without resolving reliability outright (RQ1). Structured, multi-stage reasoning components are associated with improved legal judgment prediction performance in the study that evaluates their contributions through ablation. However, the reported results do not establish that any individual component is universally responsible for the overall improvement, and one aggregate accuracy figure stated in the study\'s abstract requires clarification (RQ2). Transformer-based and domain- or language-adapted representations are associated with improved performance relative to generic or non-adapted alternatives across the reviewed studies, but the size and consistency of this advantage depends on the specific dataset, language, task, and metric, and does not license a claim that any single embedding is universally superior (RQ3). Ensembling in the strict statistical sense, distinct from multi-stage pipelines and retrieval-generation architectures, is evaluated in one study using stacked classifiers and is reported to improve performance relative to individual classifiers on the study\'s legal dataset (RQ4). Finally, the reviewed literature leaves open a consistent set of limitations: dataset scale, reliance on proprietary and unaudited evaluators, absent legal-domain-expert human evaluation, narrow language and jurisdiction coverage, unaddressed temporal drift in legal knowledge, and an unresolved accuracy/computational-cost trade-off (RQ5). These observations are synthesized from four specific studies and the literature they themselves cite. This pattern is observed across the tasks and languages represented in the reviewed studies, although broader evaluation is required to determine whether it generalizes across legal domains, jurisdictions, and deployment settings. The four approaches should not be interpreted as directly ranked against one another because they address different tasks, languages, datasets, and experimental settings. Future work that closes the gaps identified in Section XIII-particularly multilingual retrieval-grounded question answering, human-expert evaluation, cross-jurisdictional robustness testing, and shared evaluation benchmarks - would strengthen the empirical basis for any future claim about which architectural strategies are most effective for legal AI, and under what conditions.
[1] R. S. M. Wahidur, S. Kim, H. Choi, D. S. Bhatti, and H.-N. Lee, \"Legal Query RAG,\" IEEE Access, vol. 13, pp. 36978-36994, 2025, doi: 10.1109/ACCESS.2025.3542125. [2] X. Wang, X. Zhang, V. Hoo, Z. Shao, and X. Zhang, \"LegalReasoner: A multi-stage framework for legal judgment prediction via large language models and knowledge integration,\" IEEE Access, vol. 12, pp. 166843-166854, 2024, doi: 10.1109/ACCESS.2024.3496666. [3] S. Gupta, V. Agrawal, Y. Negi, S. Karunanithi, and A. Balakrishnan, \"Reducing hallucinations in legal AI: A retrieval augmented generation-based model for accurate legal guidance,\" IEEE Access, early access, 2026, doi: 10.1109/ACCESS.2026.3675624. [4] P. Prabhakar, T. Gaddam, D. Chennupalle, and P. B. Pati, \"Transformer-based embeddings and stacking classifiers for judgment outcome prediction in multilingual legal texts,\" Scientific Reports, vol. 16, art. 18305, 2026, doi: 10.1038/s41598-026-47652-x. [5] P. Lewis et al., \"Retrieval-augmented generation for knowledge-intensive NLP tasks,\" in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2020, pp. 9459-9474. [6] S. H. Varshini, G. Sarayu Varma, P. Prabhakar, and P. B. Pati, \"AI in the courtroom: Enhancing legal decision-making through predictive modelling,\" Journal of Information & Knowledge Management, 2024. [7] V. Malik et al., \"ILDC for CJPE: Indian legal documents corpus for court judgment prediction and explanation,\" arXiv:2105.13562, 2021. [8] I. Chalkidis, I. Androutsopoulos, and N. Aletras, \"Neural legal judgment prediction in English,\" in Proc. 57th Annu. Meeting Assoc. Comput. Linguistics, 2019, pp. 4317-4323. [9] C. Xiao et al., \"CAIL2018: A large-scale legal dataset for judgment prediction,\" arXiv:1807.02478, 2018. [10] I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, \"LEGAL-BERT: The muppets straight out of law school,\" in Findings Assoc. Comput. Linguistics: EMNLP, 2020, pp. 2898-2904. [11] J. E. Hu et al., \"LoRA: Low-rank adaptation of large language models,\" in Proc. Int. Conf. Learn. Represent., 2021, pp. 1-53. [12] H. Touvron et al., \"Llama 2: Open foundation and fine-tuned chat models,\" arXiv:2307.09288, 2023. [13] AI@Meta, \"Llama 3 Model Card,\" 2024. [Online]. Available: https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md. [14] A. Vaswani et al., \"Attention is all you need,\" in Proc. Adv. Neural Inf. Process. Syst., 2017, pp. 1-11.
Copyright © 2026 Neeraja R A, Sherin Shaji, Devende M R , Ajmalsha S R , Mrs. Viji C. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET85064
Publish Date : 2026-10-05
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
