Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Adepegba Akindayomi Akintade
DOI Link: https://doi.org/10.22214/ijraset.2026.84530
Certificate: View Certificate
Artificial intelligence for information technology operations has progressed from alert correlation and anomaly detection to root-cause analysis, mitigation recommendation, and agents that can invoke infrastructure tools. This progression creates an assurance problem: analytical performance does not establish authority to change a live cloud system. This critical narrative review examines the evidence required before AI output may influence or execute failure-management action in mission-critical cloud infrastructure. Searches through 24 July 2026 covered scholarly databases, major systems venues, standards sources, and authoritative production reports. Forty-four sources were coded by operational task, evidence setting, authority, safety control, rollback, and outcome verification. Production evidence is substantial for detection, triage, diagnosis, and several narrowly bounded mitigation systems, but remains weak for general-purpose autonomous action. Recent agentic AIOps surveys emphasize contracts, bounded tools, canary deployment, and rollback; the unresolved need is an assessment method that separates what a system can infer, what it may do, and what evidence shows that the action remained safe. The proposed Operational Decision-Readiness and Verification framework represents a deployment claim through analytical capability, operational authority, and assurance maturity. Twelve domains use explicit 0-3 evidence anchors, while endpoint validity, evidence integrity, intervention risk, authorization, reversibility, and recovery verification operate as non-compensatory gates. Worked assessments of six systems illustrate distinctions among analytical support, testbed execution, and narrow production autonomy. The framework is conceptual and requires prospective and inter-rater validation; it structures an action-specific readiness case rather than certifying a model or product.
The central thesis is that technical capability, operational authority, and assurance are different things. A system may detect incidents accurately, diagnose causes plausibly, or recommend effective mitigations without having sufficient evidence to be trusted with infrastructure-changing actions.
The paper therefore argues against a single linear “AI maturity” ladder. Instead, readiness should be represented using three independent dimensions:
The proposed ODRV framework turns these ideas into an action-specific readiness assessment. It introduces:
A particularly important principle is that strong evidence in one dimension cannot compensate for a fundamental deficiency in another. For example, excellent detection accuracy does not compensate for inadequate rollback, excessive privileges, poor evidence integrity, or lack of service-level verification.
The review identifies a clear gradient in the literature:
Detection and triage: relatively mature, including systems such as Warden and COMET.
Diagnosis/RCA: increasingly capable, with systems such as RCACopilot and LLM-based incident-analysis approaches, but generally retaining humans in the decision loop.
Prediction: stronger evidence exists in narrowly defined failure domains, such as Narya's prediction of Azure host failures.
Automated remediation: credible but substantially narrower. RESIN is an important production example because it automatically detected, diagnosed, and mitigated memory leaks over an extended period.
The important conclusion is that the strongest production examples are not necessarily general-purpose autonomous agents. Their reliability comes partly from constraining the failure class, available actions, operational context, and verification conditions.
The paper also extends conventional AIOps reliability thinking by treating observability itself as an attack surface.
Telemetry can be:
For an LLM agent, manipulated telemetry can become an input that causes an inappropriate infrastructure action. Similarly, broad tool permissions can give an agent more authority than its natural-language instructions appear to imply.
Consequently, safe agent evaluation must examine not only whether the agent reaches the correct answer, but also:
evidence integrity → reasoning/decision → tool selection → authorization → execution trajectory → service outcome.
The review deliberately uses a critical narrative rather than systematic-review/meta-analysis design because the subject spans several research communities. It nevertheless attempts reproducibility through:
Importantly, the author acknowledges that the review is single-author and not independently dual-screened, and that ODRV ratings are illustrative rather than independent audits.
The strongest idea running through the paper can be expressed simply:
A model evaluation answers “Can it do the task?” A readiness assessment asks “Should it be allowed to do this action here, with this authority, based on this evidence?”
That distinction gives the paper a meaningful position relative to conventional AIOps surveys and recent agentic-operations work.
The section beginning “Failure pathways, observability, and operational context” appears to be moving toward an important part of the framework but ends mid-sentence at “The relevant evidence packag…”. This section would benefit from explicitly connecting observability to the ODRV domains—for example, evidence completeness, provenance/integrity, temporal validity, topology/context awareness, and service-level verification.
AI-enabled cloud failure management is moving toward agents that can investigate incidents and invoke operational tools. The decisive question is no longer whether such systems can act, but when the evidence justifies allowing a particular action in a particular environment. Existing AIOps reviews define tasks and methods, and recent agentic-operations surveys correctly emphasize contracts, bounded tools, gates, canary execution, and rollback. The remaining assessment problem is to separate analytical capability from permission and assurance. The revised ODRV framework represents readiness as an action-specific tuple of capability, authority, assurance, and twelve evidence domains. Six non-compensatory gates prevent strong performance in one area from masking a critical failure in endpoint validity, evidence integrity, intervention risk, authorization, reversibility, or recovery verification. Worked examples show that production detection, testbed execution, and narrow autonomous mitigation are different achievements and should be reported as such. ODRV is not a certification standard and should not be used as a simple total score. Its value is to make the readiness claim explicit, falsifiable, and revisable. The next step is prospective and inter-rater validation across organizations and failure classes. Until that evidence exists, the safest principle remains clear: model capability may support an operational decision, but only systems evidence can justify operational authority.
[1] Ahmed, T., Ghosh, S., Bansal, C., Zimmermann, T., Zhang, X., and Rajmohan, S. (2023). Recommending root-cause and mitigation steps for cloud incidents using large language models. In Proceedings of the 45th IEEE/ACM International Conference on Software Engineering, 1737-1749. https://doi.org/10.1109/ICSE48619.2023.00149 [2] Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., and Zimmermann, T. (2019). Software engineering for machine learning: A case study. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice, 291-300. https://doi.org/10.1109/ICSE-SEIP.2019.00042 [3] Avizienis, A., Laprie, J.-C., Randell, B., and Landwehr, C. (2004). Basic concepts and taxonomy of dependable and secure computing. IEEE Transactions on Dependable and Secure Computing, 1(1), 11-33. https://doi.org/10.1109/TDSC.2004.2 [4] Beyer, B., and Davidovic, S. (2018). Canary analysis service. ACM Queue, 16(5), 70-95. https://doi.org/10.1145/3291278.3291290 [5] Bilal, M., Crowcroft, J., Wang, R., Xu, X., and Dustdar, S. (2026). Large language models for agentic NetOps and AIOps: Architectures, evaluation, and safety. arXiv:2605.12729. https://doi.org/10.48550/arXiv.2605.12729 [6] Breck, E., Cai, S., Nielsen, E., Salib, M., and Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data, 1123-1132. https://doi.org/10.1109/BigData.2017.8258038 [7] Chen, Y., Shetty, M., Somashekar, G., Ma, M., Simmhan, Y., Mace, J., Bansal, C., Wang, R., and Rajmohan, S. (2025). AIOpsLab: A holistic framework to evaluate AI agents for enabling autonomous clouds. Proceedings of Machine Learning and Systems, 7. https://arxiv.org/abs/2501.06706 [8] Chen, Y., Yang, H., Zhang, C., Ma, M., Kang, Y., Li, Z., He, S., Zhang, X., Wang, R., Rajmohan, S., Lin, Q., and Zhang, D. (2024). Automatic root cause analysis via large language models for cloud incidents. In Proceedings of the Nineteenth European Conference on Computer Systems, 674-688. https://doi.org/10.1145/3627703.3650083 [9] Cheng, Q., Sahoo, D., Saha, A., Yang, W., Liu, C., Woo, G., Singh, M., Savarese, S., and Hoi, S. C. H. (2023). AI for IT operations on cloud platforms: Reviews, opportunities and challenges. arXiv:2304.04661. https://doi.org/10.48550/arXiv.2304.04661 [10] Dogga, P., Parayil, A., Ghosh, S., Bansal, C., Zhang, X., Mace, J., Nath, S., and Rajmohan, S. (2023). AutoARTS: Taxonomy, insights and tools for root cause labelling of incidents in Microsoft Azure. In 2023 USENIX Annual Technical Conference, 359-375. USENIX Association. [11] Gama, J., Zliobaite, I., Bifet, A., Pechenizkiy, M., and Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), Article 44. https://doi.org/10.1145/2523813 [12] Ghosh, S., Shetty, M., Bansal, C., and Nath, S. (2022). How to fight production incidents? An empirical study on a large-scale cloud service. In Proceedings of the 13th Symposium on Cloud Computing, 126-141. https://doi.org/10.1145/3542929.3563472 [13] Goel, D., Husain, F., Singh, A., Ghosh, S., Parayil, A., Bansal, C., Zhang, X., and Rajmohan, S. (2024). X-lifecycle learning for cloud incident management using large language models. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, 417-428. https://doi.org/10.1145/3663529.3663861 [14] Goel, H. (2026). Security risks in tool-enabled AI agents: A systematic analysis of privileged execution environments. arXiv:2605.09721. https://doi.org/10.48550/arXiv.2605.09721 [15] Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 1321-1330. [16] Henderson, T., Kondareddy, A., Azad, S., and Nickell, E. (2024). SafeRevert: When can breaking changes be automatically reverted? In 2024 IEEE Conference on Software Testing, Verification and Validation. IEEE. [17] Jiang, Y., Zhang, C., He, S., Yang, Z., Ma, M., Qin, S., Kang, Y., Dang, Y., Rajmohan, S., Lin, Q., and Zhang, D. (2024). Xpert: Empowering incident management with query recommendations via large language models. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, Article 92, 1-13. https://doi.org/10.1145/3597503.3639081 [18] Kephart, J. O., and Chess, D. M. (2003). The vision of autonomic computing. Computer, 36(1), 41-50. https://doi.org/10.1109/MC.2003.1160055 [19] Levy, S., Yao, R., Wu, Y., Dang, Y., Huang, P., Mu, Z., Zhao, P., Ramani, T., Govindaraju, N., Li, X., Lin, Q., Shafriri, G. L., and Chintalapati, M. (2020). Predictive and adaptive failure mitigation to avert production cloud VM interruptions. In 14th USENIX Symposium on Operating Systems Design and Implementation, 1155-1170. USENIX Association. [20] Li, L., Zhang, X., Zhao, X., Zhang, H., Kang, Y., Zhao, P., Qiao, B., He, S., Lee, P., Sun, J., Gao, F., Yang, L., Lin, Q., Rajmohan, S., Xu, Z., and Zhang, D. (2021). Fighting the fog of war: Automated incident detection for cloud systems. In 2021 USENIX Annual Technical Conference, 131-146. USENIX Association. [21] Li, P., Wang, S., Huang, Y., Shi, Y., Zhang, C., Li, Q., Lyu, Y., Shan, C., Li, F., Feng, C., Zhu, C., and Chen, L. (2026). AgentCanary: A security evaluation framework for autonomous AI agents in real executable environments. arXiv:2606.10484. https://doi.org/10.48550/arXiv.2606.10484 [22] Lou, C., Chen, C., Huang, P., Dang, Y., Qin, S., Yang, X., Li, X., Lin, Q., and Chintalapati, M. (2022). RESIN: A holistic service for dealing with memory leaks in production cloud infrastructure. In 16th USENIX Symposium on Operating Systems Design and Implementation, 109-125. USENIX Association. [23] Lyell, D., and Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of the American Medical Informatics Association, 24(2), 423-431. https://doi.org/10.1093/jamia/ocw105 [24] Namyar, P., Ghavidel, A., Crankshaw, D., Berger, D. S., Hsieh, K., Kandula, S., Govindan, R., and Arzani, B. (2025). Enhancing network failure mitigation with performance-aware ranking. In 22nd USENIX Symposium on Networked Systems Design and Implementation, 335-357. USENIX Association. [25] Notaro, P., Cardoso, J., and Gerndt, M. (2021). A survey of AIOps methods for failure management. ACM Transactions on Intelligent Systems and Technology, 12(6), Article 81. https://doi.org/10.1145/3483424 [26] Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., and Snoek, J. (2019). Can you trust your model\'s uncertainty? Evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems, 32. [27] Palatin, T., and Beyer, B. (2021). Prodspec and Annealing: Intent-based actuation in Google production. USENIX ;login:. [28] Parasuraman, R., Sheridan, T. B., and Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, 30(3), 286-297. https://doi.org/10.1109/3468.844354 [29] Pasquini, D., Kornaropoulos, E. M., Ateniese, G., Akgul, O., Theocharis, A., and Efstathopoulos, P. (2026). When AIOps become \"AI Oops\": Subverting LLM-driven IT operations via telemetry manipulation. 35th USENIX Security Symposium, accepted prepublication paper. https://www.usenix.org/conference/usenixsecurity26/presentation/pasquini [30] Rao, A., Keller, A., Kalra, N., Steed, R., Kwegyir-Aggrey, K., Klyman, K., Staheli, D., and Bergman, A. (2026). Challenges to the monitoring of deployed AI systems. NIST AI 800-4. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-4 [31] Remil, Y., Bendimerad, A., Mathonat, R., and Kaytoue, M. (2024). AIOps solutions for incident management: Technical guidelines and a comprehensive literature review. arXiv:2404.01363. https://doi.org/10.48550/arXiv.2404.01363 [32] Rouholamini, S. R., Mirabi, M., Farazkish, R., and Sahafi, A. (2024). Proactive self-healing techniques for cloud computing: A systematic review. Concurrency and Computation: Practice and Experience, 36(24), e8246. https://doi.org/10.1002/cpe.8246 [33] Roy, D., Zhang, X., Bhave, R., Bansal, C., Las-Casas, P., Fonseca, R., and Rajmohan, S. (2024). Exploring LLM-based agents for root cause analysis. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, 208-219. https://doi.org/10.1145/3663529.3663841 [34] Salehie, M., and Tahvildari, L. (2009). Self-adaptive software: Landscape and research challenges. ACM Transactions on Autonomous and Adaptive Systems, 4(2), Article 14. https://doi.org/10.1145/1516533.1516538 [35] Schneider, C., Barker, A. D., and Dobson, S. A. (2015). A survey of self-healing systems frameworks. Software: Practice and Experience, 45(10), 1375-1398. https://doi.org/10.1002/spe.2250 [36] Schulman, C., and Perot, E. (2018). Help protect your datacenters with safety constraints. Google Research. [37] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., and Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28. [38] Shetty, M., Chen, Y., Somashekar, G., Ma, M., Simmhan, Y., Zhang, X., Mace, J., Las-Casas, P., Gupta, S., Nath, S., Bansal, C., and Rajmohan, S. (2024). Building AI agents for autonomous clouds: Challenges and design principles. In Proceedings of the 15th ACM Symposium on Cloud Computing. [39] Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1 [40] Wang, Z., Li, J., Ma, M., Li, Z., Kang, Y., Zhang, C., Bansal, C., Chintalapati, M., Rajmohan, S., Lin, Q., Zhang, D., Pei, C., and Xie, G. (2024). Large language models can provide accurate and interpretable incident triage. In 2024 IEEE 35th International Symposium on Software Reliability Engineering, 523-534. https://doi.org/10.1109/ISSRE62328.2024.00056 [41] Zhang, L., Jia, T., Jia, M., Wu, Y., Liu, A., Yang, Y., Wu, Z., Hu, X., Yu, P. S., and Li, Y. (2026). A survey of AIOps in the era of large language models. ACM Computing Surveys, 58(2), Article 44. https://doi.org/10.1145/3746635 [42] Zhang, D., Zhang, X., Bansal, C., Las-Casas, P. H. B., Fonseca, R., and Rajmohan, S. (2023). PACE-LM: Prompting and augmentation for calibrated confidence estimation with GPT-4 in cloud incident root cause analysis. arXiv:2309.05833. https://doi.org/10.48550/arXiv.2309.05833 [43] Zhang, X., Ghosh, S., Bansal, C., Wang, R., Ma, M., Kang, Y., and Rajmohan, S. (2024). Automated root causing of cloud incidents using in-context learning with GPT-4. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, 266-277. https://doi.org/10.1145/3663529.3663846 [44] Zhai, E., Chen, A., Piskac, R., Balakrishnan, M., Tian, B., Song, B., and Zhang, H. (2020). Check before you change: Preventing correlated failures in service updates. In 17th USENIX Symposium on Networked Systems Design and Implementation, 575-589. USENIX Association.
Copyright © 2026 Adepegba Akindayomi Akintade. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84530
Publish Date : 2026-08-04
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
