Modern cloud-native software systems, including banking applications, cloud platforms, enterprise resource planning architectures, and real-time online services, continuously produce large volumes of operational data such as application logs, performance metrics, and system traces. Monitoring these systems using conventional rule-based or threshold-driven methods introduces operational bottlenecks because static thresholds fail to adapt to dynamic workloads and cannot identify unseen failure modes. Furthermore, relying on human operators for diagnosis and remediation increases the Mean Time to Recovery during outages.
This paper introduces an autonomous, closed-loop telemetry and incident management framework titled AI-Driven Self-Healing System for Real-Time Software Incident Management. The architecture implements memory-efficient stream processing by embedding probabilistic data structures—specifically Count-Min Sketches and Cuckoo Filters—within an in-memory Redis Stack layer. An automated stream aggregator compiles multi-route frequencies and HTTP status code distributions into normalized 18-dimensional feature vectors over 5-second sliding windows. An unsupervised machine learning engine combining an Isolation Forest algorithm with dynamic Z-score statistical profiling learns operational baselines and isolates multivariate anomalies in real time without requiring labeled historical datasets.
When anomalous deviations are detected, the system executes automated self-healing actions such as service restarts, dynamic resource allocation, and traffic management to restore operational stability. An interactive operations dashboard visualizes real-time core health, incident alerts, service mesh topologies, and forensic audit trails. Empirical benchmarking demonstrates that the platform sustains a live ingestion throughput of 124 messages per second across 24 monitored endpoints with negligible CPU saturation (0.0%) and bounded memory pressure, confirming its effectiveness for distributed cloud environments.
Introduction
The proposed system addresses reliability and monitoring challenges in distributed microservice architectures, where dynamic service interactions can produce cascading failures, traffic spikes, latency variations, and resource contention. Conventional monitoring approaches suffer from high storage and network overhead, static threshold limitations, insufficient multivariate analysis, and dependence on manual incident recovery.
The proposed framework establishes an autonomous monitoring and self-healing architecture consisting of five major components: edge telemetry ingestion, in-memory probabilistic aggregation, stream aggregation, unsupervised anomaly detection, and automated remediation. An edge reverse proxy captures inter-service HTTP traffic and records endpoint and HTTP status-code frequencies using Count-Min Sketches and Cuckoo Filters, thereby reducing the need to store raw textual logs.
Every five seconds, the system converts the collected telemetry into an 18-dimensional normalized feature fingerprint, comprising seven endpoint-frequency features and eleven HTTP status-code features. These fingerprints are stored in a Redis Stream and analyzed using an Isolation Forest model to identify multivariate anomalies. A dynamic Z-score mechanism additionally detects sudden deviations in request volumes.
When an anomaly is detected, the orchestration layer automatically initiates predefined recovery actions such as circuit breaking, rate limiting, or state resets. This creates a closed-loop process of telemetry collection → feature extraction → anomaly detection → automated remediation, with the objective of reducing monitoring overhead and Mean Time to Recovery while improving service availability.
The main novelty is therefore not simply anomaly detection, but the integration of memory-efficient telemetry collection, multivariate unsupervised detection, and autonomous recovery into a single self-healing monitoring loop.
Conclusion
This paper presented the design, implementation, and empirical validation of an autonomous, low-overhead monitoring and self-healing system for distributed microservice architectures. By combining in-memory probabilistic data structures with unsupervised machine learning, the framework decouples telemetry collection costs from incoming request volumes. Count-Min Sketches and Cuckoo Filters enable continuous request frequency and service mesh dependency tracking with constant-time updates and bounded memory consumption.
The unsupervised Isolation Forest and dynamic Z-score deviation profiling accurately detect operational anomalies in real time without requiring pre-labeled training datasets. When service degradations occur, the autonomous orchestration engine executes corrective actions, maintains an immutable forensic audit log, and synchronizes state updates in real time via WebSockets. Benchmarks demonstrate sustained throughput of 124 messages per second with 0.0% host CPU load, validating the framework as an efficient and scalable solution for self-healing software infrastructure.
References
[1] V. Chandola, A. Banerjee, and V. Kumar, \"Anomaly Detection: A Survey,\" ACM Computing Surveys, vol. 41, no. 3, pp. 1–58, 2009.
[2] M. Breunig, H. Kriegel, R. Ng, and J. Sander, \"LOF: Identifying Density-Based Local Outliers,\" in Proc. ACM SIGMOD Int. Conf. Management of Data, 2000, pp. 93–104.
[3] S. Hochreiter and J. Schmidhuber, \"Long Short-Term Memory,\" Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
[4] F. T. Liu, K. M. Ting, and Z. Zhou, \"Isolation Forest,\" in Proc. IEEE Int. Conf. Data Mining, 2008, pp. 413–422.
[5] G. Cormode and S. Muthukrishnan, \"An Improved Data Stream Summary: The Count-Min Sketch and its Applications,\" Journal of Algorithms, vol. 55, no. 1, pp. 58–75, 2005.
[6] B. Fan, D. G. Andersen, M. Kaminsky, and M. D. Mitzenmacher, \"Cuckoo Filter: Practically Better Than Bloom,\" in Proc. 10th ACM Int. Conf. Emerging Netw. Exp. Technol., 2014, pp. 75–88.
[7] P. Barham et al., \"Xen and the Art of Virtualization,\" in Proc. ACM Symposium on Operating Systems Principles, 2003.
[8] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, \"Borg, Omega, and Kubernetes,\" Communications of the ACM, vol. 59, no. 5, pp. 50–57, 2016.
[9] T. Erl, Service-Oriented Architecture: Concepts, Technology, and Design, Prentice Hall, 2005.
[10] G. Xu, Y. Pei, and Y. Yang, \"Automated Root Cause Analysis in Cloud Systems Using Machine Learning,\" IEEE Transactions on Cloud Computing, vol. 8, no. 2, pp. 452–465, 2020.
[11] M. Ahmed, A. N. Mahmood, and J. Hu, \"A Survey of Network Anomaly Detection Techniques,\" Journal of Network and Computer Applications, vol. 60, pp. 19–31, 2016.
[12] Redis Ltd., \"Redis Stack Documentation: Probabilistic Data Structures and Streams,\" Redis Official Technical Manual, 2024.
[13] F. Pedregosa et al., \"Scikit-learn: Machine Learning in Python,\" Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.