Modern cloud applications are built using microservices that support various sectors such as healthcare, finance, e-commerce and communication. Microservices’ distributed and scalable nature provides high performance and flexibility but also increases system complexity. When failures occur, they propagate rapidly through interconnected microservices, leading to performance degradation, service downtime and financial losses. Traditional fault detection approaches mainly depend on threshold-based rules and reactive fault tracing, which are time-consuming, difficult to scale and often fail to proactively detect faults in dynamic service environments. Existing research has proposed a Graph Neural Network (GNN)-based framework that models microservice interactions as graphs for fault prediction and critical component identification. The framework, trained on simulated data generated using Markov Decision Processes (MDPs), has demonstrated strong performance in controlled environments. However, its use of synthetic data limits its effectiveness on real-world cloud systems, which exhibit unpredictable behaviour and dynamic workloads. The proposed work addresses this limitation using the Alibaba Cluster Trace – Microservices v2021 dataset collected from a real production environment. This work incorporates an Edge-Attention Graph Convolutional Network (EA-GCN), a hybrid loss function and Critical Node Removal Analysis (CNRA) to improve fault prediction and identify critical microservices. Experimental results demonstrate that the proposed approach achieves an F1-score of 0.7265 and an AUC-ROC of 0.9583, improving F1-score by 47.9% over plain GCN. CNRA identifies the most critical microservices that contribute to system-wide fault risk, prioritizing fault mitigation. The results show that the framework is effective for fault prediction and critical microservice identification in real-world cloud environments.
Introduction
The text presents an Edge-Attention Graph Convolutional Network (EA-GCN) for predicting faults in cloud microservice systems. The main goal is to move from traditional reactive monitoring to proactive fault prediction by considering both microservice performance metrics and the dependencies between services.
Key points
Problem: Modern microservice applications depend on many interconnected services. A failure or slowdown in one service can propagate through its dependencies and cause cascading failures across the application.
Limitations of traditional monitoring: Tools such as Prometheus and Nagios generally rely on fixed thresholds and monitor services individually. They may detect problems only after they occur and do not adequately model relationships between dependent services.
Use of GNNs: The proposed approach represents microservices as graph nodes and service-to-service calls as edges. This allows the model to learn how faults may propagate through the service network.
Real-world dataset: The study uses the Alibaba Cluster Trace — Microservices v2021, containing millions of real production records from thousands of microservices.
Graph construction: The study selects the top 30 microservices based on connectivity and creates a graph with 60 high-frequency dependency edges. Edge weights represent actual call frequency between services.
Node features: Each microservice is represented using seven features:
CPU usage
Memory usage
Error rate
Response latency
Request rate
Total call count
Degree centrality
EA-GCN architecture: The proposed model combines:
Edge-weighted GCN to incorporate real call frequencies.
Graph Attention Network (GAT) to learn which neighbouring services are most important.
Skip-connected GCN to preserve original service metrics and improve learning.
An MLP output layer that produces a fault probability for each microservice.
Hybrid loss function: The model combines standard Binary Cross-Entropy (BCE) with a graph-smoothness term. This encourages connected services to have reasonably consistent fault-risk predictions when appropriate.
Fault scenarios: Because the Alibaba dataset does not contain public ground-truth incident labels, the researchers generate 300 fault-injection scenarios, including:
Normal operation
CPU spikes
Memory leaks
Latency spikes
Error surges
Cascading failures
Critical Node Removal Analysis (CNRA): The study removes each microservice from the graph one at a time and measures how the overall predicted system risk changes. This identifies services that may contribute significantly to system-wide fault risk.
Research contribution: Unlike earlier work that relied mainly on synthetic data or focused on detecting failures after they occurred, this research attempts to provide proactive, service-level fault forecasting using real production traces.
Conclusion
This work validates GNN-based fault forecasting on real-world production cloud microservice traces, addressing a key limitation of Unyi et al. [1]. The proposed EA-GCN improves the F1-score by 47.9% over a standard GCN, highlighting the importance of incorporating call-frequency-based edge attention for more accurate fault prediction. Additionally, a hybrid BCE–smoothness loss and structural analysis using CNRA improve both predictive performance and interpretability.
Future work will address three key areas: (1) scaling EA-GCN to the full 6,466-service Alibaba graph via mini-batch training with GraphSAGE-style neighbourhood sampling; (2) incorporating temporal modelling through spatiotemporal GNNs to capture time-evolving fault patterns such as gradual memory leaks; and (3) integrating real incident labels from Alibaba or equivalent production systems to enable fully supervised training and rigorous evaluation on actual failure events.
References
[1] D. Unyi, E. Rigó, B. Gyires-Tóth, and R. Lovas, \"Explainable GNN-Based Approach to Fault Forecasting in Cloud Service Debugging,\" IEEE Trans. Netw. Serv. Manag., 2025, doi: 10.1109/TNSM.2025.3602223.
[2] T. N. Kipf and M. Welling, \"Semi-Supervised Classification with Graph Convolutional Networks,\" in Proc. ICLR, 2017, arXiv:1609.02907.
[3] P. Veli?kovi?, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, \"Graph Attention Networks,\" in Proc. ICLR, 2018, arXiv:1710.10903.
[4] L. Breiman, \"Random Forests,\" Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.
[5] S. Luo et al., \"Characterizing Microservice Dependency and Performance: Alibaba Trace Analysis,\" in Proc. ACM SoCC, 2021, doi: 10.1145/3472883.3487003.
[6] Alibaba Group, \"Cluster Trace — Microservices v2021,\" Available: http://aliopentrace.oss-cn-beijing.aliyuncs.com/v2021MicroservicesTraces, Accessed: Mar. 2026.
[7] A. Paszke et al., \"PyTorch: An Imperative Style, High-Performance Deep Learning Library,\" in Proc. NeurIPS, 2019.
[8] F. Pedregosa et al., \"Scikit-learn: Machine Learning in Python,\" J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.
[9] A. Maroudis et al., \"Leveraging Graph Neural Networks for SLA Violation Prediction in Cloud Computing,\" IEEE Trans. Netw. Serv. Manag., vol. 21, no. 1, pp. 605–620, 2024, doi: 10.1109/TNSM.2023.3292392.
[10] L. Wu, J. Tordsson, E. Elmroth, and O. Kao, \"MicroRCA: Root Cause Localization of Performance Issues in Microservices,\" in NOMS, 2020, pp. 1–9.
[11] Y. Gan et al., \"Sleuth: A Trace-Based Root Cause Analysis System for Large-Scale Microservices with GNNs,\" in Proc. ASPLOS, 2023, pp. 324–337, doi: 10.1145/3623278.3624758.
[12] X. Zhou et al., \"Latent Error Prediction and Fault Localization for Microservice Applications by Learning from System Trace Logs,\" in Proc. ESEC/FSE, 2019, pp. 26–30.
[13] R. Rouf et al., \"InstantOps: A Joint Approach to System Failure Prediction and Root Cause Identification in Microservices,\" in Proc. ACM/SPEC Int. Conf. Performance Engineering (ICPE), 2024, pp. 119–129, doi: 10.1145/3629526.3645047.