Remaining useful life (RUL) prediction on NASA C-MAPSS is mature, yet two requirements of safety-critical deployment remain neglected: calibrated uncertainty quantification and verification, rather than mere reporting, of model explanations. This work adds a trustworthiness layer to standard degradation models as pure post-processing, without retraining. Split conformal prediction and conformalized quantile regression (CQR) are benchmarked on gradient boosting and LSTM models trained on FD001. Fixed-width conformal intervals reach 84.7% coverage at the 90% nominal level and are miscalibrated across the lifetime; CQR restores 89.3% coverage, narrows mean width from 60.3 to 56.1 cycles, and adapts width 2.8-fold. A two-part audit confirms SHAP attributions are faithful: permuting the three top-ranked sensors raises validation RMSE by 8.22 cycles versus 0.65 for the three lowest-ranked, and SHAP agrees with permutation importance (? = 0.868). Centrally, per-instance attribution instability correlates strongly with interval width (? = 0.733, n = 300), so conformal interval width serves as a single distribution-free trust signal flagging unreliable predictions and explanations alike.
Introduction
This text presents a study on Remaining Useful Life (RUL) prediction for aircraft turbofan engines using NASA’s C-MAPSS FD001 dataset. As machine components gradually deteriorate, engine performance decreases through losses in compressor flow capacity and turbine efficiency, reduced stall margins, and reduced operating temperature margins. Accurate RUL prediction is therefore important for safe operation, maintenance planning, cost reduction, and fleet readiness.
Although machine-learning methods such as LSTM/recurrent neural networks, ensemble models, and convolutional neural networks have achieved very good RUL prediction accuracy on C-MAPSS, the authors argue that the main challenges are now trustworthiness rather than prediction accuracy. Two important problems are identified:
Uncertainty: A point estimate of RUL does not indicate how reliable the prediction is. Maintenance decisions require calibrated prediction intervals with known error/coverage rates.
Explainability: Methods such as SHAP are commonly used to identify which sensors influence predictions, but these explanations are often accepted without checking whether they truly represent the model's behavior.
The study therefore investigates the relationship between prediction uncertainty and explanation reliability. It asks whether a model that is uncertain about its RUL prediction also produces unstable or unreliable sensor-level explanations. If such a relationship exists, engineers could use the width of a prediction interval as a warning that both the prediction and its explanation should be treated cautiously.
Research Questions
The research is organized around three main questions:
RQ1 – Calibration: Do conformal prediction methods produce RUL prediction intervals with the expected validity throughout the engine's life?
RQ2 – Faithfulness: Do SHAP explanations accurately represent the model's dependence on different sensor inputs?
RQ3 – Agreement: Can prediction uncertainty, measured using conformal prediction, indicate when model explanations are confused or unstable?
Main Contributions
The paper makes four major contributions:
Develops a leakage-controlled and reproducible FD001 benchmark using engine-level train/test separation and competitive Gradient Boosting and LSTM models.
Evaluates Split Conformal Prediction and Conformalized Quantile Regression (CQR) for both overall and life-stage-specific prediction coverage.
Conducts a quantitative faithfulness audit of SHAP explanations rather than assuming that SHAP rankings are automatically trustworthy.
Introduces a new per-instance attribution instability measure and shows a statistically significant relationship between attribution instability and conformal prediction interval width.
C-MAPSS FD001 Dataset
The FD001 subset contains:
100 training engines with complete run-to-failure histories.
100 test engines with truncated histories.
One flight/operating condition.
One primary fault mode: high-pressure compressor (HPC) degradation.
21 sensor channels and 3 operational settings.
20,631 training records and 13,096 test records.
Training engine lifetimes ranging from 128 to 362 cycles.
Failure occurs when the simulated engine health index reaches zero because of increasing efficiency and flow losses in the HPC.
Feature Selection
The authors analyze sensor variability and remove sensors that provide little useful information. Seven channels—sensors 1, 5, 6, 10, 16, 18, and 19—are constant or nearly constant and are therefore discarded. The three operational settings are also removed because FD001 operates under essentially one flight condition.
Consequently, the machine-learning models use 14 informative sensor features as their input.
Conformal Prediction and SHAP
Conformal prediction is used to convert point RUL predictions into prediction intervals with a finite-sample coverage guarantee. Split conformal prediction is simple and does not require retraining, while CQR provides intervals whose widths can adapt to different regions of the input space.
SHAP (Shapley Additive Explanations) is used to determine how individual sensors contribute to each RUL prediction. However, the study emphasizes that SHAP explanations should be tested for faithfulness rather than automatically trusted. The authors examine attribution stability under input perturbations and compare SHAP results with permutation-based importance.
Overall Significance
Conclusion
This paper cast RUL prediction on C-MAPSS as a problem of prediction trustworthiness rather than accuracy alone, and provided the previously unavailable link between its two components. Conformalized quantile regression was demonstrated to provide near-nominal, locally adaptive prediction intervals on a leakage-controlled FD001 benchmark, relative to fixed-width conformal prediction, which was substantially miscalibrated across life stages; a two-part audit showed SHAP values on this dataset to be faithful, with a 12.6-fold difference in perturbation effect between highest- and lowest-ranked sensors and a rank correlation of ? = 0.868 between methods; finally, this calibrated interval width was shown to correlate strongly with the variation in those attributions (? = 0.733, p = 7.8 × 10?52) when tested against input jittering. Not only do calibrated uncertainty estimates and feature attributions agree, but their agreement endows the width of the conformal prediction interval with utility as a single, distribution-free, retraining-free trust metric for safety-critical mechanical system maintenance.
References
[1] Saxena A, Goebel K, Simon D, Eklund N. Damage propagation modeling for aircraft engine run-to-failure simulation. In: Proceedings of the International Conference on Prognostics and Health Management; Denver, CO. 2008. p. 1-9. https://doi.org/10.1109/PHM.2008.4711414
[2] Frederick D, DeCastro J, Litt J. User\'s guide for the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS). NASA Technical Report TM-2007-215026. Cleveland, OH: NASA Glenn Research Center; 2007.
[3] Heimes FO. Recurrent neural networks for remaining useful life estimation. In: Proceedings of the International Conference on Prognostics and Health Management; Denver, CO. 2008. p. 1-6. https://doi.org/10.1109/PHM.2008.4711422
[4] Zheng S, Ristovski K, Farahat A, Gupta C. Long short-term memory network for remaining useful life estimation. In: Proceedings of the IEEE International Conference on Prognostics and Health Management (ICPHM); 2017. p. 88-95.
https://doi.org/10.1109/ICPHM.2017.7998311
[5] Hu C, Youn BD, Wang P, Yoon JT. Ensemble of data-driven prognostic algorithms for robust prediction of remaining useful life. Reliability Engineering & System Safety 2012; 103: 120-135. https://doi.org/10.1016/j.ress.2012.03.008
[6] Li X, Ding Q, Sun JQ. Remaining useful life estimation in prognostics using deep convolution neural networks. Reliability Engineering & System Safety 2018; 172: 1-11. https://doi.org/10.1016/j.ress.2017.11.021
[7] Lundberg SM, Lee SI. A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems 30; 2017. p. 4765-4774.
[8] Lundberg SM, Erion G, Chen H, DeGrave A, Prutkin JM, Nair B, Katz R, Himmelfarb J, Bansal N, Lee SI. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence 2020; 2(1): 56-67. https://doi.org/10.1038/s42256-019-0138-9
[9] Alvarez-Melis D, Jaakkola TS. On the robustness of interpretability methods. In: Proceedings of the ICML Workshop on Human Interpretability in Machine Learning; 2018.
[10] Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. In: Advances in Neural Information Processing Systems 31; 2018. p. 9505-9515.
[11] Slack D, Hilgard S, Jia E, Singh S, Lakkaraju H. Fooling LIME and SHAP: adversarial attacks on post hoc explanation methods. In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society; 2020. p. 180-186. https://doi.org/10.1145/3375627.3375830
[12] Baptista ML, Goebel K, Henriques EMP. Relation between prognostics predictor evaluation metrics and local interpretability SHAP values. Artificial Intelligence 2022; 306: 103667. https://doi.org/10.1016/j.artint.2022.103667
[13] Vovk V, Gammerman A, Shafer G. Algorithmic Learning in a Random World. New York, NY: Springer; 2005.
[14] Lei J, G\'Sell M, Rinaldo A, Tibshirani RJ, Wasserman L. Distribution-free predictive inference for regression. Journal of the American Statistical Association 2018; 113(523): 1094-1111. https://doi.org/10.1080/01621459.2017.1307116
[15] Angelopoulos AN, Bates S. Conformal prediction: a gentle introduction. Foundations and Trends in Machine Learning 2023; 16(4): 494-591. https://doi.org/10.1561/2200000101
[16] Romano Y, Patterson E, Candès E. Conformalized quantile regression. In: Advances in Neural Information Processing Systems 32; 2019. p. 3543-3553.
[17] Javanmardi A, Hüllermeier E. Conformal prediction intervals for remaining useful lifetime estimation. International Journal of Prognostics and Health Management 2023; 14(2). https://doi.org/10.36001/ijphm.2023.v14i2.3417
[18] Cordier T, Blot V, Lacombe L, Morzadec T, Capitaine A, Brunel N. Flexible and systematic uncertainty estimation with conformal prediction via the MAPIE library. In: Proceedings of the Twelfth Symposium on Conformal and Probabilistic Prediction with Applications, Proceedings of Machine Learning Research 204; 2023. p. 549-581.
[19] Friedman JH. Greedy function approximation: a gradient boosting machine. Annals of Statistics 2001; 29(5): 1189-1232.
https://doi.org/10.1214/aos/1013203451
[20] Hochreiter S, Schmidhuber J. Long short-term memory. Neural Computation 1997; 9(8): 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735
[21] Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E. Scikit-learn: machine learning in Python. Journal of Machine Learning Research 2011; 12: 2825-2830.
[22] Abadi M, Barham P, Chen J, Chen Z, Davis A, Dean J, Devin M, Ghemawat S, Irving G, Isard M, Kudlur M, Levenberg J, Monga R, Moore S, Murray DG, Steiner B, Tucker P, Vasudevan V, Warden P, Wicke M, Yu Y, Zheng X. TensorFlow: a system for large-scale machine learning. In: Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI); Savannah, GA. 2016. p. 265-283.
[23] Breiman L. Random forests. Machine Learning 2001; 45(1): 5-32. https://doi.org/10.1023/A:1010933404324