Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Adepegba Akindayomi Akintade
DOI Link: https://doi.org/10.22214/ijraset.2026.84531
Certificate: View Certificate
Disaster-recovery plans commonly prioritize infrastructure components, virtual machines, or recovery tasks, although a recovered component does not by itself restore a usable business service. This study formulates hybrid-cloud recovery as dependency-complete restoration of weighted business capabilities. A service is counted as available only after its required network, identity, security, data, and application dependencies are jointly recovered and validated. The problem is modelled as a resource-constrained, multi-mode recovery schedule with recovery-point limits and security gating, and it is solved with a dependency-aware multi-objective optimizer, MVS-NSGA-II, that minimizes weighted capability downtime, recovery cost, residual risk, and data-loss exposure. Against an exact reference model on small cases and against conventional priority policies on matched medium and large scenarios, the optimizer reduces weighted capability downtime substantially under moderate resource scarcity while respecting recovery-point and security constraints. The benefit is largest where dependencies are dense and resources are constrained, and it narrows when capacity is abundant or a single objective dominates. Sensitivity analysis over duration uncertainty and resource scarcity shows stable and interpretable behaviour, and planning runtime remains practical for operational use. The results indicate that recovery should be measured and optimized at the capability boundary rather than at the component boundary, and that dependency-aware, multi-objective scheduling with an explicit decision policy provides a defensible basis for recovery runbooks and orchestration platforms in multi-site hybrid clouds.
This second text presents a business-capability-centred optimization framework for cloud disaster recovery (DR). Its central argument is that restoring infrastructure components is not equivalent to restoring a usable business service.
Modern services depend on long chains of infrastructure and external dependencies—identity, networking, storage, databases, messaging, applications, security controls, and third-party services. Consequently, an individual component can be “recovered” while the actual business capability remains unavailable.
The paper therefore defines recovery around dependency-complete business capabilities, rather than individual servers, VMs, applications, or workflows.
The key distinction is:
Technical recovery is not the same as service recovery.
A database being operational, for example, does not mean the customer-facing service is restored if identity, networking, application dependencies, or integrity validation are still unavailable.
The study represents recovery as a directed acyclic graph (DAG) of recovery tasks.
Each task can have multiple recovery modes, such as:
Each mode has:
The scheduler must therefore simultaneously decide which recovery mode to use and in what order to execute the tasks, while respecting dependencies and limited resources.
A business capability k consists of a set of required technical tasks Rk?.
Its activation time is:
Tk?=maxi∈Rk??Ci?
In other words, the capability is not restored until all of its required dependencies and validation conditions are complete.
This leads to two important recovery metrics:
Thus:
TMVS?=maxk∈K∗?Tk?
where K∗ is the set of essential capabilities.
This is arguably the paper's strongest conceptual contribution because it prevents an optimizer from declaring “recovery” merely because a collection of infrastructure tasks has completed.
The proposed search simultaneously considers:
The paper proposes a multi-objective vector:
(WCD, MVS, Cost, Risk + 2×RPO exposure, Cmax)
The authors explicitly acknowledge that the coefficient on RPO exposure is a synthetic experimental scaling choice, rather than a universal business valuation.
The review positions the approach between several established research areas:
| Research area | Typical unit | Main limitation addressed here |
|---|---|---|
| Business continuity standards | Mission/business functions | No computational scheduling method |
| Proactive cloud recovery | Faults, hosts, predefined workflows | Usually not general capability graphs |
| VM disaster-recovery optimization | VMs/resources | Does not directly measure business-service restoration |
| Service-chain orchestration | Resources/service functions | Focuses on QoS, latency, cost, stability |
| Dependency management | Production service graphs | Supports auditing/mitigation rather than recovery optimization |
| RCPSP/DAG scheduling | Activities/workflows | Usually emphasizes makespan/resource objectives |
| This study | Dependency-complete business capabilities | Connects technical recovery scheduling directly to business-service availability |
The study makes five principal contributions:
The paper identifies a gap between two worlds:
Business continuity asks:
Which services must be restored, and how quickly?
Scheduling research asks:
How can constrained technical tasks be sequenced efficiently?
Cloud recovery research asks:
How can infrastructure or workflows be restored automatically?
The proposed framework attempts to connect them:
business impact → capability dependencies → recovery tasks → resource-constrained scheduling → capability restoration → MVS.
That gives the work a clearer research identity than simply proposing another disaster-recovery scheduling heuristic.
The paper's strongest claim should probably remain carefully scoped:
MVS is a modelling construct, not evidence that the optimized schedule will actually restore a real organization's service.
Its validity depends on the quality of:
Disaster recovery is not complete when a collection of components has restarted. It is complete at a chosen service level when the dependencies and assurance conditions required for that service are functioning. This study formalized that principle through weighted capability downtime and dependency-complete minimum viable service time, embedded it in a multi-mode resource-constrained recovery model, and evaluated a capability-aware NSGA-II procedure across transparent synthetic hybrid-cloud scenarios. MVS-NSGA-II matched the exact scalarized solution in the small deterministic case and, across 90 matched primary instances, reduced weighted capability downtime, MVS time, and makespan relative to the best conventional policy. The gains were meaningful but not universal, and they involved modest cost and larger synthetic risk increases. These trade-offs are a feature of the decision problem, not a defect to conceal. The practical implication is direct: recovery plans should define capabilities, dependencies, alternative modes, validation tasks, and resource constraints in a machine-readable form. Optimization can then support incident commanders with feasible trade-offs, while policy determines which plan is authorized. Future validation should use real recovery exercises and production dependency data, but the released model and scenarios provide a reproducible foundation for that work.
[1] Bartock, M., Cichonski, J., Souppaya, M., Smith, M. C., Witte, G., & Scarfone, K. (2016). Guide for cybersecurity event recovery (NIST SP 800-184). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-184 [2] Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.). (2016). Site reliability engineering: How Google runs production systems. O\'Reilly Media. [3] Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. (2002). A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2), 182-197. https://doi.org/10.1109/4235.996017 [4] Hartmann, S., & Briskorn, D. (2022). An updated survey of variants and extensions of the resource-constrained project scheduling problem. European Journal of Operational Research, 297(1), 1-14. https://doi.org/10.1016/j.ejor.2021.05.004 [5] Hartmann, S., & Kolisch, R. (2000). Experimental evaluation of state-of-the-art heuristics for the resource-constrained project scheduling problem. European Journal of Operational Research, 127(2), 394-407. https://doi.org/10.1016/S0377-2217(99)00485-3 [6] International Organization for Standardization. (2019). ISO 22301:2019 Security and resilience - Business continuity management systems - Requirements. ISO. [7] International Organization for Standardization. (2025). ISO/IEC 27031:2025 Cybersecurity - Information and communication technology readiness for business continuity. ISO. [8] Kolisch, R., & Hartmann, S. (2006). Experimental investigation of heuristics for resource-constrained project scheduling: An update. European Journal of Operational Research, 174(1), 23-37. https://doi.org/10.1016/j.ejor.2005.01.065 [9] Levy, S., Yao, R., Wu, Y., Dang, Y., Huang, P., Mu, Z., Zhao, P., Ramani, T., Govindaraju, N., Li, X., Lin, Q., Shafriri, G. L., & Chintalapati, M. (2020). Predictive and adaptive failure mitigation to avert production cloud VM interruptions. In 14th USENIX Symposium on Operating Systems Design and Implementation (pp. 1155-1170). USENIX Association. [10] Li, J., Zhong, Y., Zhu, S., & Zhang, X. (2026). Grey-number PPO for uncertainty-aware multi-objective DAG scheduling in multi-cloud environments. Journal of King Saud University Computer and Information Sciences. https://doi.org/10.1007/s44443-026-00894-1 [11] Lou, C., Chen, C., Huang, P., Dang, Y., Qin, S., Yang, X., Li, X., Lin, Q., & Chintalapati, M. (2022). RESIN: A holistic service for dealing with memory leaks in production cloud infrastructure. In 16th USENIX Symposium on Operating Systems Design and Implementation (pp. 109-125). USENIX Association. [12] Meng, K., Li, M., Ding, J., & Zhou, H. (2026). Cloud disaster recovery model based on failure prediction. Journal of Cloud Computing, 15, Article 33. https://doi.org/10.1186/s13677-026-00846-0 [13] Quinn, S., Ivy, N., Chua, J., Barrett, M., Feldman, L., Topper, D., Witte, G., & Gardner, R. (2025). Using business impact analysis to inform risk prioritization and response (NIST IR 8286D-upd1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.IR.8286D-upd1 [14] Roberts, M. K., Shanmugam, S. K., Alhammad, S. M., & Khafaga, D. S. (2026). Risk-aware resilient cloud orchestration under correlated faults using adaptive dual-regime search with self-healing control. Engineering Science and Technology, an International Journal, 77, Article 102365. https://doi.org/10.1016/j.jestch.2026.102365 [15] Stamenkov, G. (2022). Layered business continuity and disaster recovery model. Continuity & Resilience Review, 4(3), 267-279. https://doi.org/10.1108/CRR-05-2022-0008 [16] Swanson, M., Bowen, P., Phillips, A. W., Gallup, D., & Lynes, D. (2010). Contingency planning guide for federal information systems (NIST SP 800-34 Rev. 1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-34r1 [17] Tomarchio, O., Calcaterra, D., & Di Modica, G. (2020). Cloud resource orchestration in the multi-cloud landscape: A systematic review of existing frameworks. Journal of Cloud Computing, 9, Article 49. https://doi.org/10.1186/s13677-020-00194-7 [18] Topcuoglu, H., Hariri, S., & Wu, M.-Y. (2002). Performance-effective and low-complexity task scheduling for heterogeneous computing. IEEE Transactions on Parallel and Distributed Systems, 13(3), 260-274. https://doi.org/10.1109/71.993206 [19] Wang, L., He, J., Peng, J., Zhou, L., & Zhang, Z. (2025). NSGA-II based multi-objective disaster recovery scheduling for virtual cloud platforms. Informatica, 49(36). https://doi.org/10.31449/inf.v49i36.11126 [20] Wang, W., Ren, Q., Yao, G., & Zhu, S. (2026). Reliable cloud workflow scheduling with uncertain task execution time. Journal of King Saud University Computer and Information Sciences, 38, Article 216. https://doi.org/10.1007/s44443-026-00585-x [21] Xie, Y., Wu, K., Jiang, Y., Zhang, X., & Cui, W. (2026). Hierarchical service chain orchestration for multi-cloud environments enabled by deep reinforcement learning. Journal of Cloud Computing, 15, Article 69. https://doi.org/10.1186/s13677-026-00874-w [22] Yang, T., Li, B., Shen, J., Su, Y., Yang, Y., & Lyu, M. R. (2022). Managing service dependency for cloud reliability: The industrial practice. In 33rd IEEE International Symposium on Software Reliability Engineering Workshops (pp. 67-68). IEEE. [23] Zhai, E., Chen, A., Piskac, R., Balakrishnan, M., Tian, B., Song, B., & Zhang, H. (2020). Check before you change: Preventing correlated failures in service updates. In 17th USENIX Symposium on Networked Systems Design and Implementation. USENIX Association.
Copyright © 2026 Adepegba Akindayomi Akintade. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84531
Publish Date : 2026-08-04
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
