Data-Enabled Fault Detection and Reliability Assessment in Modern Engineering Infrastructure

Authors

  • Keepak Heokhale Department of Computer Science, Colorado State University, Fort Collins, CO, USA.
  • Massimo D. Larsen Department of Computer Science, University of Central Florida, Orlando, FL, USA.
  • Zhengzihan Chen Department of Computer Science and Engineering, University at Buffalo, Buffalo, NY, USA.
  • Arthur Bearnett Department of Computer Science and Engineering, University of Nevada, Reno, Reno, NV, USA.

Keywords:

Data-driven fault detection, reliability engineering, infrastructure systems, system resilience, condition monitoring, socio-technical governance

Abstract

The accelerating digitalization of critical engineering infrastructure has created an unprecedented opportunity to embed data-enabled fault detection and reliability assessment into the operational fabric of large-scale systems. Moving beyond traditional periodic inspection and rule-based alarms, modern infrastructure increasingly relies on streaming sensor data, machine learning models, and networked decision-support platforms to anticipate degradation, localize anomalies, and optimize maintenance in near real time. This paper provides a system-level examination of the architectural, governance, and policy dimensions that determine whether such data-driven capabilities can be sustained, scaled, and trusted. We argue that the migration from physics-based models to hybrid and data-intensive paradigms fundamentally reshapes the structural trade-offs between specificity and generalizability, between centralized cloud analytics and edge autonomy, and between model accuracy and explainability. The discussion integrates perspectives from reliability engineering, cyber-physical systems, machine learning lifecycles, and sociotechnical governance to identify the points of friction that emerge when algorithmic fault detection meets regulatory frameworks originally designed for electromechanical assets. Special emphasis is placed on deployment robustness under nonstationary operating conditions, fairness in predictive maintenance decisions that affect multiple user communities, environmental sustainability of pervasive sensing, and the institutional adaptations required for shared data infrastructures. By synthesizing cross-domain case illustrations and prior frameworks, the paper proposes a forward-looking research agenda that treats reliability assessment not as a stand-alone technical computation but as an infrastructure-wide governance challenge requiring new standards, lifecycle accountability mechanisms, and inclusive policy designs. The analysis concludes that sustainable data-enabled infrastructure reliability depends as much on organizational alignment and adaptive regulation as on sensor density or algorithmic sophistication.

References

1. Zio, E. (2016). Challenges in the vulnerability and risk analysis of critical infrastructures. Reliability Engineering & System Safety, 152, 137–150. https://doi.org/10.1016/j.ress.2016.02.009

2. Jardine, A. K. S., Lin, D., & Banjevic, D. (2006). A review on machinery diagnostics and prognostics implementing condition-based maintenance. Mechanical Systems and Signal Processing, 20(7), 1483–1510. https://doi.org/10.1016/j.ymssp.2005.09.012

3. Rausand, M., & Høyland, A. (2004). System reliability theory: Models, statistical methods, and applications (2nd ed.). John Wiley & Sons.

4. Tao, F., Zhang, H., Liu, A., & Nee, A. Y. C. (2019). Digital twin in industry: State-of-the-art. IEEE Transactions on Industrial Informatics, 15(4), 2405–2415. https://doi.org/10.1109/TII.2018.2873186

5. Lee, J., Bagheri, B., & Kao, H.-A. (2015). A cyber-physical systems architecture for industry 4.0-based manufacturing systems. Manufacturing Letters, 3, 18–23. https://doi.org/10.1016/j.mfglet.2014.12.001

6. Montgomery, D. C. (2007). Introduction to statistical quality control (5th ed.). John Wiley & Sons.

7. Isermann, R. (2005). Model-based fault-detection and diagnosis – status and applications. Annual Reviews in Control, 29(1), 71–85. https://doi.org/10.1016/j.arcontrol.2004.12.002

8. Lei, Y., Yang, B., Jiang, X., Jia, F., Li, N., & Nandi, A. K. (2020). Applications of machine learning to machine fault diagnosis: A review and roadmap. Mechanical Systems and Signal Processing, 138, 106587. https://doi.org/10.1016/j.ymssp.2019.106587

9. Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58. https://doi.org/10.1145/1541880.1541882

10. Venkatasubramanian, V., Rengaswamy, R., Yin, K., & Kavuri, S. N. (2003). A review of process fault detection and diagnosis: Part I: Quantitative model-based methods. Computers & Chemical Engineering, 27(3), 293–311. https://doi.org/10.1016/S0098-1354(02)00160-6

11. Stamatelatos, M., Dezfuli, H., Apostolakis, G., Everline, C., Guarro, S., Mathias, D., Mosleh, A., & Youngblood, R. (2011). Probabilistic risk assessment procedures guide for NASA managers and practitioners (NASA/SP-2011-3421). NASA.

12. Weber, P., Medina-Oliva, G., Simon, C., & Iung, B. (2012). Overview on Bayesian networks applications for dependability, risk analysis and maintenance areas. Engineering Applications of Artificial Intelligence, 25(4), 671–682. https://doi.org/10.1016/j.engappai.2010.06.002

13. Labeau, P. E., Smidts, C., & Swaminathan, S. (2000). Dynamic reliability: Towards an integrated platform for probabilistic risk assessment. Reliability Engineering & System Safety, 68(3), 219–254. https://doi.org/10.1016/S0951-8320(00)00004-9

14. Francis, R., & Bekera, B. (2014). A metric and frameworks for resilience analysis of engineered and infrastructure systems. Reliability Engineering & System Safety, 121, 90–103. https://doi.org/10.1016/j.ress.2013.07.004

15. Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2016). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 3(5), 637–646. https://doi.org/10.1109/JIOT.2016.2579198

16. Khatri, V., & Brown, C. V. (2010). Designing data governance. Communications of the ACM, 53(1), 148–152. https://doi.org/10.1145/1629175.1629210

17. ISO 13374-1:2003. Condition monitoring and diagnostics of machines — Data processing, communication and presentation — Part 1: General guidelines. International Organization for Standardization.

18. De Bruijn, H., & Herder, P. M. (2009). System and actor perspectives on sociotechnical systems. IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, 39(5), 981–992. https://doi.org/10.1109/TSMCA.2009.2025452

19. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607

20. Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR). arXiv:1412.6572.

21. Hauschild, M. Z., Rosenbaum, R. K., & Olsen, S. I. (Eds.). (2018). Life cycle assessment: Theory and practice. Springer. https://doi.org/10.1007/978-3-319-56475-3

Downloads

Published

2026-03-15