TRUST-CALIBRATED EXPLAINABLE MULTI-AGENT AI FOR SAFE PREDICTIVE CYBER DEFENCE IN CRITICAL INFRASTRUCTURE AND CLOUD-NATIVE SYSTEMS
Downloads
Objective: This article proposes TCX-MAD, a trust-calibrated explainable multi-agent defence framework for critical infrastructure and cloud-native systems that places a safety governor between collaborative detection and response execution. Method: A design-science methodology specifies the architecture, formal decision policy, threat model, public-dataset evaluation plan, and safety-centred metrics. Results: The framework reports analytical safety properties and an illustrative decision trace rather than fabricated benchmark results because no experimental observations were supplied. Its primary evaluation target is unsafe automated actions prevented without materially increasing valid response latency, supported by detection, disruption, rollback, explanation, and human-approval metrics. Novelty: TCX-MAD integrates separate threat and response-risk estimation, tiered autonomy, dynamic manipulation-aware agent trust, an explanation-sufficiency gate, and rollback-bounded execution. It reframes autonomous cyber defence as constrained and accountable action selection under dual uncertainty.
C. Pascoe, S. Quinn, and K. Scarfone, “The NIST Cybersecurity Framework (CSF) 2.0,” NIST Cybersecurity White Paper NIST CSWP 29, National Institute of Standards and Technology, Gaithersburg, MD, USA, Feb. 2024, doi: 10.6028/NIST.CSWP.29.
A. Kott, P. Théron, M. Drašar, E. Dushku, B. LeBlanc, P. Losiewicz, A. Guarino, L. Mancini, A. Panico, M. Pihelgas, K. Rzadca, and F. De Gaspari, “Autonomous Intelligent Cyber-defense Agent (AICA) Reference Architecture, Release 2.0,” arXiv:1803.10664, 2018, doi: 10.48550/arXiv.1803.10664.
E. Tabassi, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, National Institute of Standards and Technology, Gaithersburg, MD, USA, Jan. 2023, doi: 10.6028/NIST.AI.100-1.
M. Standen, M. Lucas, D. Bowman, T. J. Richer, J. Kim, and D. Marriott, “CybORG: A Gym for the Development of Autonomous Cyber Agents,” in Proc. 30th Int. Joint Conf. Artificial Intelligence (IJCAI), 2021, pp. 4957–4959, doi: 10.48550/arXiv.2108.09118.
TTCP CAGE Working Group, “CAGE Challenge 4: A Multi-Agent Reinforcement Learning Cyber Defence Challenge,” 2024. [Online]. Available: https://github.com/cage-challenge/cage-challenge-4
P. Cichonski, T. Millar, T. Grance, and K. Scarfone, “Computer Security Incident Handling Guide,” NIST Special Publication 800-61 Rev. 2, National Institute of Standards and Technology, Gaithersburg, MD, USA, Aug. 2012, doi: 10.6028/NIST.SP.800-61r2.
Z. Zhang, H. Al Hamadi, E. Damiani, C. Y. Yeun, and F. Taher, “Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research,” IEEE Access, vol. 10, pp. 93104–93139, 2022, doi: 10.1109/ACCESS.2022.3204051.
N. Capuano, G. Fenza, V. Loia, and C. Stanzione, “Explainable Artificial Intelligence in CyberSecurity: A Survey,” IEEE Access, vol. 10, pp. 93575–93600, 2022, doi: 10.1109/ACCESS.2022.3204171.
M. T. Ribeiro, S. Singh, and C. Guestrin, “‘Why Should I Trust You?’: Explaining the Predictions of Any Classifier,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining (KDD), 2016, pp. 1135–1144, doi: 10.1145/2939672.2939778.
S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in Advances in Neural Information Processing Systems 30, 2017, pp. 4765–4774.
P. J. Phillips, C. A. Hahn, P. C. Fontana, A. N. Yates, K. Greene, D. A. Broniatowski, and M. A. Przybocki, “Four Principles of Explainable Artificial Intelligence,” NISTIR 8312, National Institute of Standards and Technology, Gaithersburg, MD, USA, 2021, doi: 10.6028/NIST.IR.8312.
S. Mohseni, N. Zarei, and E. D. Ragan, “A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems,” ACM Trans. Interactive Intelligent Systems, vol. 11, nos. 3–4, Art. no. 24, 2021, doi: 10.1145/3387166.
T. Miller, “Explanation in Artificial Intelligence: Insights from the Social Sciences,” Artificial Intelligence, vol. 267, pp. 1–38, 2019, doi: 10.1016/j.artint.2018.07.007.
R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman, “Metrics for Explainable AI: Challenges and Prospects,” arXiv:1812.04608, 2018, doi: 10.48550/arXiv.1812.04608.
F. Doshi-Velez and B. Kim, “Towards a Rigorous Science of Interpretable Machine Learning,” arXiv:1702.08608, 2017, doi: 10.48550/arXiv.1702.08608.
R. Parasuraman and V. Riley, “Humans and Automation: Use, Misuse, Disuse, Abuse,” Human Factors, vol. 39, no. 2, pp. 230–253, 1997, doi: 10.1518/001872097778543886.
J. D. Lee and K. A. See, “Trust in Automation: Designing for Appropriate Reliance,” Human Factors, vol. 46, no. 1, pp. 50–80, 2004, doi: 10.1518/hfes.46.1.50_30392.
R. Parasuraman, T. B. Sheridan, and C. D. Wickens, “A Model for Types and Levels of Human Interaction with Automation,” IEEE Trans. Systems, Man, and Cybernetics—Part A: Systems and Humans, vol. 30, no. 3, pp. 286–297, 2000, doi: 10.1109/3468.844354.
A. Jøsang, Subjective Logic: A Formalism for Reasoning Under Uncertainty. Cham, Switzerland: Springer, 2016, doi: 10.1007/978-3-319-42337-1.
I. Sharafaldin, A. Habibi Lashkari, and A. A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization,” in Proc. 4th Int. Conf. Information Systems Security and Privacy (ICISSP), 2018, pp. 108–116, doi: 10.5220/0006639801080116.
A. Alsaedi, N. Moustafa, Z. Tari, A. Mahmood, and A. Anwar, “TON_IoT Telemetry Dataset: A New Generation Dataset of IoT and IIoT for Data-Driven Intrusion Detection Systems,” IEEE Access, vol. 8, pp. 165130–165150, 2020, doi: 10.1109/ACCESS.2020.3022862.
M. A. Ferrag, O. Friha, D. Hamouda, L. Maglaras, and H. Janicke, “Edge-IIoTset: A New Comprehensive Realistic Cyber Security Dataset of IoT and IIoT Applications for Centralized and Federated Learning,” IEEE Access, vol. 10, pp. 40281–40306, 2022, doi: 10.1109/ACCESS.2022.3165809.
H.-K. Shin, W. Lee, J.-H. Yun, and B.-G. Min, “Two ICS Security Datasets and Anomaly Detection Contest on the HIL-Based Augmented ICS Testbed,” in Proc. Cyber Security Experimentation and Test Workshop (CSET), 2021, pp. 36–40, doi: 10.1145/3474718.3474719.
Copyright (c) 2024 Sajidul Haque Chowdhruy, Shakila Akter, Md. Golam Mostafa

This work is licensed under a Creative Commons Attribution 4.0 International License.















