1.信息工程大学,河南郑州 450002
2.河南省网络空间内生安全重点实验室,河南郑州 450000
3.网络空间安全教育部重点实验室,河南郑州 450000
裴金川 男,1998年6月出生于河北省唐山市。现为信息工程大学博士研究生。主要研究方向为网络主动防御与新型网络体系结构。 E-mail: jinchuan_pei@163.com
胡宇翔 男,1982年出生于河南省周口市。现为信息工程大学教授,博士生导师。主要研究方向为网络空间安全与新型网络体系结构。 E-mail: huyuxiangchn@163.com
于洪涛 男,1970年出生于辽宁省丹东市。现为信息工程大学研究员,博士生导师。主要研究方向为大数据与人工智能。 E-mail: hgnn2kan@163.com
田乐 男,1985年出生于陕西省咸阳市。现为信息工程大学副研究员,硕士生导师。主要研究方向为新型网络体系结构。 E-mail: le.tian2019@outlook.com
李梦龙 男,1992年出生于河南省新乡市。现为信息工程大学助理工程师。主要研究方向为云网融合与强化学习。 E-mail: growinglml@163.com
收稿:2026-03-30,
录用:2026-04-14,
网络首发:2026-05-20,
纸质出版:2026-04-25
移动端阅览
裴金川, 胡宇翔, 于洪涛, 等. 基于时滞FlipIt博弈与多智能体强化学习的欺骗防御时机选取方法[J]. 电子学报, 2026, 54(04): 1823-1834.
PEI Jinchuan, HU Yuxiang, YU Hongtao, et al. A Deception Defense Timing Selection Method Based on FlipIt Game with Time Delay and Multi-Agent Reinforcement Learning[J]. Acta Electronica Sinica, 2026, 54(04): 1823-1834.
裴金川, 胡宇翔, 于洪涛, 等. 基于时滞FlipIt博弈与多智能体强化学习的欺骗防御时机选取方法[J]. 电子学报, 2026, 54(04): 1823-1834. DOI:10.12263/DZXB.20260221
PEI Jinchuan, HU Yuxiang, YU Hongtao, et al. A Deception Defense Timing Selection Method Based on FlipIt Game with Time Delay and Multi-Agent Reinforcement Learning[J]. Acta Electronica Sinica, 2026, 54(04): 1823-1834. DOI:10.12263/DZXB.20260221
云边协同网络在6G通信、工业互联网等关键领域中广泛应用,其开放分布式特性面临持续演进的网络攻击威胁。网络欺骗防御通过蜜罐、诱饵等资源干扰攻击者认知,但其有效性高度依赖于欺骗资源的轮换时机。然而,在实际攻防过程中,欺骗资产部署、攻击渗透扩散均存在不可忽略的状态转换时滞,且攻防双方存在持续的策略交互,现有方法或假设状态转移瞬时完成,或采用固定周期等静态机制,忽视了真实攻防过程中存在的状态转换时滞及攻防双方的动态交互,无法准确刻画攻防状态演化的时序特征,导致策略在真实高动态环境中适应性不足。针对上述问题,本文提出一种基于时滞FlipIt博弈与多智能体强化学习的欺骗防御时机选取方法。首先,分析网络安全状态演化特性,将正常、保护、感染、受损四种节点状态间的六条转移路径分别引入对应的时滞参数,构建融合状态转换时滞的时滞微分方程组,形成更贴近实际物理过程的网络状态演化模型,弥补了传统SIS、SIR等动力学模型忽略时滞因素的缺陷。其次,分别建立攻击者与防御者模型,将攻防对抗形式化为具有时序特征的博弈过程,在此基础上提出时滞FlipIt博弈模型,显式刻画策略下发到状态生效之间的延迟机制,并给出攻防双方在时滞约束下的效用函数。然后采用多智能体近端策略优化(Multi-Agent Proximal Policy Optimization,MAPPO)算法,构建中心化训练、分布式执行的强化学习框架,设计状态空间、动作空间与奖励函数,通过攻防智能体的持续交互与策略演化,求解最优欺骗防御时机轮换策略。实验在Mininet仿真环境与小规模真实测试床上进行,结果表明所提方法能够有效优化防御时机选择,保护状态节点比例稳定在83.2%,攻击者效用值较MFD-PPO方法降低44.85%,攻击拦截成功率较FP方法提升38.4%,同时,在同等保护状态节点比例下IP跳变频率显著低于对比方法,验证了所提方法在保障防御效能与降低资源开销方面的综合优势。
The cloud-edge collaborative network is widely applied in key fields such as 6G communication and industrial internet. Its open and distributed characteristics are facing continuous evolving network attack threats. Network deception defense uses resources such as honeypots and decoys to interfere with the attacker’s cognition
but its effectiveness highly depends on the timing of the rotation of deception resources. However
in actual offensive and defensive processes
there are unignorable state transition delays in the deployment of deception assets and the spread of attack penetration
and both the attacker and defender have continuous strategic interactions. Existing methods or assumptions that the state transition is instantaneous or adopt static mechanisms such as fixed cycles ignore the state transition delays and the dynamic interaction between the attacker and defender in the actual offensive and defensive process
making it impossible to accurately depict the temporal characteristics of the evolution of the offensive and defensive states
resulting in insufficient adaptability of the strategies in the real high-dynamic environment. To address these issues
this paper proposes a deception defense timing selection method based on time-delay FlipIt game and multi-agent reinforcement learning. Firstly
it analyzes the evolution characteristics of network security states and introduces six transfer paths between four node states (normal
protected
infected
and damaged) into corresponding time-delay parameters to construct a time-delay differential equation system that integrates state transition time delays
forming a network state evolution model that is closer to the actual physical process and compensates for the shortcomings of traditional SIS
SIR
etc. dynamic models that ignore time-delay factors. Secondly
it establishes models for the attacker and defender
formalizes the offensive and defensive confrontation as a game process with temporal characteristics
and proposes a time-delay FlipIt game model to explicitly depict the delay mechanism between the strategy being issued and taking effect in the state
and gives the utility functions of both the attacker and defender under the time-delay constraints. Then
using the multi-agent proximal policy optimization (MAPPO) algorithm
a centralized training and distributed execution reinforcement learning framework is constructed
designing the state space
action space
and reward function
and solving the optimal deception defense timing rotation strategy through the continuous interaction and strategy evolution of the offensive and defensive agents. Experiments are conducted in the Mininet simulation environment and a small-scale real testbed. The results show that the proposed method can effectively optimize the selection of defense timing
maintaining a stable protection state node ratio of 83.2%
reducing the attacker’s utility value by 44.85% compared to the MFD-PPO method
increasing the attack interception success rate by 38.4% compared to the FP method
and significantly reducing the IP hop frequency compared to the comparison method under the same protection state node ratio
verifying the comprehensive advantages of the proposed method in ensuring defense effectiveness and reducing resource overhead.
Wang Yingchao , Yang Chen , Lan Shulin , et al . End-edge-cloud collaborative computing for deep learning: A comprehensive survey [J ] . IEEE Communications Surveys & Tutorials , 2024 , 26 ( 4 ): 2647 - 2683 . DOI: 10.1109/comst.2024.3393230 http://dx.doi.org/10.1109/comst.2024.3393230
Chen Xiang , Guo Zhiheng , Wang Xijun , et al . Toward 6G native-AI network: Foundation model-based cloud-edge-end collaboration framework [J ] . IEEE Communications Magazine , 2025 , 63 ( 8 ): 23 - 30 . DOI: 10.1109/mcom.001.2400582 http://dx.doi.org/10.1109/mcom.001.2400582
Jamil M N , Schelén O , Monrat A A , et al . Enabling industrial internet of things by leveraging distributed edge-to-cloud computing: Challenges and opportunities [J ] . IEEE Access , 2024 , 12 : 127308 . DOI: 10.1109/access.2024.3454812 http://dx.doi.org/10.1109/access.2024.3454812
马博 , 余应洁 , 吴莎尘 , 等 . 云边端异构算力网络计算任务分割与路径优化方法研究 [J ] . 电子学报 , 2025 , 53 ( 6 ): 1847 - 1864 .
Ma Bo , Yu Yingjie , Wu Shachen , et al . Task segmentation and path optimization in heterogeneous cloud-edge-end computing power network [J ] . Acta Electronica Sinica , 2025 , 53 ( 6 ): 1847 - 1864 . (in Chinese)
Stephanie V , Khalil I , Rahman M S , et al . Privacy-preserving ensemble infused enhanced deep neural network framework for edge cloud convergence [J ] . IEEE Internet of Things Journal , 2023 , 10 ( 5 ): 3763 - 3773 . DOI: 10.1109/jiot.2022.3151982 http://dx.doi.org/10.1109/jiot.2022.3151982
Javadpour A , Ja’fari F , Taleb T , et al . A comprehensive survey on cyber deception techniques to improve honeypot performance [J ] . Computers & Security , 2024 , 140 : 103792 . DOI: 10.1016/j.cose.2024.103792 http://dx.doi.org/10.1016/j.cose.2024.103792
Gill K S , Sharma A , Saxena S . A systematic review on game-theoretic models and different types of security requirements in cloud environment: Challenges and opportunities [J ] . Archives of Computational Methods in Engineering , 2024 , 31 ( 7 ): 3857 - 3890 . DOI: 10.1007/s11831-024-10095-6 http://dx.doi.org/10.1007/s11831-024-10095-6
蒋侣 , 张恒巍 , 王晋东 . 基于多阶段Markov信号博弈的移动目标防御最优决策方法 [J ] . 电子学报 , 2021 , 49 ( 3 ): 527 - 535 . DOI: 10.12263/DZXB.20191070 http://dx.doi.org/10.12263/DZXB.20191070
JIANG Lü , ZHANG Hengwei , WANG Jindong . A Markov signaling game-theoretic approach to moving target defense strategy selection [J ] . Acta Electronica Sinica , 2021 , 49 ( 3 ): 527 - 535 . (in Chinese) . DOI: 10.12263/DZXB.20191070 http://dx.doi.org/10.12263/DZXB.20191070
Cai Feiyang , Koutsoukos X . Real-time detection of deception attacks in cyber-physical systems [J ] . International Journal of Information Security , 2023 , 22 ( 5 ): 1099 - 1114 . DOI: 10.1007/s10207-023-00677-z http://dx.doi.org/10.1007/s10207-023-00677-z
Chowdhury N H , Adam M T P , Teubner T . Time pressure in human cybersecurity behavior: Theoretical framework and countermeasures [J ] . Computers & Security , 2020 , 97 : 101963 . DOI: 10.1016/j.cose.2020.101963 http://dx.doi.org/10.1016/j.cose.2020.101963
Zhang Yunxiao , Malacaria P . Optimization-time analysis for cybersecurity [J ] . IEEE Transactions on Dependable and Secure Computing , 2022 , 19 ( 4 ): 2365 - 2383 . DOI: 10.1109/TDSC.2021.3055981 http://dx.doi.org/10.1109/TDSC.2021.3055981
Merlevede J , Johnson B , Grossklags J , et al . Time-dependent strategies in games of timing [C ] // Proceedings of the 10th International Conference on Decision and Game Theory for Security . Heidelberg : Springer , 2019 : 310 - 330 . DOI: 10.1007/978-3-030-32430-8_19 http://dx.doi.org/10.1007/978-3-030-32430-8_19
Zhang Hengwei , Tan Jinglei , Liu Xiaohu , et al . Moving target defense decision-making method: A dynamic Markov differential game model [C ] // Proceedings of the 7th ACM Workshop on Moving Target Defense . New York : ACM , 2020 : 21 - 29 . DOI: 10.1145/3411496.3421222 http://dx.doi.org/10.1145/3411496.3421222
Farhang S , Grossklags J . When to invest in security? Empirical evidence and a game-theoretic approach for time-based security[PP/OL ] . V2.arXiv ( 2017-06-06 )[ 2026-03-10 ] . https://arxiv.org/abs/1706.00302 https://arxiv.org/abs/1706.00302 . DOI: 10.1007/978-3-319-47413-7_12 http://dx.doi.org/10.1007/978-3-319-47413-7_12
Van Dijk M , Juels A , Oprea A , et al . FLIPIT: The game of “stealthy takeover” [J ] . Journal of Cryptology , 2013 , 26 ( 4 ): 655 - 713 . DOI: 10.1007/s00145-012-9134-5 http://dx.doi.org/10.1007/s00145-012-9134-5
Chen Xiaoguang , Cao Wenyuan , Chen Lili , et al . iCyberGuard: A FlipIt game for enhanced cybersecurity in IIoT [J ] . IEEE Transactions on Computational Social Systems , 2024 , 11 ( 6 ): 8005 - 8014 . DOI: 10.1109/TCSS.2024.3443174 http://dx.doi.org/10.1109/TCSS.2024.3443174
何威振 , 谭晶磊 , 张帅 , 等 . 利用深度强化学习的多阶段博弈网络拓扑欺骗防御方法 [J ] . 电子与信息学报 , 2024 , 46 ( 12 ): 4422 - 4431 .
He Weizhen , Tan Jinglei , Zhang Shuai , et al . Multi-stage game-based topology deception method using deep reinforcement learning [J ] . Journal of Electronics & Information Technology , 2024 , 46 ( 12 ): 4422 - 4431 . (in Chinese)
Tan Jinglei , Zhang Hengwei , Zhang Hongqi , et al . Optimal timing selection approach to moving target defense: A FlipIt attack-defense game model [J ] . Security and Communication Networks , 2020 , 2020 ( 1 ): 3151495 . DOI: 10.1155/2020/3151495 http://dx.doi.org/10.1155/2020/3151495
Zhu Zinan , Zhou Lin . Application of complex network attack and defense time game model in network security defense decision [J ] . Journal of Cyber Security and Mobility , 2025 , 14 ( 2 ): 311 - 338 .
熊鑫立 , 黄郡 , 姚倩 . 基于Cheat-FlipIt博弈的网络安全对抗建模与分析 [J ] . 信息对抗技术 , 2024 , 3 ( 4 ): 63 - 80 .
Xiong Xinli , Huang Jun , Yao Qian . Modeling and analysis of network security adversarial behavior based on Cheat-FlipIt game [J ] . Information Countermeasure Technology , 2024 , 3 ( 4 ): 63 - 80 . (in Chinese)
Zheng Yan , Na Zhenyu , Ji Weidong , et al . An adaptive fuzzy SIR model for real-time malware spread prediction in industrial internet of things networks [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 13 ): 22875 - 22888 . DOI: 10.1109/JIOT.2025.3550671 http://dx.doi.org/10.1109/JIOT.2025.3550671
Qi Ju . Loss and premium calculation of network nodes under the spread of SIS virus [J ] . Journal of Intelligent & Fuzzy Systems , 2023 , 44 ( 5 ): 7919 - 7933 . DOI: 10.3233/jifs-222308 http://dx.doi.org/10.3233/jifs-222308
Qiu Li , Xiang Caixia , Wen Yidan , et al . Predictive output feedback control of networked control system with Markov DoS attack and time delay [J ] . International Journal of Robust and Nonlinear Control , 2023 , 33 ( 5 ): 3376 - 3395 . DOI: 10.1002/rnc.6572 http://dx.doi.org/10.1002/rnc.6572
Sun Rongbo , Fei Jinlong , Zhu Yuefei , et al . Multi-agent reinforcement learning for moving target defense temporal decision-making approach based on Stackelberg-FlipIt games [J ] . Computers , Materials & Continua, 2025 , 84 ( 2 ): 3765 - 3786 . DOI: 10.32604/cmc.2025.064849 http://dx.doi.org/10.32604/cmc.2025.064849
He Weizhen , Tan Jinglei , Guo Yunfei , et al . Flipit game deception strategy selection method based on deep reinforcement learning [J ] . International Journal of Intelligent Systems , 2023 , 2023 ( 1 ): 5560416 . DOI: 10.1155/2023/5560416 http://dx.doi.org/10.1155/2023/5560416
Glicksberg I L . A further generalization of the Kakutani fixed point theorem, with application to Nash equilibrium points [J ] . Proceedings of the American Mathematical Society , 1952 , 3 ( 1 ): 170 - 174 . DOI: 10.2307/2032478 http://dx.doi.org/10.2307/2032478
0
浏览量
8
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621