Kuiper: Enterprise Cloud Plug-and-Play Network Management Platform
USENIX 网络系统设计与实现研讨会 (NSDI '27), 2027. 即将发表。
我的研究致力于为 四类生产场景 构建 测量、验证与态势感知基础设施:(i) 互联网——CT 证书透明度日志、BGP、RPKI、DNS、Web-PKI;(ii) ISP / 运营商网络,包括 FITI / CERNET / CERNET2 国家级试验设施与运营商骨干网;(iii) 智算 / 数据中心基础设施——超大规模云、Serverless 平台、智能计算中心;以及 (iv) 安全可信机器人,作为独立但互补的研究方向。研究工作组织为 三大方向,与下方 研究方向 一节一一对应:
这三条研究主线汇聚为一个统一议程:下一代互联网、国家级试验设施、运营智算 / 数据中心基础设施以及安全攸关的具身系统所需的可观测基底、验证工具链与信任管道。
在可观测性与路由技术方向,我开展了深入研究,并主导了 DeepShield 等项目。迄今已在 SIGCOMM、INFOCOM、CoNEXT、CCS 等国际顶级会议与期刊发表近百篇论文。研究成果已被工业界广泛采用,相关技术深度融入 中国电信骨干网、未来互联网试验设施(FITI)、CERNET / CERNET2 等国家关键网络基础设施,服务于电信运营商、国防军队及其他关乎国计民生的重要领域。
本人荣获 ACM SIGCSE 中国新星奖,担任 《通信技术》 等国内期刊编委,以及 MDPI Electronics、IEEE JSAC 等国际期刊的客座编辑;定期为 IEEE/ACM TON、IEEE TPDS、IEEE TIFS 等知名国际期刊审稿。此外,我积极参与 ACM SIGCOMM、ACM CoNEXT、IEEE INFOCOM 等国际重要会议的组织与程序委员会工作。研究形成的技术已服务于数亿用户。
“下一代 互联网 与 运营智算 / 数据中心基础设施 需要一个统一的 测量—验证—态势感知 基底。我们的使命就是把它建成。”
“作为独立但互补的方向,我们将同样的可信性原则延伸到 安全可信机器人。”
为国家级试验床(FITI / CERNET)与运营互联网构建测量基底与态势感知管线。
分布式追踪系统
SIGCOMM'23
SIGCOMM'25;
SSL / TLS 证书分析
WWW'25;
配置验证
TIFS'24;
恶意容器检测
CCS'22;
恶意流量检测
TIFS'22
TIFS'23
CCS'21
NDSS'22
SIGCOMM'25;
日志分析
TIFS'21;
APT 检测
TIFS'23;
流量测量
TPDS'21.
为安全攸关 AI、具身智能体和自主机器人系统构建 可信 ML、溯源与证书信任基底——将我们的线速检测、APT 防御与 PKI 验证工作扩展至赛博物理场景。
对抗鲁棒的流量 / 感知异常检测
SIGCOMM'25
CCS'21
NDSS'22;
APT 与溯源图异常检测
TIFS'22
TIFS'23;
面向安全攸关客户端的证书 / PKI 信任
WWW'25;
面向互联网可观测的结构-语义融合
WWW'26;
面向时延敏感遥测的网内低延迟 FEC
INFOCOM'24.
未来议程:可信多模态感知管线、加密传感器通道的异常检测、智能体 / 机器人控制器的运行时验证、自主车队的信任证书。
利用学习驱动与形式化方法,使试验床和运营网络可验证、可自诊、可韧性恢复。
深度学习加密流量异常检测
SIGCOMM'25
TIFS'22
CCS'21
NDSS'22;
AI 辅助的分布式追踪
SIGCOMM'23
SIGCOMM'25;
面向 Web 可观测的结构-语义融合
WWW'26;
SD-WAN 在线学习
TIFS'23;
APT 检测 with provenance graphs
TIFS'23
IEEE Network'26;
日志分析 with NLP
TIFS'21;
learning-driven 恶意容器检测
CCS'22;
网内 ML 加速
INFOCOM'24.
端到端生产级系统——赋能大规模可观测性、安全防护与智能运维。
生产级分布式追踪,跨异构服务重建端到端请求路径,无需修改任何应用代码。双路径 eBPF 内核 + 用户态协议解析实现方法级延迟估计,开销比侵入式基线降低约 10 倍。已部署于多个大规模智算中心集群。
端到端出入口安全防护:链路级流量可视化、实时态势感知和基于 ML 的威胁检测。运行于可编程交换机(P4 / Tofino)、DPU、SmartNIC和 eBPF——实现 Tbps 级线速加密流量分析,同时抵御对抗逃逸和概念漂移。
面向智算中心的全栈 AIOps 平台,集成故障管理、安全事件响应、配置管理和网络验证。基于 SRv6 切片并行探测实现实时拓扑感知、自动化根因分析和策略驱动修复,在生产 AI/GPU 集群环境中将 MTTR 降低 5 倍。
面向 AI 智能体时代的零代码 eBPF 可观测性:跨 LLM / 智能体流水线、Kubernetes、服务网格和无服务器平台重建因果追踪并发现安全相关事件,无需任何应用代码修改。已作为行业事实标准开源(CNCF Sandbox,4 k+ ★)。
面向 AI 数据中心和大规模微服务部署的分布式应用追踪与故障监控;同一非侵入式观测管线延伸至机器人 / 具身 AI 运行时安全——追踪控制器调用路径并检测异常,无需修改机器人软件栈。双路径内核 + 用户态协议解析,在生产集群中以 ~10 倍更少的插桩开销实现精确方法级延迟估计。
端到端出入口流量监测:链路级可视化、流量态势感知和安全事件检测。同一管线运行于可编程交换机、DPU、SmartNIC和eBPF——依托线速 ML IDS 研究(SIGCOMM '25 P4-IDS, SIGCOMM '26),已在运营商骨干网验证。
面向智算中心的统一智能运维平台,集成故障管理、安全事件响应和配置管理,构建在共享遥测与验证基座之上——依托 SRv6 切片并行探测(ICNP '25, ToN '26)、分布式追踪(SIGCOMM '25)和配置验证工具链。
持续的互联网全域主被动探测,覆盖 IP 空间、路由(BGP / RPKI)和证书(CT 日志 / Web-PKI)——构建可查询的互联网知识库,为下游测量、态势感知和信任验证管线提供支撑。
面向运营和试验网络的静态网络验证与仿真:配置验证器(TIFS '24, INFOCOM '26)、覆盖网集成检验(CoNEXT '25)和定量策略验证,加上 what-if 分析模拟器——通过环境缩减和集成检验扩展到运营商规模。
面向大规模网络诊断的切片并行 SRv6 探测:利用 SRv6 段路由在单轮探测中同时覆盖多条覆盖路径,实现全网格延迟图的 O(1) 轮次。已在国家 FITI 试验设施验证,部署于运营商网络实时瓶颈定位。
研究成果已融入国家级网络基础设施与一线产业产品, 服务电信运营商、教育科研网络、公有云与安全厂商。
线速加密流量异常检测;大规模 WAN 感知与优化。
可编程交换机 IDS 与段路由诊断已部署于国家级下一代互联网试验设施。
面向中国教育与科研计算机网的 DeepFlow 可观测性与路由异常分析。
大规模微服务系统可观测性;联合设计跨数据中心 WAN 流量工程。
面向企业 SD-WAN 产品的应用感知 WAN 传输优化。
安全测试设备与大规模流量分析的关键技术。
* 通讯作者 · † 共同一作 · 下划线 表示指导的学生。 Full list on Google Scholar.
USENIX 网络系统设计与实现研讨会 (NSDI '27), 2027. 即将发表。
ACM 数据通信专业委员会年会 (SIGCOMM '26), 2026. 即将发表。
Internet Service Providers are uniquely positioned to deliver intrusion detection at the network's choke points, but operating an IDS at carrier scale faces stringent throughput, accuracy, and robustness constraints that conventional middlebox or host-based solutions cannot meet. This work proposes an ISP-centric IDS architecture built on programmable switches: it co-designs lightweight feature extraction in the data plane with adaptive learning in the control plane to sustain Tbps-class detection while remaining robust to traffic drift and adversarial evasion. The system has been validated in real ISP backbones and demonstrates orders-of-magnitude resource savings over CPU/GPU baselines without sacrificing detection quality.
ACM 数据通信专业委员会年会 (SIGCOMM '25), pp. 1056–1069, 2025.
As microservices grow in scale and complexity, their operation and debugging become increasingly challenging. Even a single user request can involve interactions across hundreds of components. In such intricate systems, distributed tracing, which tracks the end-to-end execution flow of requests, has become a critical monitoring tool. Among these, non-intrusive tracing frameworks that do not require code modification are particularly valued for their convenience. However, existing non-intrusive solutions either have limited applicability or lack sufficient accuracy under high concurrency. To address these challenges, we propose DeepTrace, a transaction-based, non-intrusive distributed tracing framework designed for microservices. DeepTrace leverages API endpoints and transaction fields embedded within request content to categorize requests into distinct transactions, thereby reducing the likelihood of incorrectly merging traces from different transactions. Compared to state-of-the-art frameworks, DeepTrace maintains an accuracy rate of over 95% even under high concurrency. It has also been adopted by dozens of companies in their production systems for tasks such as failure diagnosis and resource optimization.
@inproceedings{Geng_2025, series={SIGCOMM ’25}, title={Low-Overhead Distributed Application Observation with DeepTrace: Achieving Accurate Tracing in Production Systems}, url={http://dx.doi.org/10.1145/3718958.3750477}, DOI={10.1145/3718958.3750477}, booktitle={Proceedings of the ACM SIGCOMM 2025 Conference}, publisher={ACM}, author={Geng, Yantao and Zhang, Han and Wu, Zhiheng and Li, Yahui and Wang, Jilong and Yin, Xia}, year={2025}, month=Aug, pages={1056–1069}, collection={SIGCOMM ’25} }ACM 数据通信专业委员会年会 (SIGCOMM '25), pp. 1254–1256, 2025.
Attacks against data centers are becoming more common as a result of the fast expansion of applications. In order to keep pace with the growing amount of data centers connected to their networks, internet service providers must offer comprehensive security services. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for the high-speed encrypted network traffic. In this paper, we design and implement Mazu, an inline network intrusion detection system with programmable switches specifically developed to protect data centers connecting to the internet service provider. Mazu proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. Mazu also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify the malicious traffic. In addition, Mazu introduces an online update mechanism aimed at dynamically adjusting the detection model in response to environmental changes. Mazu has been in production for two years, during which time it has identified over 10 critical attack events and protect more than 10 million servers for two ISPs. Our production and testbed evaluations demonstrate that Mazu can detect malicious traffic entering the data center sites with approximately 90% accuracy within minutes.
@inproceedings{Zhang_2025, series={SIGCOMM ’25}, title={Achieving High-Speed and Robust Encrypted Traffic Anomaly Detection with Programmable Switches}, url={http://dx.doi.org/10.1145/3718958.3750493}, DOI={10.1145/3718958.3750493}, booktitle={Proceedings of the ACM SIGCOMM 2025 Conference}, publisher={ACM}, author={Zhang, Han and Liu, Guyue and Shi, Xingang and Li, Yahui and He, Dongbiao and Wang, Jilong and Wang, Zhiliang and Zhu, Yongqing and Ruan, Ke and Cao, Weihua and Yin, Xia}, year={2025}, month=Aug, pages={1254–1256}, collection={SIGCOMM ’25} }ACM 数据通信专业委员会年会 (SIGCOMM '23), pp. 420–437, 2023.
Microservices are becoming more complicated, posing new challenges for traditional performance monitoring solutions. On the one hand, the rapid evolution of microservices places a significant burden on the utilization and maintenance of existing distributed tracing frameworks. On the other hand, complex infrastructure increases the probability of network performance problems and creates more blind spots on the network side. In this paper, we present DeepFlow, a network-centric distributed tracing framework for troubleshooting microservices. DeepFlow provides out-of-the-box tracing via a network-centric tracing plane and implicit context propagation. In addition, it eliminates blind spots in network infrastructure, captures network metrics in a low-cost way, and enhances correlation between different components and layers. We demonstrate analytically and empirically that DeepFlow is capable of locating microservice performance anomalies with negligible overhead. DeepFlow has already identified over 71 critical performance anomalies for more than 26 companies and has been utilized by hundreds of individual developers. Our production evaluations demonstrate that DeepFlow is able to save users hours of instrumentation efforts and reduce troubleshooting time from several hours to just a few minutes.
@inproceedings{Shen_2023, series={ACM SIGCOMM ’23}, title={Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero Code}, url={http://dx.doi.org/10.1145/3603269.3604823}, DOI={10.1145/3603269.3604823}, booktitle={Proceedings of the ACM SIGCOMM 2023 Conference}, publisher={ACM}, author={Shen, Junxian and Zhang, Han and Xiang, Yang and Shi, Xingang and Li, Xinrui and Shen, Yunxi and Zhang, Zijian and Wu, Yongxiang and Yin, Xia and Wang, Jilong and Xu, Mingwei and Li, Yahui and Yin, Jiping and Song, Jianchang and Li, Zhuofeng and Nie, Runjie}, year={2023}, month=Sept, pages={420–437}, collection={ACM SIGCOMM ’23} }ACM 计算机与通信安全会议 (CCS '22), pp. 2627–2641, 2022.
Serverless computing, or Function-as-a-Service, is gaining continuous popularity due to its pay-as-you-go billing model, flexibility, and low costs. These characteristics, however, bring additional security risks, such as the Denial-of-Wallet (DoW) attack, to serverless tenants. In this paper, we perform a real-world DoW attack on commodity serverless platforms to evaluate its severity. To identify such attacks, we design, implement, and evaluate Gringotts, an accurate, easy-to-use DoW detection system with a negligible performance overhead. Gringotts addresses the information ambiguity inherent in serverless functions by introducing a well-designed performance metrics collection agent. Then, Gringotts uses the Mahalanobis distance to discover anomalies in the distribution of the metrics. We implement Gringotts as a real system and conduct extensive experiments using a testbed to evaluate the performance of Gringotts. Our results indicate that Gringotts has a performance overhead of less than 1.1%, with an average detection delay of 1.86 seconds and an average accuracy of over 95.75%.
@inproceedings{Shen_2022, series={CCS ’22}, title={Gringotts: Fast and Accurate Internal Denial-of-Wallet Detection for Serverless Computing}, url={http://dx.doi.org/10.1145/3548606.3560629}, DOI={10.1145/3548606.3560629}, booktitle={Proceedings of the 2022 ACM 计算机与通信安全会议}, publisher={ACM}, author={Shen, Junxian and Zhang, Han and Geng, Yantao and Li, Jiawei and Wang, Jilong and Xu, Mingwei}, year={2022}, month=Nov, pages={2627–2641}, collection={CCS ’22} }ACM CoNEXT '21, 2021.
Inter-DataCenter Wide Area Network (Inter-DC WAN) that connects geographically distributed data centers is becoming one of the most critical network infrastructures. Due to limited bandwidth and inevitable link failures, it is highly challenging to guarantee network availability for services, especially those with stringent bandwidth demands, over inter-DC WAN. We present BATE, a novel Traffic Engineering (TE) framework for bandwidth availability (BA) provision, which aims to ensure that each bandwidth demand must be satisfied with a stipulated probability, when subjected to the network capacity and possible failures of the inter-DC WAN. The three core components of BATE, i.e., admission control, traffic scheduling and failure recovery, are formulated through different mathematical models and theoretically analyzed. They are also extensively compared against state-of-the-art TE schemes, using a testbed as well as real trace driven simulations across different topologies, traffic matrices and failure scenarios. Our evaluations show that, compared with the optimal admission strategy, BATE can speed up the online admission control by 30x at the expense of less than 4% false rejections. On the other hand, compared with the latest TE schemes like FFC and TEAVAR, BATE can meet the bandwidth availability targets for 23%~60% more demands under normal loads, and when network failure causes BA targets violations.
@inproceedings{Zhang_2021, series={CoNEXT ’21}, title={Boosting bandwidth availability over inter-DC WAN}, url={http://dx.doi.org/10.1145/3485983.3494843}, DOI={10.1145/3485983.3494843}, booktitle={Proceedings of the 17th International Conference on emerging Networking EXperiments and Technologies}, publisher={ACM}, author={Zhang, Han and Shi, Xingang and Yin, Xia and Wang, Jilong and Wang, Zhiliang and Guo, Yingya and Lan, Tian}, year={2021}, month=Dec, pages={297–312}, collection={CoNEXT ’21} }ACM / IEEE 国际软件工程会议 (ICSE '26). 即将发表。
Pinpointing latency culprits inside complex microservice call graphs traditionally requires invasive code instrumentation. This work introduces a fully non-intrusive tracing technique that infers method-level delays from network-side observations and lightweight runtime hints, eliminating the need to modify application code. Evaluation on production-scale microservice deployments demonstrates fault-localization accuracy on par with intrusive baselines at a fraction of the engineering cost.
ACM Web 会议论文集 (WWW '25), 2025.
The Automatic Certificate Management Environment (ACME) protocol automates the issuance and renewal of secure socket layer certificates, simplifying the management of large-scale certificate deployments. To reduce the load on Certificate Authority (CA) servers, ACME employs a caching mechanism that stores domain validation (DV) results for 30 days. However, this mechanism allows attackers to reuse previously authorized results, potentially bypassing the DV process. In this paper, we introduce the ACME Authz Cache Attack, whereby an attacker can obtain fraudulent certificates without domain control. We demonstrate that even the prominent CA, Let's Encrypt, is vulnerable to this attack. To mitigate this, we propose ACME++, an enhanced protocol that binds the client's IP address and a unique identifier to the ACME account, ensuring secure authorization for each new client and effectively preventing the ACME Authz Cache Attack. Our implementation of ACME++ shows that it introduces little overhead to the CA server.
@inproceedings{Zhang_2025, series={WWW ’25}, title={ACME++: A Secure Authorization Mechanism for ACME Clients in the Web PKI Ecosystem}, url={http://dx.doi.org/10.1145/3696410.3714763}, DOI={10.1145/3696410.3714763}, booktitle={Proceedings of the ACM on Web Conference 2025}, publisher={ACM}, author={Zhang, Tianyu and Zhang, Han and Wei, Yunze and Li, Yahui and Shi, Xingang and Wang, Jilong and Yin, Xia}, year={2025}, month=Apr, pages={1058–1067}, collection={WWW ’25} }ACM Web 会议论文集 (WWW '26), pp. 7576–7587, 2026.
Looking Glass (LG) services expose a fragmented but invaluable view of the global Internet's routing fabric, yet their heterogeneous interfaces and free-form output have long resisted automated analysis. GlassMiner proposes a structure-semantics fusion approach that jointly models the DOM-level layout and natural-language hints of LG portals to extract structured measurements at scale. The system enables continent-wide BGP path observability and uncovers previously hidden routing anomalies across thousands of operator portals.
IEEE INFOCOM 2024, pp. 420–437, 2024.
Microservices are becoming more complicated, posing new challenges for traditional performance monitoring solutions. On the one hand, the rapid evolution of microservices places a significant burden on the utilization and maintenance of existing distributed tracing frameworks. On the other hand, complex infrastructure increases the probability of network performance problems and creates more blind spots on the network side. In this paper, we present DeepFlow, a network-centric distributed tracing framework for troubleshooting microservices. DeepFlow provides out-of-the-box tracing via a network-centric tracing plane and implicit context propagation. In addition, it eliminates blind spots in network infrastructure, captures network metrics in a low-cost way, and enhances correlation between different components and layers. We demonstrate analytically and empirically that DeepFlow is capable of locating microservice performance anomalies with negligible overhead. DeepFlow has already identified over 71 critical performance anomalies for more than 26 companies and has been utilized by hundreds of individual developers. Our production evaluations demonstrate that DeepFlow is able to save users hours of instrumentation efforts and reduce troubleshooting time from several hours to just a few minutes.
@inproceedings{Shen_2023, series={ACM SIGCOMM ’23}, title={Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero Code}, url={http://dx.doi.org/10.1145/3603269.3604823}, DOI={10.1145/3603269.3604823}, booktitle={Proceedings of the ACM SIGCOMM 2023 Conference}, publisher={ACM}, author={Shen, Junxian and Zhang, Han and Xiang, Yang and Shi, Xingang and Li, Xinrui and Shen, Yunxi and Zhang, Zijian and Wu, Yongxiang and Yin, Xia and Wang, Jilong and Xu, Mingwei and Li, Yahui and Yin, Jiping and Song, Jianchang and Li, Zhuofeng and Nie, Runjie}, year={2023}, month=Sept, pages={420–437}, collection={ACM SIGCOMM ’23} }IEEE 国际网络协议会议 (ICNP '25), pp. 1–12, 2025.
Existing latency-diagnosis tools for SRv6 either flood the data plane with probes or return probabilistic verdicts. SRmesh embeds deterministic in-band telemetry into segment-routing headers, enabling per-link latency attribution with bounded measurement overhead. It has been validated in a multi-vendor SRv6 testbed and identifies bottleneck links in seconds with 100% recall.
IEEE INFOCOM 2026.
IEEE INFOCOM 2025, pp. 420–437.
ACM CoNEXT 2025.
Overlay/underlay architecture is increasingly prevalent in modern networks but also introduces greater complexity and error-proneness. Existing control plane verifiers for underlay networks face scalability and interpretability challenges when extended to overlay/underlay networks due to methodological limitations. This paper presents MEV, the first control plane verifier designed for overlay/underlay networks. MEV introduces ensemble verification, a novel verification paradigm enabling independent reasoning for behaviors within each routing instance or protocol and forwarding behaviors within each virtual network. We implement and deploy MEV on a real nationwide overlay/underlay network, FITI, where it successfully identifies 22 misconfigurations that could cause isolation and reachability issues. Furthermore, we evaluate MEV against Batfish+ on FITI and a range of synthetic networks with diverse scales, and find that it achieves up to a 102× speedup over Batfish+.
@article{Liu_2025, title={Scalable and 可解释 Overlay Network Checking via Ensemble Verification}, volume={3}, ISSN={2834-5509}, url={http://dx.doi.org/10.1145/3768974}, DOI={10.1145/3768974}, number={CoNEXT4}, journal={Proceedings of the ACM on Networking}, publisher={Association for Computing Machinery (ACM)}, author={Liu, XinZhe and Li, Yahui and Zhang, Han and Yin, Xia and Shi, Xingang and Wang, Zhiliang and Ren, Gang and Wang, Jilong and Yao, Jiangyuan}, year={2025}, month=Nov, pages={1–25} }网络与分布式系统安全研讨会 (NDSS '22), 2022.
IEEE/ACM 网络汇刊, 2026.
Serverless computing, or Function-as-a-Service, continues to gain popularity due to its pay-as-you-go billing model, flexibility, and cost efficiency. However, these same features introduce significant security risks, such as the Denial-of-Wallet (DoW) attack. In this paper, we conduct real-world DoW attacks on commercial serverless platforms to evaluate their severity. To detect such attacks, we design, implement, and evaluate Clover, an accurate and user-friendly DoW detection system with negligible performance overhead. Clover addresses information ambiguity in serverless environments by deploying a request-oriented metric collection agent. At its core, Clover proposes a workload verification approach to bridge performance metrics and execution duration. Specifically, Clover uses a multivariate linear model to learn the benign relationship between metrics and execution duration, effectively characterizing normal workload behavior. It then continuously monitors runtime workloads by calculating their Mahalanobis distance from this learned benign model. Deviations identified through this distance indicate potential DoW attacks. Implemented as a practical system, Clover introduces performance overhead of less than 3.2%, maintains an average model execution time of only 0.84 microseconds, and achieves an accuracy of 92.7% under the most challenging scenario.
@article{Shen_2026, title={Clover: Workload Verification for Real-Time Detection of Contention-Induced Slowdowns in Serverless Platforms}, volume={34}, ISSN={2998-4157}, url={http://dx.doi.org/10.1109/TON.2026.3680829}, DOI={10.1109/ton.2026.3680829}, journal={IEEE Transactions on Networking}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Shen, Junxian and Zhang, Han and Lin, Weiwei and Geng, Yantao and Li, Jiawei and Wang, Jilong and Xu, Mingwei}, year={2026}, pages={4700–4715} }IEEE/ACM 网络汇刊, 2026.
IEEE Transactions on Networking, 2025.
IEEE Transactions on Networking, 2025.
IEEE/ACM 网络汇刊, 2022.
IEEE 并行与分布式系统汇刊, vol. 34, no. 8, pp. 2448–2463, 2023.
IEEE 并行与分布式系统汇刊, vol. 34, no. 4, pp. 1362–1375, 2023.
IEEE 并行与分布式系统汇刊, vol. 33, no. 10, pp. 2571–2583, 2022.
IEEE 并行与分布式系统汇刊, vol. 33, no. 5, pp. 1069–1083, 2022.
IEEE 信息取证与安全汇刊, 2025.
IEEE 信息取证与安全汇刊, vol. 19, pp. 10099–10113, 2024.
IEEE 信息取证与安全汇刊, vol. 18, pp. 2076–2090, 2023.
IEEE 信息取证与安全汇刊, vol. 18, pp. 1610–1624, 2023.
IEEE 信息取证与安全汇刊, vol. 17, pp. 3972–3987, 2022.
IEEE 信息取证与安全汇刊, vol. 16, pp. 2300–2311, 2021.
IEEE 国际网络协议会议 (ICNP '22), 2022.
ACM 计算机与通信安全会议 (CCS '21), 2021.
Unsupervised Deep Learning (DL) techniques have been widely used in various security-related anomaly detection applications, owing to the great promise of being able to detect unforeseen threats and superior performance provided by Deep Neural Networks (DNN). However, the lack of interpretability creates key barriers to the adoption of DL models in practice. Unfortunately, existing interpretation approaches are proposed for supervised learning models and/or non-security domains, which are unadaptable for unsupervised DL models and fail to satisfy special requirements in security domains. In this paper, we propose DeepAID, a general framework aiming to (1) interpret DL-based anomaly detection systems in security domains, and (2) improve the practicality of these systems based on the interpretations. We first propose a novel interpretation method for unsupervised DNNs by formulating and solving well-designed optimization problems with special constraints for security domains. Then, we provide several applications based on our Interpreter as well as a model-based extension Distiller to improve security systems by solving domain-specific problems. We apply DeepAID over three types of security-related anomaly detection systems and extensively evaluate our Interpreter with representative prior works. Experimental results show that DeepAID can provide high-quality interpretations for unsupervised DL models while meeting the special requirements of security domains. We also provide several use cases to show that DeepAID can help security operators to understand model decisions, diagnose system mistakes, give feedback to models, and reduce false positives.
@inproceedings{Han_2021, series={CCS ’21}, title={DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications}, url={http://dx.doi.org/10.1145/3460120.3484589}, DOI={10.1145/3460120.3484589}, booktitle={Proceedings of the 2021 ACM 计算机与通信安全会议}, publisher={ACM}, author={Han, Dongqi and Wang, Zhiliang and Chen, Wenqi and Zhong, Ying and Wang, Su and Zhang, Han and Yang, Jiahai and Shi, Xingang and Yin, Xia}, year={2021}, month=Nov, pages={3197–3217}, collection={CCS ’21} }