Kuiper: Enterprise Cloud Plug-and-Play Network Management Platform
USENIX Symposium on Networked Systems Design and Implementation (NSDI '27), 2027. To appear.
My research builds the measurement, verification, and situational-awareness infrastructure for four production scenarios: (i) the Internet — CT logs, BGP, RPKI, DNS, Web-PKI; (ii) ISP & carrier networks, including the FITI / CERNET / CERNET2 national testbeds and operator backbones; (iii) AI / data-center infrastructures — hyperscale clouds, serverless platforms, and intelligent-computing centers; and (iv) secure & trustworthy robotics as a separate but complementary thrust. The work is organized into three directions, mirrored in the Research Areas section below where each is broken down by venue and project:
These three lines converge into a single agenda: the observability substrate, verification toolchain, and trust pipeline that the next-generation Internet, national testbeds, operational AI / data-center infrastructures, and safety-critical embodied systems will require.
In the areas of observability and routing technologies, I have conducted in-depth research and led projects such as DeepShield. To date, I have published nearly 100 papers in top-tier international conferences and journals, including SIGCOMM, INFOCOM, CoNEXT, and CCS. My research outcomes have been widely adopted in industry, with related technologies deeply integrated into critical national network infrastructures such as China Telecom's backbone network, the Future Internet Technology Infrastructure (FITI), and CERNET / CERNET2. These contributions serve vital sectors including telecommunications operators, the military, and other areas crucial to the national economy and people's livelihoods.
I have received the ACM SIGCSE China Rising Star Award and serve as an editorial board member for domestic journals such as Communication Technology, as well as a guest editor for international journals including MDPI Electronics and IEEE JSAC. I also regularly review for prestigious international journals such as IEEE/ACM TON, IEEE TPDS, and IEEE TIFS. In addition, I have been actively involved in organizing and serving on program committees for major international conferences, including ACM SIGCOMM, ACM CoNEXT, and IEEE INFOCOM. The technologies developed through my research have been applied to serve hundreds of millions of users.
“The next generation of the Internet and operational AI / data-center infrastructures needs a unified measurement–verification–situational-awareness substrate. Our mission is to build it.”
“As a separate but complementary thrust, we extend the same trustworthiness principles to secure & trustworthy robotics.”
Building the measurement substrate and situational-awareness pipelines for national-scale testbeds (FITI/CERNET) and operational Internet.
Distributed tracing systems
SIGCOMM'23
SIGCOMM'25;
SSL/TLS certificate analysis
WWW'25;
configuration verification
TIFS'24;
malicious container detection
CCS'22;
malicious traffic detection
TIFS'22
TIFS'23
CCS'21
NDSS'22
SIGCOMM'25;
log analysis
TIFS'21;
APT detection
TIFS'23;
flow measurement
TPDS'21.
Building the trustworthy ML, provenance, and certificate-trust substrate needed for safety-critical AI, embodied agents, and autonomous robotic systems — extending our line-rate detection, APT defense, and PKI verification work to cyber-physical settings.
Adversarial-robust traffic / perception anomaly detection
SIGCOMM'25
CCS'21
NDSS'22;
APT & provenance-graph anomaly detection
TIFS'22
TIFS'23;
certificate / PKI trust for safety-critical clients
WWW'25;
structure-semantics fusion for Internet observability
WWW'26;
in-network low-latency FEC for time-critical telemetry
INFOCOM'24.
Forward agenda: trustworthy multi-modal perception pipelines, encrypted sensor-channel anomaly detection, runtime verification of agentic / robotic controllers, and certificate-of-trust for autonomous fleets.
Making testbed and operational networks verifiable, self-diagnosable, and resilient using learning- and formal-method-driven techniques.
Deep-learning encrypted-traffic anomaly detection
SIGCOMM'25
TIFS'22
CCS'21
NDSS'22;
AI-assisted distributed tracing
SIGCOMM'23
SIGCOMM'25;
structure-semantics fusion for web observability
WWW'26;
online learning for SD-WAN
TIFS'23;
APT detection with provenance graphs
TIFS'23
IEEE Network'26;
log analysis with NLP
TIFS'21;
learning-driven malicious container detection
CCS'22;
in-network ML acceleration
INFOCOM'24.
End-to-end systems deployed in production environments — powering observability, security, and intelligent operations at scale.
Production-grade distributed tracing that reconstructs end-to-end request paths across heterogeneous services without any application code changes. Dual-path eBPF kernel + user-space protocol parsing delivers method-level delay estimation at ~10× less overhead than intrusive baselines. Deployed across multiple large-scale AI data-center clusters.
End-to-end perimeter security for AI data-centers: link-level traffic visualization, real-time situational awareness, and ML-based threat detection. Runs natively on programmable switches (P4 / Tofino), DPUs, SmartNICs, and eBPF — achieving Tbps-class line-rate encrypted-traffic analysis while remaining robust to adversarial evasion and concept drift.
Full-stack AIOps platform integrating fault management, security-event response, configuration management, and network verification into a single control plane. Built on SRv6 slice-parallel probing for real-time topology awareness, automated root-cause analysis, and policy-driven remediation — reducing MTTR by 5× in production AI/GPU cluster environments.
A security-first runtime observation and protection framework for autonomous robots and embodied-AI agents. Monitors controller call paths, detects anomalous behavior patterns, and enforces safety invariants in real-time — without modifying the robot software stack. Designed for industrial manipulators, autonomous vehicles, and multi-agent swarm deployments.
Open-source code, datasets, and research artifacts contributed back to the community.
Zero-code eBPF observability repurposed for the AI-agent era: reconstructs causal traces and surfaces security-relevant events across LLM / agent pipelines, Kubernetes, service mesh and serverless platforms without any application code change. Open-sourced as an industry de-facto standard (CNCF Sandbox, 4 k+ ★).
Distributed application tracing & fault monitoring for AI data centers and large-scale microservice deployments; the same non-intrusive observation pipeline extends to robotic / embodied-AI runtime security — tracing controller call paths and detecting anomalies without modifying robot software stacks. Dual-path kernel + user-space protocol parsing yields accurate method-level delay estimation in production clusters at ~10× less instrumentation overhead.
End-to-end ingress / egress traffic monitoring for AI / data-center sites: link-level visualization, traffic-state awareness, and security-event detection. The same pipeline runs across programmable switches, DPUs, SmartNICs, and eBPF — backed by our line-rate ML-based IDS work (SIGCOMM '25 P4-IDS, SIGCOMM '26) and validated on real operator backbones.
Unified intelligent operations platform for AI / data-center infrastructures, integrating fault management, security-event response, and configuration management over the same telemetry & verification substrate — underpinned by our SRv6 slice-parallel probing (ICNP '25, ToN '26), distributed tracing (SIGCOMM '25), and configuration-verification toolchains.
Continuous Internet-wide active & passive probing covering IP space, routing (BGP / RPKI), and certificates (CT logs / Web-PKI) — assembled into a queryable Internet knowledge base that powers downstream measurement, situational-awareness, and trust-verification pipelines. Builds on ACME++ (WWW '25) for PKI trust and GlassMiner (WWW '26) for Looking-Glass routing intelligence.
Static network verification and simulation for operational and testbed networks: a family of verifiers for configurations (TIFS '24, INFOCOM '26), overlay networks (CoNEXT '25), and quantitative policies, plus a simulator for what-if analysis — scaling to carrier-sized state via environment reduction and ensemble checking.
Slice-parallel SRv6 probing for large-scale network diagnosis: exploits SRv6 segment routing to simultaneously probe multiple overlay paths in a single pass, achieving O(1) rounds for full-mesh latency maps. Validated on the national FITI testbed and deployed for real-time bottleneck identification in carrier networks.
Structured corpus of public Looking-Glass service responses across hundreds of ASes (WWW '26).
Labeled encrypted-flow datasets used in SiteGuard (SIGCOMM '25 / '26) and prior CCS '21 / NDSS '22 anomaly-detection studies.
Undergraduate course materials & labs (32 h, Autumn '23 / '24), covering routing, traffic engineering, observability, and Internet measurement.
Research outcomes integrated into national network infrastructures and tier-1 industry products, serving telecom operators, education networks, public clouds, and security vendors.
Encrypted-traffic anomaly detection at line rate; large-scale WAN sensing & optimization.
Programmable-switch IDS & segment-routing diagnosis deployed on the national next-generation Internet testbed.
DeepFlow-based observability and routing-anomaly analysis across the China Education and Research Network.
Observability for large-scale microservice systems; co-designed traffic engineering for inter-DC WAN.
Application-aware WAN transmission optimization for enterprise SD-WAN products.
Key technologies for security-testing facilities and large-scale traffic analysis.
* corresponding author · † co-primary authors · underline marks supervised students. Full list on Google Scholar.
USENIX Symposium on Networked Systems Design and Implementation (NSDI '27), 2027. To appear.
Annual Conference of the ACM Special Interest Group on Data Communication (SIGCOMM '26), 2026. To appear.
Internet Service Providers are uniquely positioned to deliver intrusion detection at the network's choke points, but operating an IDS at carrier scale faces stringent throughput, accuracy, and robustness constraints that conventional middlebox or host-based solutions cannot meet. This work proposes an ISP-centric IDS architecture built on programmable switches: it co-designs lightweight feature extraction in the data plane with adaptive learning in the control plane to sustain Tbps-class detection while remaining robust to traffic drift and adversarial evasion. The system has been validated in real ISP backbones and demonstrates orders-of-magnitude resource savings over CPU/GPU baselines without sacrificing detection quality.
Annual Conference of the ACM Special Interest Group on Data Communication (SIGCOMM '25), pp. 1056–1069, 2025.
As microservices grow in scale and complexity, their operation and debugging become increasingly challenging. Even a single user request can involve interactions across hundreds of components. In such intricate systems, distributed tracing, which tracks the end-to-end execution flow of requests, has become a critical monitoring tool. Among these, non-intrusive tracing frameworks that do not require code modification are particularly valued for their convenience. However, existing non-intrusive solutions either have limited applicability or lack sufficient accuracy under high concurrency. To address these challenges, we propose DeepTrace, a transaction-based, non-intrusive distributed tracing framework designed for microservices. DeepTrace leverages API endpoints and transaction fields embedded within request content to categorize requests into distinct transactions, thereby reducing the likelihood of incorrectly merging traces from different transactions. Compared to state-of-the-art frameworks, DeepTrace maintains an accuracy rate of over 95% even under high concurrency. It has also been adopted by dozens of companies in their production systems for tasks such as failure diagnosis and resource optimization.
@inproceedings{Geng_2025, series={SIGCOMM ’25}, title={Low-Overhead Distributed Application Observation with DeepTrace: Achieving Accurate Tracing in Production Systems}, url={http://dx.doi.org/10.1145/3718958.3750477}, DOI={10.1145/3718958.3750477}, booktitle={Proceedings of the ACM SIGCOMM 2025 Conference}, publisher={ACM}, author={Geng, Yantao and Zhang, Han and Wu, Zhiheng and Li, Yahui and Wang, Jilong and Yin, Xia}, year={2025}, month=Aug, pages={1056–1069}, collection={SIGCOMM ’25} }Annual Conference of the ACM Special Interest Group on Data Communication (SIGCOMM '25), pp. 1254–1256, 2025.
Attacks against data centers are becoming more common as a result of the fast expansion of applications. In order to keep pace with the growing amount of data centers connected to their networks, internet service providers must offer comprehensive security services. However, existing network intrusion detection systems (NIDS) are either ineffective or inefficient for the high-speed encrypted network traffic. In this paper, we design and implement Mazu, an inline network intrusion detection system with programmable switches specifically developed to protect data centers connecting to the internet service provider. Mazu proposes a dual-plane feature extraction model to extract extensive traffic features at near line-speed. Mazu also proposes a lightweight one-class classification model that trains the best parameters exclusively on benign traffic to identify the malicious traffic. In addition, Mazu introduces an online update mechanism aimed at dynamically adjusting the detection model in response to environmental changes. Mazu has been in production for two years, during which time it has identified over 10 critical attack events and protect more than 10 million servers for two ISPs. Our production and testbed evaluations demonstrate that Mazu can detect malicious traffic entering the data center sites with approximately 90% accuracy within minutes.
@inproceedings{Zhang_2025, series={SIGCOMM ’25}, title={Achieving High-Speed and Robust Encrypted Traffic Anomaly Detection with Programmable Switches}, url={http://dx.doi.org/10.1145/3718958.3750493}, DOI={10.1145/3718958.3750493}, booktitle={Proceedings of the ACM SIGCOMM 2025 Conference}, publisher={ACM}, author={Zhang, Han and Liu, Guyue and Shi, Xingang and Li, Yahui and He, Dongbiao and Wang, Jilong and Wang, Zhiliang and Zhu, Yongqing and Ruan, Ke and Cao, Weihua and Yin, Xia}, year={2025}, month=Aug, pages={1254–1256}, collection={SIGCOMM ’25} }Annual Conference of the ACM Special Interest Group on Data Communication (SIGCOMM '23), pp. 420–437, 2023.
Microservices are becoming more complicated, posing new challenges for traditional performance monitoring solutions. On the one hand, the rapid evolution of microservices places a significant burden on the utilization and maintenance of existing distributed tracing frameworks. On the other hand, complex infrastructure increases the probability of network performance problems and creates more blind spots on the network side. In this paper, we present DeepFlow, a network-centric distributed tracing framework for troubleshooting microservices. DeepFlow provides out-of-the-box tracing via a network-centric tracing plane and implicit context propagation. In addition, it eliminates blind spots in network infrastructure, captures network metrics in a low-cost way, and enhances correlation between different components and layers. We demonstrate analytically and empirically that DeepFlow is capable of locating microservice performance anomalies with negligible overhead. DeepFlow has already identified over 71 critical performance anomalies for more than 26 companies and has been utilized by hundreds of individual developers. Our production evaluations demonstrate that DeepFlow is able to save users hours of instrumentation efforts and reduce troubleshooting time from several hours to just a few minutes.
@inproceedings{Shen_2023, series={ACM SIGCOMM ’23}, title={Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero Code}, url={http://dx.doi.org/10.1145/3603269.3604823}, DOI={10.1145/3603269.3604823}, booktitle={Proceedings of the ACM SIGCOMM 2023 Conference}, publisher={ACM}, author={Shen, Junxian and Zhang, Han and Xiang, Yang and Shi, Xingang and Li, Xinrui and Shen, Yunxi and Zhang, Zijian and Wu, Yongxiang and Yin, Xia and Wang, Jilong and Xu, Mingwei and Li, Yahui and Yin, Jiping and Song, Jianchang and Li, Zhuofeng and Nie, Runjie}, year={2023}, month=Sept, pages={420–437}, collection={ACM SIGCOMM ’23} }ACM SIGSAC Conference on Computer and Communications Security (CCS '22), pp. 2627–2641, 2022.
Serverless computing, or Function-as-a-Service, is gaining continuous popularity due to its pay-as-you-go billing model, flexibility, and low costs. These characteristics, however, bring additional security risks, such as the Denial-of-Wallet (DoW) attack, to serverless tenants. In this paper, we perform a real-world DoW attack on commodity serverless platforms to evaluate its severity. To identify such attacks, we design, implement, and evaluate Gringotts, an accurate, easy-to-use DoW detection system with a negligible performance overhead. Gringotts addresses the information ambiguity inherent in serverless functions by introducing a well-designed performance metrics collection agent. Then, Gringotts uses the Mahalanobis distance to discover anomalies in the distribution of the metrics. We implement Gringotts as a real system and conduct extensive experiments using a testbed to evaluate the performance of Gringotts. Our results indicate that Gringotts has a performance overhead of less than 1.1%, with an average detection delay of 1.86 seconds and an average accuracy of over 95.75%.
@inproceedings{Shen_2022, series={CCS ’22}, title={Gringotts: Fast and Accurate Internal Denial-of-Wallet Detection for Serverless Computing}, url={http://dx.doi.org/10.1145/3548606.3560629}, DOI={10.1145/3548606.3560629}, booktitle={Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security}, publisher={ACM}, author={Shen, Junxian and Zhang, Han and Geng, Yantao and Li, Jiawei and Wang, Jilong and Xu, Mingwei}, year={2022}, month=Nov, pages={2627–2641}, collection={CCS ’22} }ACM CoNEXT '21, 2021.
Inter-DataCenter Wide Area Network (Inter-DC WAN) that connects geographically distributed data centers is becoming one of the most critical network infrastructures. Due to limited bandwidth and inevitable link failures, it is highly challenging to guarantee network availability for services, especially those with stringent bandwidth demands, over inter-DC WAN. We present BATE, a novel Traffic Engineering (TE) framework for bandwidth availability (BA) provision, which aims to ensure that each bandwidth demand must be satisfied with a stipulated probability, when subjected to the network capacity and possible failures of the inter-DC WAN. The three core components of BATE, i.e., admission control, traffic scheduling and failure recovery, are formulated through different mathematical models and theoretically analyzed. They are also extensively compared against state-of-the-art TE schemes, using a testbed as well as real trace driven simulations across different topologies, traffic matrices and failure scenarios. Our evaluations show that, compared with the optimal admission strategy, BATE can speed up the online admission control by 30x at the expense of less than 4% false rejections. On the other hand, compared with the latest TE schemes like FFC and TEAVAR, BATE can meet the bandwidth availability targets for 23%~60% more demands under normal loads, and when network failure causes BA targets violations.
@inproceedings{Zhang_2021, series={CoNEXT '21}, title={Boosting bandwidth availability over inter-DC WAN}, url={http://dx.doi.org/10.1145/3485983.3494843}, DOI={10.1145/3485983.3494843}, booktitle={Proceedings of the 17th International Conference on emerging Networking EXperiments and Technologies}, publisher={ACM}, author={Zhang, Han and Shi, Xingang and Yin, Xia and Wang, Jilong and Wang, Zhiliang and Guo, Yingya and Lan, Tian}, year={2021}, month=Dec, pages={297–312}, collection={CoNEXT '21} }ACM/IEEE International Conference on Software Engineering (ICSE '26). To appear.
Pinpointing latency culprits inside complex microservice call graphs traditionally requires invasive code instrumentation. This work introduces a fully non-intrusive tracing technique that infers method-level delays from network-side observations and lightweight runtime hints, eliminating the need to modify application code. Evaluation on production-scale microservice deployments demonstrates fault-localization accuracy on par with intrusive baselines at a fraction of the engineering cost.
Proceedings of the ACM Web Conference (WWW '25), 2025.
The Automatic Certificate Management Environment (ACME) protocol automates the issuance and renewal of secure socket layer certificates, simplifying the management of large-scale certificate deployments. To reduce the load on Certificate Authority (CA) servers, ACME employs a caching mechanism that stores domain validation (DV) results for 30 days. However, this mechanism allows attackers to reuse previously authorized results, potentially bypassing the DV process. In this paper, we introduce the ACME Authz Cache Attack, whereby an attacker can obtain fraudulent certificates without domain control. We demonstrate that even the prominent CA, Let's Encrypt, is vulnerable to this attack. To mitigate this, we propose ACME++, an enhanced protocol that binds the client's IP address and a unique identifier to the ACME account, ensuring secure authorization for each new client and effectively preventing the ACME Authz Cache Attack. Our implementation of ACME++ shows that it introduces little overhead to the CA server.
@inproceedings{Zhang_2025, series={WWW ’25}, title={ACME++: A Secure Authorization Mechanism for ACME Clients in the Web PKI Ecosystem}, url={http://dx.doi.org/10.1145/3696410.3714763}, DOI={10.1145/3696410.3714763}, booktitle={Proceedings of the ACM on Web Conference 2025}, publisher={ACM}, author={Zhang, Tianyu and Zhang, Han and Wei, Yunze and Li, Yahui and Shi, Xingang and Wang, Jilong and Yin, Xia}, year={2025}, month=Apr, pages={1058–1067}, collection={WWW ’25} }Proceedings of the ACM Web Conference (WWW '26), pp. 7576–7587, 2026.
Looking Glass (LG) services expose a fragmented but invaluable view of the global Internet's routing fabric, yet their heterogeneous interfaces and free-form output have long resisted automated analysis. GlassMiner proposes a structure-semantics fusion approach that jointly models the DOM-level layout and natural-language hints of LG portals to extract structured measurements at scale. The system enables continent-wide BGP path observability and uncovers previously hidden routing anomalies across thousands of operator portals.
IEEE INFOCOM 2024, pp. 420–437, 2024.
Microservices are becoming more complicated, posing new challenges for traditional performance monitoring solutions. On the one hand, the rapid evolution of microservices places a significant burden on the utilization and maintenance of existing distributed tracing frameworks. On the other hand, complex infrastructure increases the probability of network performance problems and creates more blind spots on the network side. In this paper, we present DeepFlow, a network-centric distributed tracing framework for troubleshooting microservices. DeepFlow provides out-of-the-box tracing via a network-centric tracing plane and implicit context propagation. In addition, it eliminates blind spots in network infrastructure, captures network metrics in a low-cost way, and enhances correlation between different components and layers. We demonstrate analytically and empirically that DeepFlow is capable of locating microservice performance anomalies with negligible overhead. DeepFlow has already identified over 71 critical performance anomalies for more than 26 companies and has been utilized by hundreds of individual developers. Our production evaluations demonstrate that DeepFlow is able to save users hours of instrumentation efforts and reduce troubleshooting time from several hours to just a few minutes.
@inproceedings{Shen_2023, series={ACM SIGCOMM '23}, title={Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero Code}, url={http://dx.doi.org/10.1145/3603269.3604823}, DOI={10.1145/3603269.3604823}, booktitle={Proceedings of the ACM SIGCOMM 2023 Conference}, publisher={ACM}, author={Shen, Junxian and Zhang, Han and Xiang, Yang and Shi, Xingang and Li, Xinrui and Shen, Yunxi and Zhang, Zijian and Wu, Yongxiang and Yin, Xia and Wang, Jilong and Xu, Mingwei and Li, Yahui and Yin, Jiping and Song, Jianchang and Li, Zhuofeng and Nie, Runjie}, year={2023}, month=Sept, pages={420–437}, collection={ACM SIGCOMM '23} }IEEE International Conference on Network Protocols (ICNP '25), pp. 1–12, 2025.
Existing latency-diagnosis tools for SRv6 either flood the data plane with probes or return probabilistic verdicts. SRmesh embeds deterministic in-band telemetry into segment-routing headers, enabling per-link latency attribution with bounded measurement overhead. It has been validated in a multi-vendor SRv6 testbed and identifies bottleneck links in seconds with 100% recall.
IEEE INFOCOM 2026.
IEEE INFOCOM 2025, pp. 420–437.
ACM CoNEXT 2025.
Overlay/underlay architecture is increasingly prevalent in modern networks but also introduces greater complexity and error-proneness. Existing control plane verifiers for underlay networks face scalability and interpretability challenges when extended to overlay/underlay networks due to methodological limitations. This paper presents MEV, the first control plane verifier designed for overlay/underlay networks. MEV introduces ensemble verification, a novel verification paradigm enabling independent reasoning for behaviors within each routing instance or protocol and forwarding behaviors within each virtual network. We implement and deploy MEV on a real nationwide overlay/underlay network, FITI, where it successfully identifies 22 misconfigurations that could cause isolation and reachability issues. Furthermore, we evaluate MEV against Batfish+ on FITI and a range of synthetic networks with diverse scales, and find that it achieves up to a 102× speedup over Batfish+.
@article{Liu_2025, title={Scalable and Interpretable Overlay Network Checking via Ensemble Verification}, volume={3}, ISSN={2834-5509}, url={http://dx.doi.org/10.1145/3768974}, DOI={10.1145/3768974}, number={CoNEXT4}, journal={Proceedings of the ACM on Networking}, publisher={Association for Computing Machinery (ACM)}, author={Liu, XinZhe and Li, Yahui and Zhang, Han and Yin, Xia and Shi, Xingang and Wang, Zhiliang and Ren, Gang and Wang, Jilong and Yao, Jiangyuan}, year={2025}, month=Nov, pages={1–25} }Network and Distributed System Security Symposium (NDSS '22), 2022.
IEEE/ACM Transactions on Networking, 2026.
Serverless computing, or Function-as-a-Service, continues to gain popularity due to its pay-as-you-go billing model, flexibility, and cost efficiency. However, these same features introduce significant security risks, such as the Denial-of-Wallet (DoW) attack. In this paper, we conduct real-world DoW attacks on commercial serverless platforms to evaluate their severity. To detect such attacks, we design, implement, and evaluate Clover, an accurate and user-friendly DoW detection system with negligible performance overhead. Clover addresses information ambiguity in serverless environments by deploying a request-oriented metric collection agent. At its core, Clover proposes a workload verification approach to bridge performance metrics and execution duration. Specifically, Clover uses a multivariate linear model to learn the benign relationship between metrics and execution duration, effectively characterizing normal workload behavior. It then continuously monitors runtime workloads by calculating their Mahalanobis distance from this learned benign model. Deviations identified through this distance indicate potential DoW attacks. Implemented as a practical system, Clover introduces performance overhead of less than 3.2%, maintains an average model execution time of only 0.84 microseconds, and achieves an accuracy of 92.7% under the most challenging scenario.
@article{Shen_2026, title={Clover: Workload Verification for Real-Time Detection of Contention-Induced Slowdowns in Serverless Platforms}, volume={34}, ISSN={2998-4157}, url={http://dx.doi.org/10.1109/TON.2026.3680829}, DOI={10.1109/ton.2026.3680829}, journal={IEEE Transactions on Networking}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Shen, Junxian and Zhang, Han and Lin, Weiwei and Geng, Yantao and Li, Jiawei and Wang, Jilong and Xu, Mingwei}, year={2026}, pages={4700–4715} }IEEE/ACM Transactions on Networking, 2026.
IEEE Transactions on Networking, 2025.
IEEE Transactions on Networking, 2025.
IEEE/ACM Transactions on Networking, 2022.
IEEE Transactions on Parallel and Distributed Systems, vol. 34, no. 8, pp. 2448–2463, 2023.
IEEE Transactions on Parallel and Distributed Systems, vol. 34, no. 4, pp. 1362–1375, 2023.
IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 10, pp. 2571–2583, 2022.
IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 5, pp. 1069–1083, 2022.
IEEE Transactions on Information Forensics and Security, 2025.
IEEE Transactions on Information Forensics and Security, vol. 19, pp. 10099–10113, 2024.
IEEE Transactions on Information Forensics and Security, vol. 18, pp. 2076–2090, 2023.
IEEE Transactions on Information Forensics and Security, vol. 18, pp. 1610–1624, 2023.
IEEE Transactions on Information Forensics and Security, vol. 17, pp. 3972–3987, 2022.
IEEE Transactions on Information Forensics and Security, vol. 16, pp. 2300–2311, 2021.
IEEE International Conference on Network Protocols (ICNP '22), 2022.
ACM SIGSAC Conference on Computer and Communications Security (CCS '21), 2021.
Unsupervised Deep Learning (DL) techniques have been widely used in various security-related anomaly detection applications, owing to the great promise of being able to detect unforeseen threats and superior performance provided by Deep Neural Networks (DNN). However, the lack of interpretability creates key barriers to the adoption of DL models in practice. Unfortunately, existing interpretation approaches are proposed for supervised learning models and/or non-security domains, which are unadaptable for unsupervised DL models and fail to satisfy special requirements in security domains. In this paper, we propose DeepAID, a general framework aiming to (1) interpret DL-based anomaly detection systems in security domains, and (2) improve the practicality of these systems based on the interpretations. We first propose a novel interpretation method for unsupervised DNNs by formulating and solving well-designed optimization problems with special constraints for security domains. Then, we provide several applications based on our Interpreter as well as a model-based extension Distiller to improve security systems by solving domain-specific problems. We apply DeepAID over three types of security-related anomaly detection systems and extensively evaluate our Interpreter with representative prior works. Experimental results show that DeepAID can provide high-quality interpretations for unsupervised DL models while meeting the special requirements of security domains. We also provide several use cases to show that DeepAID can help security operators to understand model decisions, diagnose system mistakes, give feedback to models, and reduce false positives.
@inproceedings{Han_2021, series={CCS ’21}, title={DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security Applications}, url={http://dx.doi.org/10.1145/3460120.3484589}, DOI={10.1145/3460120.3484589}, booktitle={Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security}, publisher={ACM}, author={Han, Dongqi and Wang, Zhiliang and Chen, Wenqi and Zhong, Ying and Wang, Su and Zhang, Han and Yang, Jiahai and Shi, Xingang and Yin, Xia}, year={2021}, month=Nov, pages={3197–3217}, collection={CCS ’21} }