II Intrusion detection systems (IDS)
Cybersecurity aims to protect computer systems, networks, and data from unauthorized access, misuse, disruption, and damage. As modern services increasingly depend on interconnected infrastructures, detecting malicious activities before they cause serious impact has become a core requirement for network defense. Intrusion Detection Systems (IDSs) address this requirement by monitoring system events or network traffic and identifying behaviors that may indicate security violations. The concept of intrusion detection was first introduced by James Anderson in the early 1980s [1]. Building on this foundation, Denning proposed one of the earliest functional IDS models, which formalized intrusion detection as the process of monitoring audit records and identifying abnormal or suspicious system behavior [2]. In general, an IDS can be implemented as a software- or hardware-based security mechanism. Its primary goal is to detect potential attacks, policy violations, or abnormal activities and provide alerts that support timely investigation and response.
Over the past few decades, IDS research has evolved from rule-based and statistical methods to machine learning and deep learning approaches. Existing IDS studies differ not only in the detection algorithms they use, but also in the attack assumptions, deployment environments, data sources, and explanation mechanisms they consider. Therefore, before reviewing generative AI and FL-based IDS, it is necessary to summarize the main IDS research directions that provide the technical context for this survey. In this section, we review representative IDS-related studies from five perspectives. Section II-A discusses adversarial machine learning for IDS, focusing on the vulnerability of ML-based detectors to adversarial manipulation. Section II-B reviews anomaly-based IDS, where attacks are detected by modeling deviations from normal behavior. Section II-C summarizes IDS studies for IoT environments, where resource constraints and heterogeneous devices introduce additional challenges. Section II-D reviews explainable IDS, which aims to make detection results more interpretable for security analysts. Finally, Section II-E discusses commonly used IDS datasets, which are fundamental for training, evaluating, and comparing IDS models.
II-A Adversarial Machine Learning for IDS
Machine learning (ML) and deep learning (DL) models have been widely adopted in IDS because they can learn discriminative patterns from network traffic and detect attacks beyond manually defined signatures. However, the use of ML also introduces a new attack surface. An adversary may intentionally perturb input traffic features or manipulate training data to mislead the detector, causing malicious traffic to be classified as benign or forcing benign traffic to be reported as malicious. These threats are generally studied under adversarial machine learning, which focuses on understanding the vulnerability of ML models and improving their robustness against adversarial manipulation.
For IDS, adversarial robustness is especially important because network attackers can actively adapt their behaviors after observing or probing the detection system. Even small modifications to traffic characteristics may change the prediction of an ML-based IDS while preserving the attack objective. Therefore, researchers have investigated both adversarial attacks against IDS models and defense strategies that make IDS models more stable under adversarial conditions.
Jmila and Khedher [4] evaluated the vulnerability of IDS models built with shallow ML classifiers, including Decision Tree, Random Forest, and Logistic Regression. Their experiments on NSL-KDD [5] and UNSW-NB15 [6] show that different adversarial attacks affect classifiers differently, indicating that IDS robustness depends on both the attack strategy and the underlying model architecture. Alotaibi and Rassam [7] further surveyed adversarial attacks against ML-based IDS and summarized representative defense strategies. These studies show that adversarial machine learning is an important research direction for IDS, particularly as IDS models become more data-driven and are deployed in adaptive threat environments.
II-B Anomaly-based IDS
IDSs are commonly categorized as signature-based, anomaly-based, or hybrid systems according to their detection strategy. A signature-based IDS detects intrusions by matching observed activities with predefined attack signatures or rules. This approach is effective for known attacks, but it is limited when facing new or evolving threats that do not match existing signatures. In contrast, an anomaly-based IDS first models normal system or network behavior and then identifies activities that deviate from this learned baseline. This capability makes anomaly-based IDS particularly useful for detecting unknown or zero-day attacks, although it may also introduce higher false-positive rates when benign behavior changes over time.
Several surveys have reviewed anomaly-based IDS from different perspectives. Yang et al. [8] conducted a systematic review of network IDS studies and analyzed commonly used data processing techniques, evaluation metrics, benchmark datasets, and detection models. Hajj et al. [9] presented a taxonomy of network attacks and discussed attack tools, relevant detection features, IDS data sources, dataset types, system architectures, and detection modes. They also summarized key challenges that affect the effectiveness of anomaly-based IDS, including dataset quality, feature selection, evaluation consistency, and deployment constraints.
These studies show that anomaly-based IDS is a central direction in IDS research because it directly addresses the limitation of signature-based detection under unknown attacks. However, its performance depends heavily on how normal behavior is modeled, how representative the training data are, and how deviations are distinguished from benign traffic variations.
II-C IDS for IoT
The Internet of Things (IoT) connects sensors, actuators, embedded devices, and coordinator nodes to provide networked services in domains such as smart homes, healthcare, transportation, and industrial systems. Compared with conventional networks, IoT environments are more heterogeneous because devices may differ in hardware capability, operating system, communication protocol, and deployment context. Many IoT devices also have limited computation, memory, and energy resources, which makes it difficult to deploy complex security mechanisms directly on the devices. In addition, vulnerabilities in firmware, hardware, and protocol implementations can expose IoT systems to malware propagation, denial-of-service attacks, spoofing, and unauthorized access. These characteristics make IDS design for IoT different from traditional IDS design.
Several studies have reviewed IDS techniques specifically for IoT environments. Kumar et al. [10] provided a taxonomy of ML-based IDS for secure IoT communication and compared different categories according to their advantages, limitations, and resource-related evaluation metrics. They also proposed an IDS model that combines convolutional neural networks (CNNs) with fuzzy rules to improve performance under IoT constraints, including energy consumption and packet delivery ratio. Jayalaxmi et al. [11] analyzed ML- and DL-based IDS and Intrusion Prevention Systems (IPS) for IoT. They further proposed a risk factor analyzer and a hybrid Intrusion Detection and Prevention System (IDPS) framework to address limitations of purely anomaly-based or signature-based methods.
These studies show that IoT-oriented IDS must consider both detection accuracy and deployment constraints. Effective IDS for IoT should be able to detect diverse attacks while remaining lightweight, adaptive to heterogeneous devices, and practical for resource-constrained environments.
II-D Explainable IDS (X-IDS)
Machine learning (ML) and deep learning (DL) models have been widely used in IDS because they can learn complex traffic patterns and improve attack detection performance. However, many DL-based IDS models operate as black-box systems, meaning that their internal decision process is difficult for human analysts to interpret. This lack of transparency limits the ability of administrators and security experts to understand why an alert is generated, identify the root cause of an attack, and determine an appropriate response. Explainable Artificial Intelligence (XAI) addresses this issue by providing interpretable evidence or explanations for model predictions. In the IDS context, explainability is important not only for improving user trust, but also for supporting incident analysis, model debugging, and security decision-making.
Several surveys have studied explainable IDS from different perspectives. Moustafa et al. [12] presented a comprehensive survey on XAI methods for cyber defense, with a particular focus on anomaly-based IDS in IoT networks. Their work reviewed studies at the intersection of XAI, anomaly-based intrusion detection, IoT security, summarized open challenges, and future directions for explainable cyber-defense systems. Neupane et al. [13] proposed a taxonomy of X-IDS techniques and categorized existing methods into white-box and black-box approaches. They also introduced a three-layered X-IDS architecture inspired by the DARPA XAI program [14] and discussed key challenges in developing explainable IDS models.
These studies indicate that X-IDS is an important direction for practical IDS deployment. A detection model with high accuracy may still be difficult to use in real security operations if its alerts cannot be interpreted. Therefore, explainable IDS aims to bridge the gap between automated detection and human-centered cyber-defense analysis.
II-E Datasets for IDS
| Dataset Name | Year | Training instances | Testing instances | # of classes | # of Features | FL support |
|---|---|---|---|---|---|---|
| KDD CUP 99 [15] | 1999 | 1,074,992 | 311,029 | 5 | 41 | No |
| ISCX NSL-KDD [5] | 2009 | 4,898,431 | 311,027 | 4 | 42 | No |
| UNSW-NB15 [6] | 2015 | 447,915 | - | 10 | 49 | No |
| CICIDS2017 [16] | 2017 | 2,830,743 | - | 8 | 84 | No |
| CSE-CIC-IDS2018 [16] | 2018 | 16,136,255 | - | 12 | 76 | No |
| CICDDoS2019 [17] | 2019 | 50,063,112 | - | 11 | 84 | No |
| Car Hacking [18] | 2020 | 8,694,507 | - | 5 | 12 | No |
| FLNET2023 [19] | 2023 | 6,807,107 | - | 11 | 84 | Yes |
| CICIoT2023 [20] | 2023 | 46,556,613 | - | 25 | 39 | No |
| CIC IoV [21] | 2024 | 1,408,219 | - | 6 | 12 | No |
| X-CANIDS [22] | 2024 | 286,293 | - | 5 | 689 | No |
TABLE II: Representative IDS datasets. Datasets in bold are widely used benchmark datasets in recent IDS research.
Benchmark datasets are essential for developing, evaluating, and comparing IDS models because they provide shared traffic records, attack labels, and feature representations for reproducible experiments [19, 17, 16, 6]. In IDS research, dataset quality directly affects the reliability of model evaluation. A dataset should contain representative benign and malicious traffic, diverse attack types, clear labeling rules, and sufficient feature information for downstream detection tasks. However, constructing realistic IDS datasets is difficult because real network traffic may contain sensitive user information, organization-specific configurations, and security-critical infrastructure details.
Several studies have reviewed the limitations and design considerations of IDS datasets. Ring et al. [23] analyzed network-based IDS datasets from the perspective of attack scenarios and discussed the relationships among different datasets. Their work also provided recommendations for dataset selection and future dataset construction. Khraisat et al. [24] reviewed IDS datasets together with detection techniques, data collection methods, evaluation practices, and dataset limitations. These studies show that dataset selection is not only an experimental detail but also a key factor that determines whether IDS results can generalize to realistic deployment environments.
Table II summarizes commonly used IDS datasets and recent domain-specific datasets. The listed datasets cover traditional network intrusion detection, distributed denial-of-service detection, IoT traffic, in-vehicle network security, and FL-oriented IDS evaluation. The bold datasets, including UNSW-NB15, CICIDS2017, CICDDoS2019, and CICIoT2023, are emphasized because they are widely used in recent IDS studies and provide relatively large-scale traffic records with multiple attack categories. In contrast, datasets such as Car Hacking, CIC IoV, and X-CANIDS are more domain-specific and are mainly designed for vehicle or controller area network security. FLNET2023 is particularly relevant to FL-based IDS because it explicitly supports federated evaluation, while most existing IDS datasets are centrally collected and do not naturally reflect client-level data distribution.
Although these datasets have supported substantial IDS research, several limitations remain. Many benchmark datasets are collected in controlled environments and may not fully capture real network topology, user behavior, temporal dynamics, or evolving attack strategies. In addition, most datasets are designed for centralized learning and provide limited support for studying non-independent and identically distributed (non-IID) clients in FL-based IDS. These limitations motivate the development of more realistic, privacy-aware, and federated IDS benchmarks.