OER·harvester

← Back to the library
arXiv HTML resource

Generative AI and Federated Learning for Intrusion Detection Systems: A Survey

Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Internet of Things (IoT), enterprise, and distributed network environments. However, developing reliable IDS models remains challenging because attack behaviors evolve over time, realistic datasets are difficult to obtain, traffic records may be incomplete, attack classes are…

Licence
OPEN CC-BY-4.0
Authors
Jiefei Liu, Abu Saleh Md Tayeen, Pratyay Kumar, Qixu Gong, Wenbin Jia…
Published
2026-07-01 · arXiv
Language
en
Length
20383 words
Type
narrative text

Cites 94 works

inferred
Open original ↗

III Generative models

{forest}

Fig. 1: Generative AI with intrusion detection system

Generative AI refers to a class of models that learn the underlying structure or probability distribution of observed data and use the learned representation to generate new samples. In the IDS domain, generative models are mainly used to address data-related limitations, including limited attack samples, class imbalance, missing values, and insufficient traffic diversity. They are also used for anomaly detection, adversarial traffic generation, and explanation of IDS outputs. As shown in Figure 1, this survey organizes generative AI techniques for IDS into four representative model families: Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), diffusion models, and Large Language Models (LLMs).

Section III-A first reviews VAE-based IDS studies. VAEs encode input data into a latent distribution and reconstruct or generate samples from this latent space, making them useful for anomaly detection, data generation, and data augmentation. Section III-B then discusses GAN-based IDS methods. GANs use a generator and a discriminator in an adversarial training process and have been widely applied to synthetic traffic generation, class-imbalance mitigation, and adversarial attack generation. Section III-C reviews diffusion-based methods, which generate data through a denoising process and have recently been extended from image and text generation to tabular data generation. Section III-D discusses LLM-based IDS applications, including traffic-log analysis, tabular data generation, attack classification, and IDS alert explanation.

Network traffic data used in IDS can appear in different formats, including flow-level tabular features, packet-level byte sequences, traffic logs, and raw packet capture (PCAP) files. Therefore, different generative model families can be applied depending on how the traffic is represented. Tabular generative models are suitable for flow-based IDS datasets, sequence and language models can be used for logs or byte-level representations, and image-based generative models can be applied when traffic records are transformed into image-like formats. The following subsections summarize how these generative models have been used in IDS and what technical challenges remain for each model family.

III-A Variational Autoencoder in IDS

Variational Autoencoders (VAEs) are generative neural networks introduced by Kingma and Welling [25]. A VAE encodes input data into a latent space and reconstructs the input from the learned latent representation. Different from a conventional autoencoder, a VAE models the latent representation as a probability distribution rather than a fixed vector. This probabilistic design allows the model to sample from the latent space and generate new data points that follow the learned data distribution. Therefore, VAEs are suitable for both representation learning and synthetic data generation.

In IDS research, VAEs are mainly used for anomaly detection, data generation, and data augmentation. For anomaly detection, a VAE learns the distribution of normal or known traffic and identifies samples that are difficult to reconstruct. For data generation and augmentation, a VAE can generate synthetic network traffic samples to improve data diversity, mitigate class imbalance, or support downstream IDS training. This subsection reviews VAE-based IDS studies according to these major application objectives.

III-A1 Anomaly Detection

{forest}

Fig. 2: Variational autoencoder anomaly detection in intrusion detection system

VAE-based anomaly detection relies on the assumption that normal traffic can be reconstructed more accurately than abnormal or unseen traffic. After training, the VAE reconstructs an input sample from its latent representation, and the difference between the original input and the reconstructed output is measured as the reconstruction error. A large reconstruction error indicates that the input does not follow the learned traffic distribution and may correspond to malicious or abnormal behavior. As shown in Figure 2, this mechanism has been widely used in anomaly-based IDS.

Several studies have applied VAEs to network-flow anomaly detection. Zavrak and Iskefiyeli [26] used VAE-based reconstruction error to detect anomalies from network flow features. Xu et al. [27] proposed a Log-Cosh Conditional VAE, where the log-cosh loss improves robustness to outliers during reconstruction. These studies show that reconstruction-based VAE models can provide an effective unsupervised or semi-supervised mechanism for detecting abnormal traffic patterns.

VAEs have also been extended to specific IDS scenarios. In IoT environments, Khanam et al. [28] incorporated focal loss into a VAE-based model to improve the detection of abnormal traffic under imbalanced conditions. Nguyen et al. [29] proposed the Gradient-based Explainable Variational Autoencoder (GEE), which combines anomaly detection with gradient-based explanations to help interpret detected anomalies. For insider threat detection, Pantelidis et al. [30] used autoencoder-based models, including VAEs, to identify abnormal user behavior within an organization. In critical infrastructure, VAE-based methods have been applied to power systems and Advanced Metering Infrastructure (AMI) networks to detect data corruption and stealth cyber-attacks [31, 32].

Evaluation studies further show that VAE-based anomaly detection can be improved through model design and loss-function selection. Conditional VAEs have been evaluated for network intrusion detection and shown to be effective when the model is designed to capture class- or condition-specific traffic patterns [33, 34]. Other studies have explored latent-space clustering and inferential autoencoder designs to improve anomaly separation and adapt to changing network conditions [35, 36]. Overall, VAE-based anomaly detection is useful for IDS because it can model normal traffic without requiring extensive labeled attack samples, but its performance depends on the quality of the learned latent representation and the threshold used to distinguish normal and abnormal reconstruction behavior.

III-A2 Data Generation

{forest}

Fig. 3: Variational autoencoder in intrusion detection system

VAEs can be used to generate synthetic IDS data by sampling from the learned latent distribution and decoding the sampled representations into new traffic records. This capability is useful when real network traffic is limited, sensitive, or imbalanced. In IDS, VAE-based data generation is mainly used to enrich training datasets, simulate network traffic, and generate samples for underrepresented attack classes. Ren et al. [37] and Lin et al. [38] showed that VAE-generated data can improve IDS training by increasing data diversity. Martin et al. [39, 40] further developed VAE-based generative models for intrusion detection and demonstrated their ability to produce synthetic traffic samples for downstream detection tasks.

Another application of VAE-based data generation is traffic simulation. In this setting, the generated samples are used to represent different network behaviors, including benign activities and malicious attacks, under controlled experimental conditions. Dinh et al. [41] and Yang et al. [42] used VAE-based models to support IDS training by generating or reconstructing traffic patterns that improve the representation of network behaviors. This is useful for evaluating IDS models when collecting sufficient real-world attack traffic is difficult.

VAEs have also been used to address class imbalance by generating synthetic samples for minority attack classes. Class imbalance is common in IDS datasets because benign traffic and frequent attack types often dominate the data, while rare attacks have limited training samples. Chuang and Huang [43] and Yang et al. [44] showed that VAE-based balancing strategies can improve IDS performance on underrepresented classes. Evaluation studies also indicate that the usefulness of generated data should be measured through downstream IDS performance, such as detection accuracy, recall, and robustness, rather than only by visual or statistical similarity [45, 46].

III-A3 Data Augmentation

{forest}

Fig. 4: Variational autoencoder in intrusion detection system

VAE-based data augmentation aims to expand IDS training data by generating additional samples or feature variations from the learned latent distribution. This is different from general data generation because the main objective is not only to create synthetic traffic, but also to improve downstream IDS training. In particular, VAE-based augmentation is useful when the original dataset is limited, imbalanced, or insufficiently diverse.

One common strategy is feature-level augmentation, where a VAE learns latent representations of traffic records and generates variations that enrich the original feature space. This strategy can help the IDS model observe broader traffic patterns during training and reduce overfitting to limited samples. For example, Sabeel et al. [47] applied VAE-based augmentation to improve the detection of atypical attack flows, where rare or unusual attack behaviors are difficult to learn from the original data alone.

Another important strategy is oversampling, where VAEs generate synthetic samples for underrepresented classes. As noted above, when minority attack classes have too few samples, the detector may become biased toward majority classes and fail to recognize rare attacks. VAE-based oversampling mitigates this issue by increasing the number of minority-class samples while preserving the learned data distribution [48, 42].

Several studies have evaluated the impact of VAE-based augmentation on IDS performance. Compared with traditional augmentation or resampling methods, VAE-based approaches can generate samples that better reflect the underlying feature relationships in network traffic [49, 50]. Empirical results also show that VAE-augmented data can improve detection accuracy, robustness, and class-level performance, especially under imbalanced or limited-data settings [51, 52].

VAE-based augmentation has also been explored for improving IDS robustness. By exposing the detector to reconstructed, perturbed, or boundary-related traffic samples, the augmented training data can help the model become less sensitive to small variations in traffic patterns. For example, VAE-based methods have been used in two-stage detection and known/unknown intrusion detection settings to improve the model’s ability to handle abnormal or unseen traffic [53, 54]. Overall, VAE-based data augmentation supports IDS by improving training data diversity, but its effectiveness still depends on the quality of generated samples and their consistency with realistic network behavior.

III-B GAN in IDS

{forest}

Fig. 5: Generative Adversarial Nets (GAN).

Generative Adversarial Networks (GANs) were introduced by Goodfellow et al. [55]. A GAN contains two neural networks: a generator and a discriminator. The generator learns to produce synthetic samples, while the discriminator learns to distinguish generated samples from real samples. Through this adversarial training process, the generator gradually improves its ability to produce data that resemble the original distribution. Later studies improved the stability and quality of GAN training, including Wasserstein GAN and improved Wasserstein GAN [56, 57]. As summarized in Figure 5, GANs have been applied to IDS mainly for tabular data generation, data augmentation, data imputation, network traffic generation, and adversarial traffic generation.

III-B1 Tabular GAN

Many IDS datasets are represented as tabular data, where each row corresponds to a traffic record and each column corresponds to a feature such as duration, packet count, byte count, protocol type, or flow statistic. Therefore, tabular GAN models are relevant to IDS because they can generate synthetic flow-level records for training or evaluation. Park et al. [58] introduced Table-GAN for synthetic table generation, aiming to reduce the risk of exposing real data during data sharing. Xu et al. [59] proposed CTGAN and TVAE to model complex tabular distributions, including mixed continuous and discrete features. These tabular generation methods provide the methodological basis for applying GANs to flow-based IDS datasets.

Conditional GANs have also been used to generate class-specific IDS samples. Li et al. [60] proposed a BERT-enhanced Conditional GAN for multi-class intrusion detection. In this framework, the CGAN generates additional samples for minority attack classes, while BERT is embedded in the discriminator to strengthen the dependency between input features and output labels. The method addresses class imbalance and improves multi-class IDS performance on several datasets, including CSE-CIC-IDS2018, NF-ToN-IoT-V2, and NF-UNSW-NB15-v2.

III-B2 IDS Data Generation

In IDS, GAN-based data generation is mainly used for three purposes. The first purpose is data augmentation, where GANs generate additional samples to enrich the training set and improve downstream detection performance. This is useful when attack samples are limited or when minority attack classes are underrepresented. Park et al. [61] and Huang and Lei [62] applied GAN-based augmentation to address class imbalance in IDS and improve detection performance. Related surveys also show that GANs have been widely studied for imbalance learning and synthetic data generation in broader machine learning settings [63, 64].

The second purpose is data imputation, where GANs estimate missing or incomplete feature values while preserving the structure of the original data. Missing values can occur because of packet loss, incomplete collection, preprocessing errors, or unavailable traffic attributes. GAN-based imputation methods have been reviewed in [65, 66], and they are relevant to IDS because incomplete traffic records can reduce the reliability of model training and evaluation.

The third purpose is adversarial traffic generation. In this setting, GANs generate malicious traffic records that preserve attack functionality while appearing similar to benign or normal traffic. Such generated samples can be used offensively to bypass IDS models or defensively to evaluate and improve IDS robustness. De Araujo-Filho et al. [67], Aldhaheri and Alhuzali [68], and Shu et al. [69] studied GAN-based adversarial generation against IDS. In addition, Zhao et al. [70] investigated GAN-based network traffic generation for improving IDS performance. These studies show that GANs can support IDS model training, but they also introduce security concerns because generated traffic can be used to test or evade detection systems.

III-C Diffusion in IDS

{forest}

Fig. 6: Generative diffusion on surveys.

Diffusion models are generative models that learn to generate data through a gradual adding noising and denoising process. Sohl-Dickstein et al. [71] introduced the diffusion-based generative framework in 2015, where the forward process progressively adds noise to data and the reverse process learns to recover clean samples from noisy inputs. Later, Denoising Diffusion Probabilistic Models (DDPMs) formalized this process as an effective deep generative modeling approach [72]. Although diffusion models were first widely studied in image generation, they have also been extended to text and tabular data generation [73, 74, 75]. Dhariwal and Nichol [76] further showed that diffusion models can achieve strong image generation performance compared with GAN-based methods.

For IDS, diffusion models are relevant because network traffic can be represented in several forms, including tabular flow features, packet sequences, logs, and transformed image-like representations [77, 78]. Therefore, diffusion models can be used for different IDS-related tasks, including adversarial attack generation, adversarial purification, adversarial training, and synthetic data generation. The following subsections summarize these applications and highlight their relevance to IDS.

III-C1 Adversarial Attacks

Diffusion models have been studied in cybersecurity partly because they can generate high-quality synthetic samples that resemble real data. Existing surveys and studies have discussed the role of generative models, including diffusion-related methods, in creating adversarial examples and evaluating security vulnerabilities [79, 80, 81, 82]. In adversarial attack settings, the generated sample is designed to remain close to a valid input while causing a target ML model to make an incorrect prediction. This idea is relevant to IDS because attackers may modify traffic features while preserving malicious behavior, making the attack harder to detect.

Some diffusion-based adversarial methods combine the denoising process with gradient-based perturbation strategies [83, 84, 85], such as Projected Gradient Descent (PGD). The purpose is to generate samples that appear realistic but still mislead the target model. Although much of this work has been developed in image domains, the same principle is important for IDS because generated or perturbed traffic can be used to test whether a detector is robust to adaptive attacks.

In response to diffusion-generated fake or adversarial samples, Hooda et al. [86] proposed Disjoint Diffusion Deepfake Detection (D4). D4 is designed to detect fake images generated by diffusion models and to generalize to unseen data distributions and generative techniques. While this work is not specific to IDS, it reflects a broader security problem: as generative models become stronger, detection systems must also be evaluated against generated and previously unseen samples.

III-C2 Adversarial Purification

{forest}

Fig. 7: Generative diffusion on adversarial purification.

Adversarial purification (AP) is a defense strategy that uses a generative model to remove or reduce adversarial perturbations before classification. The main idea is to map a potentially perturbed input back toward the clean data distribution so that the downstream classifier receives a less corrupted sample. Early work such as PixelDefend used generative modeling for purification [87]. More recent studies have applied diffusion models to adversarial purification because the denoising process can naturally remove small perturbations from input data [88, 89, 90].

Diffusion-based AP has been evaluated in several domains, including image, 3D point cloud, text, and audio data [91, 92, 93, 94, 95, 96, 97]. Carlini et al. [98] further studied the robustness guarantees of diffusion-based purification, while other studies examined its limitations and evaluation reliability [99, 100]. For IDS, adversarial purification is a potential direction because network traffic may be intentionally perturbed to evade detection. However, applying AP to IDS requires preserving protocol validity and attack semantics, not only removing statistical noise.

III-C3 Adversarial Training

{forest}

Fig. 8: Generative diffusion on adversarial training.

Adversarial training (AT) improves model robustness by training the model with adversarially perturbed examples. Goodfellow et al. [101] introduced adversarial examples and demonstrated that including such examples during training can improve resistance to adversarial attacks. Later studies showed that robust generalization often requires more training data, because the model must learn stable decision boundaries under both clean and adversarial conditions [102, 103]. This observation motivated the use of external data and generative models to expand training sets for robust learning [104, 105, 106].

Diffusion models can support adversarial training by generating additional training samples that improve data diversity. Wang et al. [107] used an Elucidating Diffusion Model (EDM) to generate high-quality synthetic image data for adversarial training and showed that generated data can improve robustness without relying only on external real data. Yu et al. [108] proposed the Adversarial Denoising Diffusion Model (ADDM) for unsupervised anomaly detection and showed that it can improve performance compared with DDPM-based anomaly detection methods under reduced sample settings. These studies suggest that diffusion-generated data may be useful for improving IDS robustness, especially when real adversarial or rare attack samples are limited.

{forest}

Fig. 9: Generative diffusion on tabular data generation.

III-C4 Data Generation

Data generation is the most direct application of diffusion models to IDS [109]. Many IDS datasets are represented as tabular flow-level features, where each record describes a network flow using statistical attributes such as packet counts, byte counts, duration, and protocol information. Therefore, tabular diffusion models are particularly relevant to IDS. TabDDPM [110] extended diffusion modeling to tabular data and provided a basis for generating structured records with both numerical and categorical features.

Recent studies have applied diffusion models to intrusion detection and related security tasks [111, 112, 113]. In these settings, diffusion-generated samples can be used to increase data diversity, supplement limited attack samples, or mitigate class imbalance. Wang et al. [112] further showed that diffusion-based generation can help address data imbalance by generating additional samples for underrepresented categories. Compared with GAN-based generation, diffusion models may provide more stable training [109], but their usefulness for IDS still depends on whether the generated traffic preserves realistic feature relationships, protocol constraints, and attack semantics.

III-D LLM in IDS

{forest}

Fig. 10: Large Language Model (LLM).

Large Language Models (LLMs) are transformer-based models trained to process and generate sequential data. In IDS research, LLMs and related transformer-based language models can be applied when network traffic is represented as logs, packet-byte sequences, flow records converted into textual formats, or structured tabular records. Compared with conventional ML models, LLMs provide two potential advantages for IDS: they can model contextual relationships in sequential traffic representations, and they can generate natural-language explanations for security analysts. However, their use in IDS also introduces practical challenges, including high computation cost, data formatting sensitivity, limited interpretability, and possible hallucination when explanations are generated without sufficient grounding.

III-D1 IDS Applications

LLM-based and transformer-based IDS studies mainly focus on using language-model architectures to classify network activities. Lira et al. [114] proposed BERTIDS, a BERT-based model for network intrusion detection. In this method, network logs are converted into tokenized sequences that can be processed by BERT. The model is then fine-tuned to distinguish normal traffic from different attack categories. Their experiments on NSL-KDD reported an accuracy of 98.01% and a false positive rate of 1.48%, showing that transformer-based language models can be adapted to IDS classification tasks.

Manocchio et al. [115] presented FlowTransformer, a modular framework for transformer-based Network Intrusion Detection Systems (NIDSs). The framework allows different components, including input encoding, transformer architecture, classification head, and evaluation dataset, to be replaced and compared. Their evaluation across public flow-based NIDS datasets showed that the classification head has a substantial effect on detection performance. This result indicates that applying transformers to IDS is not only a matter of selecting a large model; the representation of traffic features and the design of the output classifier are also important.

Although these studies show the potential of transformer-based models for IDS, several limitations remain. First, fine-tuning large models requires substantial computation and memory resources, which may be impractical for resource-constrained security environments. Second, IDS data are often numerical or mixed-type tabular records, while LLMs are originally designed for textual sequences. Therefore, the performance of LLM-based IDS depends strongly on how traffic records are encoded into model-readable inputs. Third, model complexity can make it difficult to interpret why a specific traffic record is classified as malicious, which limits direct use in high-stakes security operations.

III-D2 LLM for Explainable IDS

Another important use of LLMs in IDS is explanation generation. IDS alerts often contain technical information such as attack labels, traffic features, source and destination attributes, and model confidence scores. These alerts may be difficult for non-expert users to interpret. LLMs can be used as an explanation layer that converts IDS outputs into natural-language descriptions, summarizes possible causes, and suggests response actions. In this setting, the LLM does not necessarily replace the detector; instead, it helps users understand and act on detection results.

Rjoub et al. [116] discussed the role of explainable AI in cybersecurity and highlighted the need for human-understandable explanations in security systems. Juttner et al. [117] proposed ChatIDS, which uses ChatGPT to explain IDS alerts and provide suggestions to non-expert users. This type of approach can improve the usability of IDS outputs, especially when alerts must be interpreted by operators who do not have deep knowledge of network security or ML models.

LLMs have also been combined with traditional ML classifiers and XAI tools. Ali et al. [118] introduced HuntGPT, an intrusion detection dashboard that integrates a Random Forest classifier, XAI methods such as SHAP and LIME, and GPT-3.5 Turbo. The classifier detects anomalies, the XAI methods identify important features behind the prediction, and the LLM presents the result in a more understandable form for analysts. Other studies have also explored explainable AI for attack classification and cybersecurity analysis [119, 120]. These works show that LLMs are useful for improving the communication between IDS models and human users, but the generated explanations should be grounded in detector outputs and verified evidence to avoid unsupported conclusions.

III-D3 LLM Tabular Data Generation

In 2024, Kim et al. [121] investigated the effectiveness of using LLM to generate synthetic data that addresses the class imbalance in tabular data. The paper found that using CSV-style prompting (compared to sentence-style in GReaT) can significantly improve the ability of LLM to generate accurate and balanced data, enhancing ML performance for minor classes in imbalanced data.

Not all data generated by LLMs is equally valuable and useful to downstream model performance; some samples may be harmful. Thus, assessment of the generated data is vital for any generative model. In 2024, Seedat et al. [122] introduced a method called Curated LLM (CLLM), which aims to generate synthetic tabular data in environments where data is scarce ($n<100$) and to apply a rigorous data curation process to ensure the quality of the generated data. The CLLM first harnesses the prior knowledge embedded in LLMs, using them to generate synthetic datasets based on a small number of real examples. Then, it relies on a curation mechanism that uses metrics like predictive confidence and uncertainty to filter and refine the generated data, improving its utility for downstream ML tasks. The paper used several real-world datasets to demonstrate the superior performance of CLLM over conventional generators (including CTGAN, TVAE, TabDDPM, SMOTE, and GReaT) in the low-data regimes.

In many domains where data privacy is crucial, synthetic data generated from real datasets can be used to avoid exposing real-world data. However, synthetic data can still contain the original dataset pattern or details. Differential privacy (DP), which introduces randomness in the data generation, is a promising approach to reducing the risk of re-identifying individuals. In 2024, Tran et al. [123] introduced DP-LLMTGen (Differentially Private LLM-based Tabular data Generators), a novel framework designed to generate synthetic tabular data while preserving DP. The framework utilizes a two-stage fine-tuning procedure with a novel loss function specifically designed for tabular data. The first stage focuses on learning the data format using non-sensitive, randomly generated data with original sensitive data. The second stage fine-tunes the LLM with DP mechanisms to ensure that the generated data maintains privacy while accurately capturing the feature distributions and dependencies of the original dataset. Then, synthetic data are generated by sampling from the fine-tuned LLM. The experiment result showed that the proposed DP-LLMTGen framework is able to effectively generate high-fidelity synthetic tabular data while preserving differential privacy.

To address the inefficiency and high computational costs associated with using LLMs for tasks involving tabular data, Einy et al. [124] proposed a selective enrichment approach. Their method aims to use LLMs to enrich tabular data to enhance the performance of classical ML models. LLMs are applied only to specific parts of the data that benefit the most from the additional contextual knowledge that LLMs provide. The result demonstrates that this approach can significantly enhance the performance of ML models on tabular data while maintaining cost-effectiveness.

However, in 2024, Xu et al. [125] demonstrated that LLMs are generally inadequate for tabular data generation when used directly or even after traditional fine-tuning. They suggest that due to their autoregressive nature, LLMs struggle to model the complex conditional dependencies and mixture distributions that exist in real-world tabular data. Further, feature ordering becomes more important when the dataset grows, and incorrect feature order can significantly degrade the quality of the generated data. The authors proposed a novel approach called Permutation-aided Fine-tuning (PAFT). Although the results show that PAFT can reproduce underlying relationships in generated data, there is still a significant gap between the current capabilities of LLMs and the requirements for generating realistic synthetic tabular data.

(a) Centralized machine learning

(b) Federated learning

(a) Centralized machine learning

III-E Challenges and Valuable Research Directions

Generative AI has been applied to IDS mainly to address data-related limitations. When labeled attack samples are limited, generative models can create additional samples for training. When datasets contain missing or incomplete records, generative models can support data imputation by estimating missing values from learned feature relationships. When datasets are imbalanced, generative models can generate samples for minority attack classes and reduce the bias of IDS models toward majority classes. Although these applications are useful, the use of generated data in IDS also introduces several challenges.

The first challenge is the quality of synthetic data. Generated samples are usually evaluated through their effect on downstream IDS models, such as whether they improve classification accuracy, recall, or robustness. However, synthetic data do not always improve downstream performance. Low-quality samples may introduce noise, distort class boundaries, or cause the detector to overfit artificial patterns. In addition, the performance of generative models is sensitive to design choices and parameter settings, such as latent-space design in VAEs, training stability in GANs, and noise schedules or sampling steps in diffusion models [126, 127, 128, 76, 129]. Therefore, selecting appropriate generative model configurations remains an important problem for IDS applications.

The second challenge is the reliability and realism of synthetic network traffic. Common distribution-level metrics, such as Kullback-Leibler (KL) Divergence [130], Jensen-Shannon (JS) Divergence [131], Wasserstein Distance [132], Fréchet Inception Distance (FID) [133], Maximum Mean Discrepancy (MMD) [134], Perceptual Path Length (PPL) [135], Energy Distance [136], and Precision and Recall for Distributions [137], can measure similarity between real and generated data distributions. However, distributional similarity alone is not sufficient for IDS. Synthetic traffic should also preserve protocol constraints, temporal dependencies, attack semantics, and network-topology relationships. For example, generated traffic may be statistically similar to real traffic but still invalid in a real testbed if packet sequences violate protocol behavior or if flows do not match the underlying network topology. Future work should develop IDS-specific evaluation methods that assess both statistical fidelity and network-level validity.

The third challenge is the limited use of Large Language Models (LLMs) for IDS-specific data generation. Existing LLM-based generation methods have shown potential for text and tabular data generation, including prompt-based synthetic data generation [138, 139, 140]. However, IDS data often contain numerical features, protocol-dependent relationships, temporal patterns, and topology-aware constraints, which are difficult for general-purpose LLMs to model directly. At the same time, LLMs provide a potential advantage because they can incorporate textual descriptions, domain knowledge, network configurations, and attack procedures during generation. This makes LLMs a promising direction for topology-aware and knowledge-guided IDS data generation. Future research should investigate how to adapt LLMs to network-security domains, how to ground generated traffic in valid network behavior, and how to evaluate whether LLM-generated samples are useful for IDS training and testing.

Overall, generative AI can support IDS by improving data availability, diversity, and robustness. However, future studies should move beyond simply generating more samples. More attention is needed on sample quality control, IDS-specific realism evaluation, privacy preservation, adversarial misuse, and domain-specific generative models that understand network protocols, attack behaviors, and deployment environments.