IV Generative AI Embedded Federated Learning based Intrusion Detection System
{forest}
Fig. 12: Generative AI with intrusion detection system
Federated Learning (FL) is a distributed learning paradigm designed for scenarios where data are generated and stored across many clients. It was originally motivated by applications such as mobile-device intelligence, where user data are distributed across devices and cannot be easily collected at a central server because of privacy and communication constraints [3]. Instead of transferring raw data to the server, FL allows clients to train models locally and share model updates for aggregation. Figure 11 illustrates the difference between centralized Machine Learning (ML) and FL.
In centralized ML, as shown in Figure 11(a), clients send their local datasets to a central server. The server then combines the collected data, trains an ML model, and distributes the trained model for future prediction. This approach is simple to implement when data can be centrally collected, but it may expose sensitive information and introduce high communication costs when the local datasets are large.
In FL, as shown in Figure 11(b), training is performed through repeated collaboration between the server and clients. The server first initializes a global model and sends it to participating clients. Each client trains the model using its local dataset and returns the updated model parameters or gradients to the server. The server then aggregates these local updates to obtain a new global model and sends the updated global model back to clients for the next training round. FedAvg is a representative aggregation method that computes a weighted average of local model updates to construct the global model [3]. By avoiding direct raw-data sharing, FL can reduce privacy risks and raw-data transmission costs, although communication overhead and potential information leakage from model updates remain important concerns.
FL is particularly relevant to IDS because network traffic is naturally distributed across routers, edge devices, organizations, and geographic regions. In a traditional centralized IDS training pipeline, clients or network devices send traffic records to a central server, and the server uses the aggregated data to train a detection model. This process can be expensive for high-volume traffic and may expose sensitive information, such as user behavior, service configurations, or organization-specific security patterns. The problem becomes more significant for high-rate attacks [141], such as denial-of-service (DoS) and distributed denial-of-service (DDoS) attacks, which can generate a large number of traffic records in a short period. FL-based IDS addresses this limitation by allowing each client to train locally while sharing only model updates with the server.
However, FL-based IDS also introduces challenges that are different from centralized IDS. Client data are often non-independent and identically distributed (non-IID) [142] because different clients may observe different traffic volumes, device types, services, and attack categories. Some clients may have limited attack samples, while others may have highly imbalanced traffic distributions. In addition, clients may have different computation and communication capabilities, and the FL process may be vulnerable to poisoning attacks or unreliable updates. These challenges motivate the integration of generative AI with FL-based IDS. Generative models can potentially augment local data, mitigate class imbalance, support privacy-preserving synthetic data generation, improve robustness, and reduce the effect of heterogeneous client distributions.
This section reviews generative AI-embedded FL-based IDS according to the same generative model families discussed in Section III. Section IV-A discusses VAE-embedded FL-IDS, where autoencoder-based models can support representation learning, anomaly detection, and communication reduction. Section IV-B reviews GAN-embedded FL-IDS, with a focus on data augmentation, adversarial traffic generation, and class-imbalance mitigation. Section IV-C discusses diffusion-embedded FL-IDS and its potential for synthetic data generation under distributed settings. Finally, Section IV-D discusses the emerging role of LLMs in FL-based IDS, including their potential use for explanation, knowledge-guided generation, and network-security analysis.
IV-A Autoencoder-embedded FL-IDS
Autoencoder-based models, including Autoencoders (AEs) and Variational Autoencoders (VAEs), have been used in FL-based IDS mainly for representation learning, anomaly detection, privacy preservation, and communication reduction. An AE learns to compress input data into a lower-dimensional latent representation and reconstruct the original input from that representation. This design is useful for IDS because network traffic often contains high-dimensional flow features, and the learned latent representation can preserve important traffic patterns while reducing feature size. A VAE extends the AE by modeling the latent representation as a probability distribution, which enables sampling and synthetic data generation from the learned distribution.
In FL-based IDS, autoencoder-based models are useful because clients can learn compact representations of local traffic without directly sharing raw network data. Instead of transmitting full datasets, clients may share model updates or compressed latent representations with the server. This can reduce communication cost and limit direct exposure of sensitive traffic records. In addition, reconstruction error from an AE or VAE can be used for anomaly detection, where traffic samples with large reconstruction errors are treated as potential intrusions [143, 54, 144, 145].
Tayeen et al. [146] proposed CAFNET, a compressed autoencoder-based federated framework for network anomaly detection. CAFNET reduces communication overhead by transmitting compact representations between clients and the server rather than transferring raw data or large model components. The reported results show that CAFNET maintains detection performance while reducing communication cost by up to 95%. NR et al. [147] also used VAE-based federated learning for intrusion detection in Industrial IoT environments, aiming to improve privacy protection and reduce communication overhead.
Overall, AE- and VAE-embedded FL-IDS methods are suitable for settings where network traffic is high-dimensional, privacy-sensitive, and distributed across multiple clients. However, these methods still face several challenges. The compressed latent representation must preserve enough information for accurate intrusion detection, while avoiding unnecessary leakage of sensitive traffic patterns. In addition, when client data are non-IID, the latent spaces learned by different clients may not be fully aligned, which can affect aggregation and global anomaly detection performance.
IV-B GAN-embedded FL-IDS
GANs can be integrated with FL-based IDS from two main perspectives. The first perspective treats GAN-generated traffic as a security threat. As discussed in Section III-B2, GANs can generate adversarial traffic records that resemble real traffic and are designed to bypass IDS models. For example, Lin et al. [148] introduced IDSGAN for generating adversarial attack traffic against IDS. Aldhaheri and Alhuzali [68] proposed SGAN-IDS, and Zhang et al. [149] studied poisoning attacks against FL-based network intrusion detection. These studies show that generated malicious traffic can be difficult to detect because it is optimized to remain close to legitimate traffic patterns while misleading the detection model.
In FL-based IDS, this problem becomes more complex because the server does not directly access clients’ raw traffic data. The server mainly observes model updates, which makes it harder to inspect whether local training data contain GAN-generated adversarial samples or poisoned records. Therefore, detecting and defending against GAN-based adversarial traffic remains an important challenge for FL-based IDS. Vy et al. [150] studied poisoning attacks and defense mechanisms in an FL-based IDS framework for Industrial IoT networks, showing that adversarial manipulation must be considered when deploying IDS models in distributed environments.
The second perspective uses GANs as a defensive tool for data augmentation. In FL-based IDS, clients often hold non-IID and imbalanced local datasets because different networks observe different services, devices, traffic volumes, and attack categories. A local GAN can generate additional samples for minority classes and improve the local training distribution without requiring clients to share raw traffic data. Tabassum et al. [151] proposed FEDGAN-IDS, where GAN-based augmentation is used with FL to address data imbalance while preserving privacy.
Despite this potential, GAN-embedded FL-IDS faces communication and training challenges. A GAN usually contains at least two model components, a generator and a discriminator, which increases training complexity and may introduce additional communication cost if the GAN parameters are shared across clients and the server. Model compression and scaling studies, such as [152, 153], show that model size, input size, and performance must be balanced carefully. For FL-based IDS, this means that GANs should be designed or deployed in a way that improves local data quality without offsetting the privacy and communication advantages of FL.
IV-C Diffusion-embedded FL-IDS
Diffusion models are an emerging direction for generative AI-embedded FL-based IDS. Compared with GANs, diffusion models often provide more stable training because they do not rely on adversarial optimization between a generator and a discriminator. This property is useful in FL, where clients may have heterogeneous and imbalanced local data. However, diffusion models can also require substantial sampling computation, so their communication and computation costs must be evaluated carefully in distributed IDS settings.
Jothiraj and Mashhadi [154] introduced Phoenix, a federated generative diffusion framework that uses diffusion models to improve training data diversity across clients. Their study compared diffusion-based generation with GAN-based generation in an FL setting and showed that diffusion models can generate high-quality samples while reducing communication cost. Although Phoenix is not limited to IDS, its design is relevant to FL-based IDS because local clients often have limited, imbalanced, and non-IID traffic data.
For FL-based IDS, diffusion models can potentially support local data augmentation, minority-class sample generation, and privacy-preserving synthetic traffic generation. Instead of sharing raw traffic data, clients may use diffusion models to enrich local training data or share generative knowledge with the server. This direction is especially useful when attack samples are rare or unevenly distributed across clients. Nevertheless, diffusion-embedded FL-IDS remains underexplored. Future studies need to examine whether diffusion-generated traffic preserves realistic network behavior, whether the generation process is efficient enough for edge or IoT clients, and how diffusion models can be integrated with FL without increasing communication overhead.
IV-D LLM-embedded FL-IDS
LLM-embedded FL-IDS remains an emerging direction with limited direct studies. In this context, an LLM can be integrated into an FL-based IDS pipeline in several ways, including traffic classification, anomaly detection, data augmentation, data imputation, alert explanation, and security knowledge extraction. These capabilities are relevant to FL-based IDS because clients often have limited, imbalanced, and heterogeneous local traffic data. For example, LLM-based data augmentation or imputation may help enrich local datasets without requiring clients to share raw traffic records.
However, applying LLMs directly in FL-based IDS is challenging. The main limitation is the high communication and computation cost of training or fine-tuning large models across distributed clients. Unlike smaller IDS models, LLMs contain a large number of parameters, making full-model transmission impractical for many edge, IoT, or organizational clients. In addition, IDS data are often represented as numerical flow features, packet sequences, or structured logs, while general-purpose LLMs are primarily trained on natural-language text. Therefore, effective input representation and domain adaptation are necessary before LLMs can be reliably used for IDS tasks.
A practical direction is to avoid transmitting full LLMs in the FL process. Instead, future studies may explore parameter-efficient tuning, adapter-based learning, knowledge distillation, or server-side LLM assistance. In these settings, clients may train lightweight local IDS models, while LLMs provide auxiliary functions such as generating domain-informed synthetic samples, explaining alerts, summarizing attack behaviors, or supporting security analysts. Overall, LLM-embedded FL-IDS has potential, but its feasibility depends on reducing communication cost, grounding LLM outputs in valid network behavior, and adapting LLMs to network-security-specific data formats.
IV-E Challenges and Valuable Research Directions
Generative AI-embedded FL-based IDS introduces both opportunities and risks. On the defensive side, generative models can support local data augmentation, mitigate class imbalance, improve robustness, and reduce the need to share raw traffic data. On the offensive side, the same generative models can be used to create synthetic malicious traffic that resembles benign traffic or bypasses IDS models. This dual-use nature makes generative AI important for both improving FL-based IDS and evaluating its vulnerability under adaptive attacks.
The first research challenge is defending FL-based IDS against generative adversarial traffic. GANs and diffusion models can generate traffic samples that are statistically close to real traffic but intentionally optimized to mislead a detector. In FL-based IDS, this threat is harder to identify because the server usually receives model updates rather than raw local traffic. As a result, adversarial or poisoned local data may influence the global model through aggregation without being directly inspected. Future work should study how to detect generative adversarial traffic in distributed settings, how to distinguish malicious client updates from benign non-IID updates, and how to design robust aggregation methods for FL-based IDS.
The second research direction is using generative AI to address non-independent and identically distributed (non-IID) client data. In FL-based IDS, each client may observe different devices, services, traffic volumes, and attack categories. This heterogeneity can reduce the quality of the global model because local updates are optimized on different data distributions. Generative models can reduce this problem by augmenting local datasets, generating minority-class samples, or improving local data diversity before model training. Existing studies have explored VAE-, GAN-, and diffusion-based methods for handling non-IID data in FL [155, 156, 157]. For IDS, this direction is especially relevant because rare attacks may appear only on a small subset of clients.
The third challenge is communication efficiency. Although generative models can improve local training data, transmitting large generative models between the server and clients may increase communication cost. This problem is more significant for GANs and LLMs because they may contain large model components or require expensive fine-tuning. In contrast, some VAE- and diffusion-based FL methods can be designed to reduce communication by sharing compact representations, selected parameters, or generated knowledge instead of full datasets or full models. Recent studies on communication-efficient federated diffusion learning show that diffusion-based strategies can reduce communication cost while maintaining model performance [158, 159]. Future FL-based IDS studies should evaluate not only detection accuracy, but also communication cost, client computation cost, and scalability.
The fourth research direction is realistic federated IDS data generation and benchmarking. As discussed in Section II-E, realistic FL-based IDS datasets remain limited. Most FL-based IDS studies use centrally collected datasets and partition them artificially across clients. Although this strategy is convenient, it may not reflect real client-level heterogeneity, network topology, temporal changes, or organization-specific attack patterns. FLNET [19] provides an important step toward FL-oriented IDS evaluation, but more datasets are needed to represent diverse federated deployment scenarios. Generative AI may help create topology-aware and client-specific traffic data, but the generated data must be evaluated for both statistical similarity and network-level validity.
The fifth research direction is LLM-assisted FL-based IDS. To the best of our knowledge, direct studies on LLM-embedded FL-based IDS are still limited. However, LLMs may support FL-based IDS through domain-informed data generation, alert explanation, attack-behavior summarization, traffic-log analysis, and response recommendation. Compared with VAEs, GANs, and diffusion models, LLMs may be better suited for incorporating textual domain knowledge, such as protocol descriptions, network configurations, and attack procedures. However, most pre-trained LLMs are general-purpose models and are not optimized for network-security data. Future studies should investigate network-domain-specific LLMs, parameter-efficient adaptation, and methods for grounding LLM outputs in valid traffic behavior and verified security evidence.
Overall, generative AI-embedded FL-based IDS should be evaluated from multiple perspectives, including detection performance, robustness to adversarial generation, privacy protection, communication efficiency, and realism of generated traffic. Future research should move beyond using generative models only as data generators and study how they can be safely integrated into distributed IDS training and deployment.