Generative AI and Federated Learning for Intrusion Detection Systems: A Survey
Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Internet of Things (IoT), enterprise, and distributed network environments. However, developing reliable IDS models remains challenging because attack behaviors evolve over time, realistic datasets are difficult to obtain, traffic records may be incomplete, attack classes are…
Generative AI and Federated Learning for Intrusion Detection Systems: A SurveyThanks: Jiefei Liu, Pratyay Kumar, Qixu Gong, Wenbin Jiang, Huiping Cao, and Satyajayant Misra are with the Department of Computer Science, New Mexico State University, Las Cruces, NM, USA (e-mail: {jiefei, pratyay, qixugong, wbjiang, hcao, misra}@nmsu.edu). Thanks: Abu Saleh Md Tayeen is with the University of Hartford, CT, USA (e-mail:tayeen@hartford.edu). Thanks: Jayashree Harikumar is with DEVCOM Analysis Center, WSMR, NM, USA (e-mail: jayashree.harikumar.civ@army.mil).
Built from what the sources declared and what the gates observed.
Nothing absent has been filled in here.
1 value read out of the text by the enrichment rules and 94 links to or from other resources — a lab's files, the pages it links, the works it cites stand under their elements, marked inferred, and are kept apart in the exports.
DCMI Metadata Terms. Dublin Core has no element that separates the original file from the text extracted out of it, and none for LOM's educational characterisation. Both survive here as provenance statements and in the record itself, not in the projection.
the standard ↗
The groups below are this library's, for reading. DCMI Terms itself has no categories; each term keeps its standard name.
Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Internet of Things (IoT), enterprise, and distributed network environments. However, developing reliable IDS models remains challenging because attack behaviors evolve over time, realistic datasets are difficult to obtain, traffic records may be incomplete, attack classes are often imbalanced, and privacy constraints limit centralized data collection. Recent advances in generative artificial intelligence (AI) and Federated Learning (FL) provide new opportunities to address these limitations. Generative models can support anomaly detection, synthetic traffic generation, data augmentation, data imputation, adversarial traffic generation, and IDS alert explanation. FL enables distributed IDS training without directly sharing local network traffic, making it suitable for privacy-sensitive and geographically distributed environments. This survey provides a structured review of generative AI and FL techniques for IDS. We first summarize representative IDS research directions, including adversarial machine learning, anomaly-based detection, IoT-oriented IDS, explainable IDS, and benchmark datasets. We then categorize generative AI applications in IDS according to model families and task objectives, covering autoencoder-based models, Generative Adversarial Networks (GANs), diffusion models, and Large Language Models (LLMs). Finally, we review emerging studies that integrate generative AI with FL-based IDS and discuss open challenges, including synthetic data quality, realistic traffic generation, dual-use adversarial risks, non-IID client distributions, communication-efficient model sharing, federated IDS benchmarking, and domain-specific LLMs for network security.
The abstract, and the depositor's additional notes after it when the source has a field for them.
Where the work was published, in the source's own words: a journal with its volume and pages, a conference, an imprint.
Relations and custody
1/9
Is part ofdcterms:isPartOf
none found — the source did not declare it
The repository, book or record this resource was found inside.
Has partdcterms:hasPart
none found — the source did not declare it
What this resource is made of, when the source lists its parts.
Is version ofdcterms:isVersionOf
none found — the source did not declare it
Has versiondcterms:hasVersion
none found — the source did not declare it
Referencesdcterms:references
inferred
read its reference list, line 421citation
doi:10.1109/tse.1987.232894
doi:10.1016/j.comnet.2022.109073
doi:10.3390/fi15020062
doi:10.1016/j.cose.2022.102675
doi:10.1109/access.2022.3220622
doi:10.1109/comst.2023.3280465
doi:10.1109/access.2022.3216617
doi:10.21227/qvr7-n418
doi:10.1109/tvt.2023.3327275
arXiv:1312.6114
doi:10.1109/access.2020.3001350
doi:10.1109/jiot.2020.3034621
and 82 more, every one of them in the export
Works this one cites, when the source declares them as relations. What its text links to and its reference list cites is inferred, and stands under it apart.
Is referenced bydcterms:isReferencedBy
none found — the source did not declare it
What links to or cites this one, when a source declares it; inferred from the texts held otherwise, and set apart.
Requiresdcterms:requires
none found — the source did not declare it
What this resource needs in order to be used, when a source declares it. The files a lab works on are inferred, and stand under it apart.
Is required bydcterms:isRequiredBy
none found — the source did not declare it
What needs this resource, when a source declares it. The lab a component belongs to is inferred, and stands under it apart.
Provenancedcterms:provenance
this engine
source
conversion
latexml-html
Retrieved from arXiv on 2026-10-09 in response to the search string “(all:"artificial intelligence" OR all:"machine learning" OR all:"generative AI" OR all:"deep learning" OR all:"reinforcement learning" OR all:"large language model") AND (all:"AI concepts" OR all:"types of AI" OR all:"AI fundamentals" OR all:"recognizing AI" OR all:"recognising AI" OR all:"general versus narrow AI" OR all:"narrow AI" OR all:"general AI" OR all:"machine intelligence" OR all:"AI strengths and weaknesses" OR all:"traditional software" OR all:"rule-based systems" OR all:"introduction to AI" OR all:"introduction to artificial intelligence" OR all:"artificial intelligence introduction" OR all:"AI primer" OR all:"foundations of artificial intelligence" OR all:"overview of AI" OR all:"understanding AI" OR all:"history of AI" OR all:"AI essentials" OR all:"AI terminology" OR all:"metaphors for AI" OR all:"AI fundamental concepts" OR all:"AI key concepts" OR all:"philosophy of AI" OR all:"critical AI literacy")”. arXiv served the resource and is not asserted to be its publisher or author.
Text extracted from latexml-html to Markdown by arxiv-html; the original is retained unchanged beside it.
Where it was collected from, what was converted, and what container it came out of — the custody statements that would otherwise be mistaken for authorship.
IEEE 1484.12.1 Learning Object Metadata. LOM has no element for an SPDX identifier or a licence URI, so both are written into 6.3 Rights.Description. Flattening this record into simple Dublin Core would lose more again, which is why the two projections exist side by side rather than one being generated from the other.
the standard ↗
Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Internet of Things (IoT), enterprise, and distributed network environments. However, developing reliable IDS models remains challenging because attack behaviors evolve over time, realistic datasets are difficult to obtain, traffic records may be incomplete, attack classes are often imbalanced, and privacy constraints limit centralized data collection. Recent advances in generative artificial intelligence (AI) and Federated Learning (FL) provide new opportunities to address these limitations. Generative models can support anomaly detection, synthetic traffic generation, data augmentation, data imputation, adversarial traffic generation, and IDS alert explanation. FL enables distributed IDS training without directly sharing local network traffic, making it suitable for privacy-sensitive and geographically distributed environments. This survey provides a structured review of generative AI and FL techniques for IDS. We first summarize representative IDS research directions, including adversarial machine learning, anomaly-based detection, IoT-oriented IDS, explainable IDS, and benchmark datasets. We then categorize generative AI applications in IDS according to model families and task objectives, covering autoencoder-based models, Generative Adversarial Networks (GANs), diffusion models, and Large Language Models (LLMs). Finally, we review emerging studies that integrate generative AI with FL-based IDS and discuss open challenges, including synthetic data quality, realistic traffic generation, dual-use adversarial risks, non-IID client distributions, communication-efficient model sharing, federated IDS benchmarking, and domain-specific LLMs for network security.
The abstract, and the depositor's additional notes after it as a second LangString when the source has a field for them.
Role, entity and date per declared contribution. A role outside LOM's vocabulary is reported in the entry's description instead.
3 Meta-metadata
4/4
Identifier3.1
this engine
resource_id
URI: tag:aim-pro.eu,2026:oer/f40027e54d10/record
The identifier of this metadata record — the resource's own, with /record after it, because the record is a description of the resource and not the resource.
Contribute3.2
this engine
source
AIM-PRO WP3 OER harvester (arXiv) — creator
Who generated this record and when — a statement about the record, not about the resource.
Metadata schema3.3
LOMv1.0
aimpro-oer-profile/1
LOMv1.0, and the profile this was built against.
Language3.4
language gate
read the declared field
en
4 Technical
3/7
Format4.1
conversion
latexml-html
text/markdown
One value per form held: the original as the source published it, and the Markdown this engine extracted.
Size4.2
conversion
latexml-html
137025
Bytes. The original's, because the resource is the file and not our conversion of it.
Location4.3
this engine
resource_id
https://arxiv.org/abs/2607.01305
Where the source serves it.
Requirement4.4
not collected — this library does not fill it
Software or hardware needed to use it. No source declares it.
Installation remarks4.5
not collected — this library does not fill it
No source declares it.
Other platform requirements4.6
not collected — this library does not fill it
No source declares it.
Duration4.7
not collected — this library does not fill it
Playing time, for audio and video. The corpus holds neither.
5 Educational
1/11
Interactivity type5.1
not collected — this library does not fill it
Active, expositive or mixed. A judgement about how the resource is used; no source declares it.
Yes unless the licence reserves nothing — attribution is a restriction. The conditions after the dash are the licence gate's reading; the export carries LOM's bare term.
The SPDX id, the licence URI and the conditions. LOM has no element for any of the three, so this is where they survive.
7 Relation
0/2
Kind7.1
none found — the source did not declare it
Resource7.2
inferred
read its reference list, line 421citation
references: doi:10.1109/tse.1987.232894
references: doi:10.1016/j.comnet.2022.109073
references: doi:10.3390/fi15020062
references: doi:10.1016/j.cose.2022.102675
references: doi:10.1109/access.2022.3220622
references: doi:10.1109/comst.2023.3280465
references: doi:10.1109/access.2022.3216617
references: doi:10.21227/qvr7-n418
references: doi:10.1109/tvt.2023.3327275
references: arXiv:1312.6114
references: doi:10.1109/access.2020.3001350
references: doi:10.1109/jiot.2020.3034621
and 82 more, every one of them in the export
The container a file was found inside, and the Markdown extracted from the original. What a lab requires, and the lab a component belongs to, are inferred and stand apart.
8 Annotation
0/3
Entity8.1
not collected — this library does not fill it
Comments on the resource's educational use, by whoever made them. The platform's review grades competencies, which are classification (9), and writes no comment here.
Date8.2
not collected — this library does not fill it
Description8.3
not collected — this library does not fill it
9 Classification
0/4
Purpose9.1
not collected — this library does not fill it
Empty in the record: no source declares a competency. The alignment reads the resource for them and stands beside the record, never in it, and a taxon path derived from the search string that found it would be a claim about the query.
Taxon path9.2
not collected — this library does not fill it
Where the competency framework goes. Empty in the record for the reason above.
Description9.3
not collected — this library does not fill it
Keyword9.4
not collected — this library does not fill it
What could not be established3
Where the source's metadata could not be carried
over as it was — missing, contradictory, with no matching term in the
standard, restructured, or taken from the repository — and what was done
instead. Without these notes, an empty element would look like something
the harvester missed.
Status
Field
Why
Not available
rights_holder
no rights holder is named at source; the licence is recorded without one rather than attributed to the platform that served it
Not available
publisher
the source named no publisher of the work; where it was collected from is recorded as collection provenance instead, which is a different claim
Not available
educational
the source declared no educational metadata — no resource type, audience, context, difficulty or learning time. Nothing here estimates them