Generative AI for Design and Manufacturing
Faez Ahmed¹, Wei “Wayne” Chen, and Mark Fuge³
1 Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, USA 2
J. ike Walker ’66 epartment of echanical Engineering, exas & University, College Station, USA 3 Department of Mechanical and Process Engineering, ETH Zürich, Switzerland E-mail: faez@mit.edu
Status
Why Can’t Machines Design Other Machines (Yet)? Engineers have long imagined a world where machines could design other machines [1,2], and recent advances in generative AI systems, such as diffusion models and large language models (LLMs), have revived that dream. They can already draft code, summarize documents, and even propose initial engineering concepts. So why aren’t aerospace companies, for example, using them to design and certify entire UAVs from scratch? This paper argues that the problem is not a lack of imagination, but a set of scientific and organizational barriers. To map that landscape, we divide the problem along two axes (Fig.1). On one axis lies design depth: how well AI can perform specialized tasks such as modeling 3D geometry, surrogate modeling of simulations,
design optimization, or uncertainty quantification. On the other axis lies design breadth: how well AI can integrate data and knowledge across domains and the entire product lifecycle. Each axis is constrained in two ways. First, there are scientific limits—where the methods themselves are not yet capable. Second, there are adoption barriers—issues of trust, data, interoperability, and organizational inertia. Taken together, these four quadrants (Fig.1) capture key reasons why AI remains more promise than practice in real-world engineering design. This paper first reviews some of those barriers, and then proposes a brief roadmap to overcome them.
Current and future challenges
1.1 Scientific Barriers
1.1.1 Scientific Barriers to Breadth: Why AI Struggles Across the Design Process
Although there are many barriers to learning design breadth, the five main issues are: (1) lack of interoperability among design representations, (2) representation heterogeneity, (3) limited model reasoning, (4) inefficient human-AI Collaboration, and (5) inability to explore high-dimensional design spaces. The first obstacle to breadth is fragmented data. Modern engineering lifecycles generate enormous amounts of information, including requirements documents, CAD models, simulations, test logs, bills of materials, and manufacturing documents, but these are scattered across incompatible systems. Without consistent, interoperable histories and formats, AI systems cannot learn end-to-end mappings or models that span the entire design lifecycle. The second problem is heterogeneity. Representations vary in different domains—meshes, splines, point clouds, bills of materials—and subsystems interact in strongly coupled, nonlinear ways. Design often requires multi-discipline, multi-fidelity models in a multi-code environment [3]. Today’s domain-specific surrogates break down in these out-of-distribution (OOD) regions. Real-world design may require modeling of behaviors outside of the original training data, such as airplanes facing flutter, control instability, or electronic–thermal interactions. The third obstacle is a gap in how deeply existing models can reason and plan. Language models can stitch together ideas from diverse fields [4], but their tendency to hallucinate [5] makes them unreliable as integrators of mission-critical systems. Currently, they are often limited to assisting in brainstorming or connecting high-level knowledge [6]. The risk of producing plausible but unsafe or underperforming designs limits their applicability. The fourth obstacle is the underdeveloped science of human–AI collaboration. We do not yet understand how engineers actually interact with modern AI tools across the lifecycle, how trust is built, how tacit goals are communicated, how authority is shared, and what roles are best played by machines versus humans [7]. Lastly, the fifth obstacle is that many current AI models struggle to generate transformative, detailed designs. Models trained on historical data interpolate well, but can they genuinely invent? As design spaces scale combinatorially, brute-force data-driven or optimization approaches become untenable. We need reliable methods that compose knowledge and search efficiently and effectively. Yet methods and benchmarks today focus almost exclusively on narrow, single-stage tasks, not on cross-stage innovation.
Figure 1. We Categorize Current Challenges and Opportunities Across Four Quadrants.
1.1.2 Scientific Barriers to Depth: Why Narrow Models Still Fail
Scientific obstacles to achieving depth fall into four main categories: (1) difficulties generalizing, (2) lack of model diagnostics, (3) fast and accurate verification of outputs, and (4) integrating AI with existing toolsets. First, even within single tasks, AI struggles to generalize. Datasets in engineering are sparse, proprietary, and task-specific. Transfer learning has shown promise in domains like computer vision to reduce data demand [8], but it is less effective in engineering design due to the high heterogeneity of data, representations, and problems. For instance, a model trained on automotive aerodynamics is unlikely to transfer to UAV wings. Engineers need ways to know when a model is operating inside its training manifold and when it is not. A second challenge is understanding when machine learning models fail and why. Simulation tools, such as CFD solvers, can be checked via methods like mesh convergence. Machine learning models, in contrast, are often difficult to debug, and their error bounds are hard to interpret, especially when the training data is hidden. Engineers require calibrated uncertainty estimates and transparent diagnostics for reliably using machine learning. A third hurdle is verification. If an AI proposes a design, how do we know it meets our needs? In some cases, physics-based checks are possible, but in most of the applications, just physics-based verification may not be sufficient, and human verification is difficult. The last hurdle is the need for integration. Too often, AI seeks to replace well-established tools rather than complement them. Hybrid tools could be helpful in such situations: AI accelerating optimizers, suggesting experiments, or guiding simulations—while integrating with the existing design process, tools, and human users.
1.1.3 Scientific Desired Future State: What Machines Must Learn
What would it take for machines to design machines? The answer lies in four broad AI capabilities.
First, composition: AI must recognize cross-domain couplings and emergent phenomena [3]. It must understand how local geometry affects global aeroelasticity, or how additive manufacturing paths influence microstructure and thus fatigue life. Second, abstraction: AI should learn to build and select the right surrogate models, at the right fidelity, and know when those abstractions fail. Like the Wright brothers, it should be able to design experiments to correct its own theories. Third, decision-making under uncertainty: AI must plan experiments and simulations intelligently, reuse knowledge across tasks, quantify transfer uncertainty, and recognize when “enough” is enough. Fourth, collaboration with humans and society: AI must elicit intent, surface tradeoffs, justify its decisions, and pass ethical and regulatory scrutiny. It must know when to ask for help. Milestones on the way include the ability to detect wrong theories, uncover emergent hazards, reframe ill- posed problems, reveal hidden tradeoffs, adapt under constraints, demonstrate genuine novelty, and make credible analogies across domains. Finally, organizations themselves must be ready to absorb the advances.
1.2 Adoption Barriers
1.2.1 Adoption Barriers to Breadth: Organizational Gravity
Even when the methods are proficient, organizations face practical roadblocks. Interoperability is one: data integration across tools, vendors, and legacy systems is expensive and difficult to manage and maintain. Intellectual property and privacy are also sensitive topics. Companies need assurance that their proprietary models and datasets will not leak into public systems or to other companies. Feedback loops are another. Many critical objectives in real-world design scenarios—ease of inspection, maintainability, evolving regulatory priorities—are not explicitly written into AI objectives. As the situation evolves, humans must be able to impose new constraints, and the AI tools must adapt. Process fit is an important barrier. Engineering workflows rely on version control, iterative reviews, and certification milestones. AI tools that don’t align with these are often abandoned. Cultural factors also play a role: engineers are skeptical of “creative” machines, and many enjoy doing the work themselves. Meanwhile, the workforce also lacks training. Engineers in subjects like civil, mechanical, or aerospace are trained to interpret CFD, FEA, or physical experimentation, not AI methods. We lack both the tools and the training for AI in design and manufacturing.
1.2.2 Adoption Barriers to Depth: Trust and Certification
For depth-focused tools, the key adoption issue is trust [9]. First, engineers want to know what went into the model: how much data, what quality, what diversity. Today, such “datasheets” and contextual information rarely exist. This problem is compounded by a lack of integration with existing tools. Second, engineers cannot rely on models that lack robustness, generalization, explainability, transparency, reproducibility, and accountability [10]. Certification adds another hurdle. Aviation regulators, for example, require auditable chains of evidence. If a design comes from a black-box AI, it is unclear who signs off. Incremental modifications “close to existing designs,” such as those obtained by iterative optimization, may be easier to certify, but that disincentivizes radical AI-supported innovation. Return on investment (ROI) is also difficult to assess due to a lack of standardized assessment. When does AI outperform classical methods, on what metrics, which problems, and by how much? Until clear comparative evidence exists and demos move to production, ROI for companies against trusted workflows might be hard to estimate.
1.2.3 Adoption Barriers to Depth: Trust and Certification
For depth-focused tools, the key adoption issue is trust [9]. First, engineers want to know what went into the model: how much data, what quality, what diversity. Today, such “datasheets” and contextual information rarely exist. This problem is compounded by a lack of integration with existing tools. Second, engineers cannot rely on models that lack robustness, generalization, explainability, transparency, reproducibility, and accountability.
Advances in science and technology to meet challenges
We discuss the roadmap for both research and deployment based on four pillars: data and infrastructure, model capability and generalization, workforce and organizations, and trust and compliance. (Four Pillars, Three Horizons, and Clear Metrics)
2.1 A Roadmap for Research
Data and Infrastructure. We need open, design-relevant datasets and benchmarks [11] that span modalities, disciplines, and lifecycle stages and measure metrics, such as manufacturability and ROI [12]. Shared experimental platforms should allow large-scale, controlled comparisons.
Model Capability and Generalization. Multimodal foundation models and agent-based architectures can interpret diverse engineering artifacts and manage interdependencies. But they must move beyond interpolation to robust OOD generalization: physics priors [13], causal rules [14], active learning [15], and meta-learning [16] could be key. Sample efficiency and navigating multi-modality are also needed in engineering.
Workforce and Organizations. We need ethnographic and human-computer interaction research on how engineers actually work, plus new interoperability standards to ease toolchain integration. To reduce the gap between academic studies and industry practice, AI integration in real engineering teams should be studied [17]. Education must also prepare “AI-fluent” engineers for hybrid workflows and real-world challenges.
Trust and Compliance. Models must produce explainable, auditable design traces that align with certification requirements [18]. Standardized evaluation frameworks should match regulatory expectations.
2.2 A Roadmap for Deployment
Data and Infrastructure. Near-term priorities are setting data governance standards, unifying platforms, and curating open datasets for benchmarking. Cross-industry investment producing large-scale synthetic and real-world datasets that span across domains and multimodal representations will expand GenAI’s scope. Ultimately, allowing AI to access solvers in a self-supervised fashion would enable its own scalable dataset augmentation.
Model Capability and Generalization. Early efforts should focus on model accuracy, efficiency, and producing constraint feasibility. Once achieved, the focus can then shift to robust generalization, handling multiple objectives, and reliable OOD performance. Ultimately, AI should move toward handling system-level design tasks while minimizing human guidance.
Workforce and Organizations. In the short term, organizations should launch AI literacy programs, update curricula, and run pilots to build skills and adapt workflows. Progress requires formal collaboration frameworks, hybrid teams, and shared best practices. Long term, the aim is an AI-fluent culture where engineers routinely co-design with GenAI, guided by workforce studies and adoption metrics.
Trust and Compliance. Short-term efforts must map AI-generated designs to existing certification standards and establish ethics/IP guidelines. Standardized validation, open benchmarks, and accountability frameworks will also be critical. Long term, formal certification pathways for AI-designed products will emerge, with regulators, industry, and academia ensuring reliability, safety, and compliance.
2.3 Measuring Progress: From Benchmarks to Real Outcomes
Metrics must track both high-level system outcomes and technical details. On the system side: reductions in design-cycle time, higher first-pass certification rates, broader cross-discipline integration, and workforce adoption and satisfaction. On the technical side: effectiveness, sample efficiency, calibrated uncertainty, OOD robustness, manufacturability, and ROI are key. Importantly, metrics must be disentangled across the four quadrants—breadth vs. depth, science vs. adoption—so stakeholders can see exactly where progress is happening and where it is not. Significant research and investments are still needed to establish a large number of diverse and realistic benchmarks for progress within Design and Manufacturing (e.g., [19]). By comparison, LLM progress has been accelerated due to a variety of diverse benchmarks [20] could the same benefits be brought to design?
Concluding remarks
Why Machines Don’t Design Machines? The reason machines don’t yet design machines is not a single missing breakthrough. It is a combination of scientific and institutional deficits. Scientifically, our models lack reliable generalization, uncertainty quantification, and system-level reasoning across heterogeneous, tightly coupled domains. Institutionally, we lack interoperable data infrastructure, certification frameworks, cultural acceptance, and workforce integration. The path forward is clear: build domain-relevant benchmarks and datasets; harden models with physics, uncertainty, and compositional reasoning; redesign workflows for human–AI collaboration; and codify trust and certification pipelines. Only by tackling science and adoption together, across both breadth and depth, will we turn AI design from clever demos into certified engineering practice.
Acknowledgements
PI Ahmed acknowledges support from the National Science Foundation under CAREER Award No. 2443429.
References
[1] Finger S and Dixon J R 1989 A review of research in mechanical engineering design. Part I: Descriptive, prescriptive, and computer-based models of design processes. Research in engineering design 1 51–67 [2] Antonsson E K and Cagan J 2001 Formal engineering design synthesis. (Cambridge University Press) [3] Antonau I, Warnakulasuriya S, Baars S, Baimuratov I, Wittenborg T, Kreuzeberg L, Attravanam A and W¨uchner R 2025 Challenges in realizing 3rd generation multidisciplinary design optimization. Advances in Computational Science and Engineering 5 1–21 [4] Bordas A, Le Masson P, Thomas M and Weil B 2024 What is generative in generative artificial intelligence? A design-based perspective. Research in Engineering Design 35 427–443 [5] Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, Ishii E, Bang Y J, Madotto A and Fung P 2023 Survey of hallucination in natural language generation. ACM computing surveys 55 1–38 [6] Ren R, Ma J and Luo J 2025 Large language model for patent concept generation. Advanced Engineering Informatics 65 103301 [7] Zhang G, Raina A, Cagan J and McComb C 2021 A cautionary tale about the impact of AI on human design teams. Design studies 72 100990 [8] Zhuang F, Qi Z, Duan K, Xi D, Zhu Y, Zhu H, Xiong H and He Q 2020 A comprehensive survey on transfer learning. Proceedings of the IEEE 109 43–76 [9] Siau K and Wang W 2018 Building trust in artificial intelligence, machine learning, and robotics. Cutter business technology journal 31 47
[10] Li B, Qi P, Liu B, Di S, Liu J, Pei J, Yi J and Zhou B 2023 Trustworthy AI: From principles to practices. ACM Computing Surveys 55 1–46 [11] Ahmed F, Picard C, Chen W, McComb C, Wang P, Lee I, Stankovic T, Allaire D and Menzel S 2025 Design by Data: Cultivating Datasets for Engineering Design. Journal of Mechanical Design 147 040301 [12] Briard T, Jean C, Aoussat A and V´eron P 2023 Challenges for data-driven design in early physical product design: A scientific and industrial perspective. Computers in Industry 145 103814 [13] Karniadakis G E, Kevrekidis I G, Lu L, Perdikaris P, Wang S and Yang L 2021 Physics-informed machine learning. Nature Reviews Physics 3 422–440 [14] Liu C, Sun X, Wang J, Tang H, Li T, Qin T, Chen W and Liu T Y 2021 Learning causal semantic representation for out-of- distribution prediction. Advances in Neural Information Processing Systems 34 6155–6170 [15] Serles P, Yeo J, Hach´e M, Demingos P G, Kong J, Kiefer P, Dhulipala S, Kumral B, Jia K, Yang S et al. 2025 Ultrahigh specific strength by Bayesian optimization of carbon nanolattices. Advanced Materials 37 2410651 [16] Hollmann N, M¨uller S, Purucker L, Krishnakumar A, K¨orfer M, Hoo S B, Schirrmeister R T and Hutter F 2025 Accurate predictions on small data with a tabular foundation model. Nature 637 319–326 [17] Edgecomb I M, Brisco R, Gunn K et al. 2025 Artificial Intelligence in engineering design: an industry perspective. Proceedings of the Design Society 5 641–650 [18] Hoffman R R, Johnson M, Bradshaw J M and Underbrink A 2013 Trust in automation. IEEE Intelligent Systems 28 84–88 [19] Felten F, Apaza G, Baunlich G, Diniz C, Dong X, Drake A, Habibi M, Hoffman N J, Keeler M, Massoudi S, VanGessel F G and Fuge M 2025 EngiBench: A Framework for Data-Driven Engineering Design Research URL http://arxiv.org/abs/2508.00831 [20] Clark P, Cowhey I, Etzioni O, Khot T, Sabharwal A, Schoenick C and Tafjord O 2018 Think you have solved question answering? try arc, the ai2 reasoning challenge (Preprint 1803.05457) URL https://arxiv.org/abs/1803.05457