OER·harvester

← Back to the library
arXiv HTML resource

Generative AI Models for Different Steps in Architectural Design: A Literature Review

Recent advances in generative artificial intelligence (AI) technologies have been significantly driven by models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and denoising diffusion probabilistic models (DDPMs). Although architects recognize the potential of generative AI in design, personal barriers often restrict their access to the latest technological developments, thereby cau…

Licence
OPEN CC-BY-4.0
Authors
Chengyuan Li, Tianyu Zhang, Xusheng Du, Ye Zhang, Haoran Xie
Published
2024-03-30 · arXiv
Language
en
Length
18594 words
Type
narrative text

Cites 186 works

inferred
Open original ↗

4 Future Research Directions

In this section, we summarize the current generative AI models of image, video, and 3D forms that could be applied to six architectural design steps in the future. According to the research discussed in this paper (Section 2), generative AI technology can be categorized into image generation models, video generation models, 3D generation models, and language generation models. This classification is derived from the outputs of various generative AI models. Since language generation models are not directly applicable in architectural design, the focus of current and future application technologies lies in image, video, and 3D generation models. Publicly available data from the Huggingface platform indicate that the number of video and 3D generation models is considerably smaller than that of image generation models. Image generation models are predominantly used by architects due to their advanced technology. As technology progresses, their application in architectural design is expected to be more effective and widespread in the future.

Currently, generative AI is limited to specific design steps and cannot yet handle the full architectural design process. In the future, technological advancements could lead to a comprehensive generative AI platform, allowing designers to flexibly select and combine modules, resulting in a more adaptable and integrated design process. Data consistency and compatibility are essential to ensure seamless information transfer between different steps in architectural generative AI platform. Leveraging the cross-step capabilities of generative AI platform can effectively link design steps by using the output of one step as the input for the next. For example, the conceptual design step can generate concept images based on the designer’s ideas by text/image-to-image models, and text/image-to-3D model generation technology can directly transform design conceptual images into initial 3D models, providing a foundation for detailed design. Floor plans and elevations designs can be automatically extracted from these 3D models. Then, AI models can generate detailed drawings that meet design specifications. In the structural and sectional design step, structural analysis and sectional drawings can also be automatically performed based on the generated 3D model, ensuring the design’s feasibility and safety. A real-time synchronization mechanism ensures that changes in one design step are instantly reflected in others, maintaining consistency and efficiency.

4.1 Future Research Directions for Architectural Concept Generation

Presently, text-to-image generation models produce creative architectural Concept images from brief descriptions or specific parameters, while image-to-image models generate images with consistent styles or features. These models enable the exploration of architectural forms beyond current human conceptions. Text/Image-to-image models such as DALL·E 3 [124], ControlNet [179], and SceneGenie [33] offer style transfer and scene control but still struggle to blend complex styles. Future improvements should focus on better transferring architectural styles, materials, and forms, enabling more flexible exploration of design concepts. Enhancing AI’s multimodal input and integrating natural language descriptions, 2D sketches, and 3D models will streamline conceptual design generation.

Generative AI-based video models open new possibilities for architectural concept generation. These models create dynamic visualizations from a single architectural image or textual description, enabling architects to refine concepts in motion. As shown in Figure 9, the first and second rows depict effect demonstration videos generated from input images using PIKA [162], where buildings undergo minor movements and scaling while maintaining consistency with the surrounding environment. DynamiCrafter [166] can generate rotating buildings, as demonstrated in the third row, where the model predicts architectural styles from different angles and ensures consistent generation. The application of these video models expands the possibilities for architectural concept generation by allowing designers to explore and present ideas dynamically.

Maintaining fine details across frames is crucial in architectural concept videos, especially for materials, textures, and structural elements. Make-A-Video [136] converts text descriptions into dynamic videos, but its ability to preserve detail and ensure smooth transitions between frames needs improvement. In the future, refining temporal convolution and attention mechanisms could enhance frame-by-frame consistency, ensuring that intricate architectural elements like textures and lighting transition throughout the video sequences. DynamiCrafter [166] adds dynamic effects like moving clouds and flowing water to static architectural images. Future advancements could incorporate physics-based algorithms to simulate realistic environmental interactions, such as wind affecting facades or light shifting throughout the day. This would significantly enhance the realism and functionality of architectural concept videos, allowing architects to test and visualize how their designs interact with the environment. Although DynamiCrafter supports high-resolution models, further enhancements could bring video quality closer to photorealism, which is crucial for presenting architectural concepts. Incorporating dynamic elements such as human activity, vehicles, and environmental changes would enrich the architectural narrative, providing a more immersive experience. PIKA [162] adjusts elements in architectural videos based on prompts while maintaining overall image integrity. In the future, it could improve in handling complex scene transformations. Architectural design often requires structural changes or material replacements, so enhancing PIKA to manage large-scale modifications would be beneficial. Using advanced prompt-to-scene transformation techniques, PIKA could allow significant design changes without compromising building integrity, leading to smoother and more natural presentations of architectural designs in video form.

Figure 9: Currently, generative models, such as PIKA [162] and DynamiCrafter [166], are capable of generating high-quality videos from images, supporting multi-angle rotation, and style transfer.

Figure 9: Currently, generative models, such as PIKA [162] and DynamiCrafter [166], are capable of generating high-quality videos from images, supporting multi-angle rotation, and style transfer.

(a) Castle

(a) Castle

(b) Abbey

(b) Abbey

(c) Chichen Itza

(c) Chichen Itza

(d) Castle

(d) Castle

(e) Pisa tower

(e) Pisa tower

(f) Opera house

(f) Opera house

(a) Castle

4.2 Future Research Directions for Architectural 3D Forms Generation

Nowadays, using architectural images or text prompts as inputs to generate 3D architectural forms can enhance modeling efficiency. As shown in Figure 10, DreamFusion [116] and Magic3D [88] facilitate the rapid creation of 3D architectural models from text descriptions. Magic123 employs a two-stage process that transforms complex real-world images into detailed 3D models , illustrated in Figure 11. GaussianEditor [17] excel in generating and refining architectural 3D models. GaussianEditor utilizes Gaussian semantic tracing and Hierarchical Gaussian Splatting for precise and intuitive architectural detail editing, as depicted in Figure 12.

DreamFusion [116] and Magic3D [88] generate 3D architectural models quickly but need improvement in high-resolution and detailed rendering. Current models need more clarity in handling complex details, textures, and fine structures. In the future, enhancing rendering capabilities will improve detail preservation, resulting in richer visual presentations, which is essential for accurately conveying architectural designs. CLIP-NeRF [161] and GaussianEditor [17] offer editing features that allow architects to adjust models using text or image prompts. In the future, real-time interaction and editing can be improved. Future technologies could enable architects to adjust specific model details in real-time, such as dynamically modifying facades or window layouts through a more intuitive interface. This would significantly accelerate the design process and improve efficiency. DreamCraft3D [145] generates personalized 3D models from text and image prompts, but personalization remains limited. Future models could enhance customization by allowing designers to input detailed parameters like room layouts, materials, and styles, generating diverse 3D forms tailored to specific needs. Supporting more complex prompts would enable a more flexible exploration of architectural forms and styles. Current models rely on single-mode inputs like text or images, but future models could support diverse inputs such as architectural floor plans and facade details for more precise 3D models. Integrating multi-modal inputs would better capture creativity while considering practical construction needs, enabling architects to make more comprehensive plans in the early design stages.

(a) Dragon: input image (left) and 3D model (right)

(a) Dragon: input image (left) and 3D model (right)

(b) Teapots: input image (left) and 3D model (right)

(b) Teapots: input image (left) and 3D model (right)

(a) Dragon: input image (left) and 3D model (right)

(a) Turn the bear into a Grizzly bear

(a) Turn the bear into a Grizzly bear

(b) Make it Autumn

(b) Make it Autumn

(a) Turn the bear into a Grizzly bear

4.3 Future Research Directions for Architectural Floor Plan Generation

The advancement of generative AI models offers innovative tools for generating and visualizing floor plans and spatial layouts through various inputs such as text and room layouts, as shown in Figure 13. Text-to-image models like StackGAN [177] and GLIDE [111] can generate architectural floor plans from text prompts, enabling quick initial design drafts. Future improvements could integrate more inputs like building regulations, functional needs, and material details, producing detailed floor plans that comply with building standards in real-time for greater precision. Image-to-image models like ControlNet [179] can control outputs using layouts, sketches, and masks, while Layout2Im [182] generates architectural floor plans based on room layouts and spatial relationships. Future improvements could enhance detail accuracy and incorporate specific design parameters like wall and window placement, resulting in more refined, practical floor plans. Allowing real-time adjustments and generating multiple design options for comparison would make these models more efficient and adaptable design tools for various projects. Video generation models like Make-A-Video [136] and DynamiCrafter [166] effectively display dynamic changes in architectural floor plans, aiding designers and clients in understanding spatial layouts. Future integration of real-time editing and interactive features could allow dynamic adjustments, such as modifying room layouts or spatial flow during video generation. This would create more intuitive and dynamic design presentations, enhancing communication efficiency in design presentations and client demonstrations.

Figure 13: Existing generative models can generate layout of rooms based on input text and can also be controlled accordingly based on input layouts, Future models will be able to generate layouts using a diverse range of information.

Figure 13: Existing generative models can generate layout of rooms based on input text and can also be controlled accordingly based on input layouts, Future models will be able to generate layouts using a diverse range of information.

4.4 Future Research Directions for Architectural Facade Generation

Currently, layout and segmentation masks represent facade information in 2D image generation. Future applications of generative AI in facade generation may integrate various data inputs—such as semantic segmentation maps, conceptual images, and textual descriptors, as shown in Figure 14. Text-to-image models like GLIGEN [83] can generate architectural facade images from textual descriptions, quickly producing low-resolution sketches with basic shapes and colors, followed by refinement to enhance details and resolution. Future improvements could incorporate more input textual information like building regulations, material information, and style requirements to generate high-resolution facades that meet design standards. Adding real-time editing features would allow designers to adjust details flexibly during the generation process. Similarly, image-to-image models like ControlNet [179] generate aesthetically appealing facades based on images, prioritizing visual appeal over strict design standards. Future enhancements could include environmental adaptability (e.g., Heatmaps representing the light, ventilation, insulation) to ensure the facades are both visually pleasing and environmentally suitable.

Figure 14: Current generative models can create facade based on input mask, Future models will be able to generate facade using a diverse range of information.

Figure 14: Current generative models can create facade based on input mask, Future models will be able to generate facade using a diverse range of information.

4.5 Future Research Directions for Architectural Structural System Generation

Text/Image-to-image models like Midjourney [189] and ControlNet [179] inspire architectural design by generating creative structural images from text, layouts, or sketches. But they often overlook building codes, material properties, and mechanical requirements for complex structures. Future developments could focus on integrating building regulations and engineering constraints into the models, ensuring that generated designs are both creative and aligned with real-world standards. Text-to-3D models like DreamFusion [116] generate 3D architectural structures, but the details often need to be more clear for complex designs. Future advancements will likely focus on improving precision and detail by refining algorithms and incorporating high-resolution 3D datasets. Integrating text-to-3D modeling with Building Information Modeling (BIM) could further enhance these models, enabling them to include not only geometric data but also technical parameters like material properties, energy consumption, and structural performance. AI-driven optimization could ensure that designs are both aesthetically pleasing and compliant with mechanical and safety standards. Future image-to-3D models like CLIP-NeRF [161], which combine text and image inputs, could enhance their ability to generate both stylized and functional outputs by incorporating users’ style preferences, functional requirements, and technical constraints. Future developments may focus on leveraging deep learning to process complex multimodal inputs better.

4.6 Future Research Directions for Architectural Section Generation

Text/Image-to-image models like GLIDE [111] and Layout2Im [182] can transform text, hand-drawn sectional sketches, and layout diagrams into detailed sectional images. Future advancements could incorporate AI-driven optimization tools to improve precision, reducing the need for manual adjustments. Supporting multimodal inputs, such as text, sketches, and technical data, would enable the generation of sectional drawings that better align with project requirements. This evolution would greatly enhance the intuitiveness and efficiency of architectural design. 3D generation models like Magic3D [88] construct 3D building models, allowing sectional images to be obtained by slicing, improving the understanding of complex spatial structures. Future advancements may integrate automated slicing and dynamic analysis tools, enabling architects to examine sections from different angles quickly. Enhanced precision will provide better insights into spatial layouts and structural details. Video generation models like DynamiCrafter [166] create dynamic visualizations of architectural sections, aiding in the comprehensive understanding of space performance. Future developments could integrate real-time simulation and environmental modeling, allowing designers to assess how building spaces perform under different conditions.