OER·harvester

← Back to the library
arXiv HTML resource

Generative AI Models for Different Steps in Architectural Design: A Literature Review

Recent advances in generative artificial intelligence (AI) technologies have been significantly driven by models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and denoising diffusion probabilistic models (DDPMs). Although architects recognize the potential of generative AI in design, personal barriers often restrict their access to the latest technological developments, thereby cau…

Licence
OPEN CC-BY-4.0
Authors
Chengyuan Li, Tianyu Zhang, Xusheng Du, Ye Zhang, Haoran Xie
Published
2024-03-30 · arXiv
Language
en
Length
18594 words
Type
narrative text

Cites 186 works

inferred
Open original ↗

2 Generative AI Models

Generative AI models are rapidly advancing with the continual emergence of new methods, as illustrated in Figure 4. Deep learning-based approaches, including VAEs, GANs, DMs, flow models, and transformer models have significantly enhanced image and text generation techniques. VAEs were pioneering generative models, using an encoder–decoder architecture with probabilistic graphical models to learn latent representations for image generation [42]. GANs marked a milestone by employing a generator and discriminator in an adversarial training process, driving the generator to create images that increasingly resembled real data [69, 171]. DMs have recently emerged as revolutionary, demonstrating outstanding image generation quality [128, 125]. Flow models achieve precise probability estimation and efficient sample generation through reversible transformations [194]. In addition, transformer models, with their self-attention mechanism, have significantly improved sequential data processing and form the basis of many generative models [193]. In this paper, we focus on the widely used GANs in architecture and the state-of-the-art Diffusion Models.

Large language models (LLMs)—like the generative pre-trained transformer (GPT) [123] series and bidirectional encoder representations from transformers (BERT) [27], based on transformer architecture—are trained on vast text data, thereby revealing strong language understanding and generation capabilities that can be applied in tasks such as text generation, answering questions, and machine translation. Large vision models use deep learning techniques for image data processing and generation, typically based on convolutional neural networks (CNNs) and transformer architectures, excelling in image classification, object detection, and image generation. These large models have gained significant attention in academia and practical applications, pushing the boundaries of image and text generation technologies and highlighting the potential and broad application of deep learning in generative models.

2.1 Generative Adversarial Networks

GANs [40] comprise a generator and a discriminator, as illustrated in Figure 5. The generator produces synthetic samples from random noise, while the discriminator evaluates whether these synthetic samples resemble real samples. The objective is for the generator to create samples indistinguishable from real data according to the discriminator. This adversarial framework enables the model to maintain a dynamic equilibrium between generation and discrimination, thereby driving the learning and optimization of the entire system. Despite the many advantages of GANs, they also face challenges, such as mode collapse during training. This occurs when the generator overly focuses on a few patterns or samples, thereby neglecting the broader diversity of the dataset and resulting in outputs that are limited, repetitive, or lack diversity. This issue impacts the robustness and generalization of the model.

Figure 4: Types of generative artificial intelligence models.

Figure 5: The framework of Generative Adversarial Networks (GAN), Variational Autoencoders (VAE), Diffusion Models (DM) and Latent Diffusion Models (LDM). Where $z$ is a compressed low-dimensional representation of the input.

Figure 5: The framework of Generative Adversarial Networks (GAN), Variational Autoencoders (VAE), Diffusion Models (DM) and Latent Diffusion Models (LDM). Where $z$ is a compressed low-dimensional representation of the input.

Conditional GAN:

To address the limited controllability of traditional GAN models, Conditional GAN (CGAN) [104] was introduced. CGAN controls the image generation process by incorporating conditional informa- tion—such as text, labels, or hand-drawn sketches—to produce images that meet specific criteria. In CGAN, the generator receives both random noise and conditional information, thereby enabling it to produce images that align more closely with the given conditions. This approach enhances the precision of the generated results. Additionally, variants such as pix2pix [63] and StyleGAN [69] have been developed to further refine and extend the capabilities of conditional image generation.

2.2 Diffusion Models

In image generation, DMs outperform GANs and VAEs [138, 55]. Most DMs that are currently used are based on denoising diffusion probabilistic model (DDPM) [55], which simplifies the DM through variational inference. As shown in Figure 5, DMs include both forward diffusion and reverse denoising (inference) processes. The forward process follows the concept of a Markov chain and turns the input image into Gaussian noise. Given a data sample $x_{0}$, the Gaussian noise is progressively increased to the data sample during $T$ steps in the forward process, producing noisy samples $x_{t}$. As the timestep increases, the distinguishable features of $x_{0}$ gradually diminish. Eventually When $T$ tends to infinity, $x_{T}$ is equivalent to a Gaussian distribution with isotropic covariance. In addition, the inference process can be understood as a sequence of denoising autoencoders with the same weights (autoencoder is typically implemented as U-Net [129], which are trained to forecast denoised images of their corresponding inputs, $x_{t}$.

Latent Diffusion Model:

Different from DDPM, LDM [128] do not directly operate on the images but operate in the latent space called perceptual compression. LDMs reduce the dimensionality of the data by projecting it into an efficient, low-dimensional latent space in which imperceptible high-frequency details are abstracted away. The framework of an LDM is illustrated in Figure 5. After the image is compressed by the encoder to latent representation, the diffusion process is performed in the latent representation space. Finally, LDM infers the data sample from the noise and the decoder restores the data to the original pixel space and obtains the generated images. To accelerate the generation speed, the latent consistency model (LCM) [95] was proposed to optimize the step of denoising inference.

2.3 Foundation Models

In computer science, foundation models, also called large-scale models use deep learning models with numerous parameters and intricate structures, particularly in uageguage processing and computer vision tasks. These models demand substantial computational resources for training but exhibit exceptional performance across diverse tasks. The evolution from basic neural networks to sophisticated DMs, as depicted in Figure 6, illustrates the continuous quest for more robust and adaptable AI systems.

Figure 6: The evolution of prominent large-scale models in computer science.

2.3.1 Large Language Models (LLM)

The transformer model has achieved remarkable success in natural language processing (NLP) which consists of several components: encoder, decoder, positional encoding, and the final linear and SoftMax layers. Both the encoder and decoder comprise multiple identical layers. Each layer contains several components of attention layers and feedforward network layers. Additionally, positional encoding is used to inject positional information into the text embeddings, thereby indicating the position of words within the sequence. Notably, transformer has paved the way for two prominent transformer models: BERT [27] and GPT [123]. The main difference is that BERT is based on a bidirectional pre-training language model and fine-tuning, while GPT is based on an autoregressive pre-training language model and prompting.

GPT aims to pre-train models using large-scale unsupervised learning to facilitate the generation and understanding of natural language. The training process involves two primary stages: Initially, a language model is trained in an unsupervised manner on extensive corpora without task-specific labels or annotations. Subsequently, supervised fine-tuning occurs during the second stage, catering to specific application domains and tasks. In addition, BERT has emerged as a breakthrough approach, achieving state-of-the-art performance across diverse language tasks. BERT’s training methodology comprises two key stages: pre-training and fine-tuning. Pre-training involves the utilization of extensive text corpora to train the language model. The primary objective of pre-training is to endow the BERT model with robust language understanding capabilities, thereby enabling it to effectively tackle various NLP tasks. Subsequently, fine-tuning utilizes the pre-trained BERT model in conjunction with smaller labeled datasets to refine the model parameters. This process facilitates the customization of the model to specific tasks, thereby enhancing its suitability and performance for targeted applications, such as sentiment analysis and text classification.

In recent years, LLMs have witnessed rapid and explosive growth. Basic language models refer to models that are only pre-trained on large-scale text corpora, without any fine-tuning. Examples of such models include the language model for dialogue applications (LaMDA) [151] and OpenAI’s GPT-3 [8].

2.3.2 Large Vision Models

In computer vision, pretrained vision-language models such as contrastive language-image pre-training (CLIP) [122] have demonstrated powerful zero-shot generalization performance across various downstream visual tasks. These models are typically trained on hundreds of millions to billions of image–text pairs collected from the web. In addition, certain research efforts also focus on large-scale base models conditioned on visual input prompts. For example, the segment anything model (SAM) [78] can perform category-agnostic segmentation from given images and visual prompts (such as boxes, points, or masks).

The current generative models based on DMs present unprecedented creative and understanding capabilities. Stable diffusion [128] uses the CLIP [122] text encoder and can adjust the model through text prompts. Its diffusion process begins with random noise and gradually denoising until a complete data sample is generated. DALLE-3 [124] utilizes DMs with massive data to generate amazing results. Midjourney excels at adapting to actual artistic styles to create images with any combination of effects the user desires.

Figure 7: Some examples of generated results from image generation and video generation models.

Figure 7: Some examples of generated results from image generation and video generation models.

2.4 Applications of 2D Generative AI

In this section, we introduce widely used applications of generative AI, including image generation (Section 2.4.1) and video generation (Section 2.4.2). Furthermore, we present results from presented models in Figure 7 as illustrative references.

2.4.1 Image Generation

Text-to-image:

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision, StackGAN [177] proposed a two-stage model to solve this issue. In the first stage, StackGAN generates the primitive shape and colors of the object based on the given text description, thus yielding initial low-resolution images. In the second stage, StackGAN takes the low-resolution result and text prompts as inputs and generates high-resolution images with photo- realistic details. It can rectify defects in the results of the first stage and add exhaustive details during the refinement process. Guided language to image diffusion for generation and editing (GLIDE) [111] extends the core concepts of DMs by adding additional text information to enhance the training process, ultimately generating text- conditioned images. Based on these foundations—utilizing DMs and extensive data—with the release of LDM [128], stable diffusion based on LDM has also sprung up. These works cover areas such as image editing and more powerful 3D generation, further advancing image generation and making it closer to human needs.

Image-to-image:

Image domain refers to a specific category or style of images characterized by distinct visual attributes, such as color, texture, or semantic content. Image-to-image translation can convert the content in an image from one image domain to another with cross-domain conversion between images. The objective of sketch-to-image generation is to ensure that the generated image maintains consistency in both context and appearance with the provided hand-drawn sketch. Pix2Pix [63] stands out as a classic GAN model capable of handling diverse image translation tasks, including the transformation of sketches into fully realized images. In addition, SketchyGAN [16] focuses on the sketch- to-image generation task and aims to achieve more diversity and realism. Currently, ControlNet [179] can control DMs by adding extra conditions, such as sketch, layout and masks. The sketch-to-image generation tasks are applied in both photo-realistic and anime-cartoon styles [115, 62]. The layout typically encompasses details such as the position, size, and relative relationships of individual objects. Layout2Im [182] is designed to take a coarse spatial layout, consisting of bounding boxes and object categories, for generating a set of realistic images. These images accurately depict the specified objects in their intended locations. To enhance the global attention in context, [51] introduced the context feature conversion module to ensure that the generated feature encoding for objects remains aware of other coexisting objects in the scene. With regard to DMs, GLIGEN [83] facilitates grounded text-to-image generation using prompts and bounding boxes as condition inputs in open worlds.

2.4.2 Video Generation

Since text prompts only generate some discrete tokens, text-to-video generation is more difficult than tasks such as image retrieval and image captioning. The video diffusion model [56] is the first to use a DM for video generation tasks. The video DM proposes 3D UNet, which can be applied on variable sequence lengths. Thus, it can be jointly trained on video and image modeling goals, thereby making it suitable for video generation tasks. Additionally, Make-A-Video [136] is based on the pre-trained text-to-image model and adds one-dimensional convolution and attention layers in the time dimension to transform it into a text2video model. By learning the connection between text and vision through the T2I model, the single-modal video data is utilized to learn the generation of temporal dynamic content. Furthermore, the consistency and controllability of video generation models have also garnered increased attention from researchers. PIKA [162] has been proposed to support dynamic transformations of elements in a scene based on prompts, without causing the overall image to collapse. DynamiCrafter [166] utilizes pre-trained video diffusion priors to add animation effects to static images based on textual prompts. This tool supports high-resolution models and, thus, provides better dynamic effects, higher resolution, and stronger consistency. Recently, Kling has brought impressive video generation quality and effects, excelling in aspects such as detailed camera work, lighting variations, adherence to physics, and aesthetic appeal.

2.5 3D Generative Models

In addition to 2D images, 3D models have a wide range of applications in architecture. In this section, we introduce various methods and representations used in computer graphics and vision to generate 3D models. We also discuss advances in text-to-3D and image-to-3D modeling techniques.

2.5.1 3D Shape Representation

Representation in 3D visual problems can generally be divided into four categories: voxel-based, point cloud-based, mesh- based, and implicit representation-based. As shown in Figure 8(a), the voxel format describes a 3D object as a matrix of volume occupancy, where the size of the matrix is fixed. Wu et al. [165] adopted voxel representation in the generation of 3D shapes. Voxel format requires high resolution to describe fine-grained details; thus, as the shape resolution increases, the computational cost also explodes. The reconstruction results of voxel-based research are limited in resolution and do not provide topological guarantees or represent sharp features. As shown in Figure 8(b), point clouds are a lightweight 3D representation composed of $(x,y,z)$ coordinate values. Point clouds are a natural means to represent shapes. PointNet [118] extracts global shape features using the max-set operations and it is used widely as an encoder for point-based generative networks [158]. However, point clouds do not represent topology and are unsuitable for generating watertight surfaces. Meshes are widely used and constructed from vertices and faces (Figure 8(c)). [163] deformed a pre-defined template to restrict a fixed topology using graph convolution. Recently, meshes have been used to represent shapes in deep learning techniques [106]. Although meshes are more suitable for describing the topological structure of objects, they usually require advanced preprocessing steps.

(a) Voxel

(a) Voxel

(b) Point

(b) Point

(c) Mesh

(c) Mesh

(d) Implicit

(d) Implicit

(a) Voxel

2.5.2 Implicit 3D Shape Representation

In the field of three-dimensional shape modeling, implicit functions are commonly represented in three ways: occupancy field, signed distance function (SDF), or unsigned distance function (UDF), and the recently emerging neural radiance fields (NeRF). As shown in Figure 8(d), implicit representation refers to describing a surface with a zero-crossing point of a volume function $\psi:R^{3}\to R$, whose value can be adjusted. Representing a 3D shape as a set of level sets of a deep network, mapping 3D coordinates to a signed distance function [114] or occupancy field [100]. Implicit representation can create a lightweight, continuous shape representation with no resolution limits.

Occupancy field is one of the implicit function methods based on deep learning [100]. Occupancy field assigns binary values to each point in three-dimensional space, determining whether the point is occupied by an object. This approach utilizes neural networks to learn the representation of occupancy fields, thereby facilitating highly detailed three-dimensional reconstruction. The advantage of occupancy field lies in its dynamic modeling of object occupancy in scenes, thus making it suitable for handling complex three-dimensional environments. Building upon occupancy field, the signed distance function (SDF) has become a crucial direction in implicit function representation within deep learning. SDF assigns a signed distance value to each point, indicating the shortest distance from the point to the object’s surface. Positive values signify points outside the object, while negative values indicate points inside the object.As shown in Figure . DeepSDF [114] provides an end-to-end approach for continuous SDF learning, thereby enabling precise modeling of irregular shapes and local geometry.

Neural radiance fields (NeRF) [102] have revolutionized the field of computer vision and graphics by introducing a novel approach to scene representation. at the heart of NeRF lies the concept of representing a scene as a continuous function capturing radiance information at every point. NeRF introduces an implicit representation, enabling the encoding of detailed and continuous volumetric information. This allows for high-fidelity recon- struction and rendering of scenes with fine-scale structures, surpassing the limitations of explicit representations. Recently, 3D Gaussian Splatting [70] was introduced by projecting 3D information onto a 2D domain using Gaussian kernels and achieved better performance than NeRF.

2.5.3 3D Model Generation

Text-to-3D:

Recent advancements in text-to-3D synthesis have demonstrated remarkable progress, with researchers em- ploying various sophisticated strategies to bridge the gap between natural language descriptions and the creation of detailed 3D content. The pioneering work DreamFusion [116] harnesses a pre-trained 2D text-to-image DM to generate 3D models without large-scale labeled 3D datasets or specialized denoising architectures. Magic3D [88] improves upon DreamFusion’s [116] limitations by implementing a two-stage coarse-to-fine approach, accelerating the optimization process through a sparse 3D representation before refining it into high-resolution textured meshes via a differentiable renderer.

Image-to-3D:

Recent 3D reconstruction techniques specifically focus on generating and reconstructing three-dimensional objects and scenes from a single or few images. NeRF [101] represents a state-of-the-art technique in which complex scene representations are modeled as continuous neural radiance fields optimized with sparse input views. Leveraging the joint language-image embedding space of the CLIP model, CLIP-NeRF [161] proposes a unified framework that allows manipulating NeRF in a user-friendly manner, using either a short text prompt or an exemplar image. DreamCraft3D [145] introduces a hierarchical process for 3D content creation that employs bootstrapped score distillation sampling from a view-dependent DMs. This two-step method refines textures through personalized DMs trained on augmented scene renderings, thereby delivering high-fidelity, coherent 3D objects. In contrast, Magic123 [119] offers a two-stage solution for generating high-quality textured 3D meshes from unposed wild images. It optimizes a neural radiance field for coarse geometry and fine-tunes details using differentiable mesh representations guided by both 2D and 3D diffusion priors.