2 Current Advanced AI Technologies
This paper provides a review of AI in the creative industries, building on our previous publication in 2022 Anantrasirichai and Bull (2022). The reader is referred to that work for an introduction to AI, basic neurons, convolutional neural networks (CNNs), generative adversarial networks (GANs), recurrent neural networks (RNNs) and deep reinforcement learning (DRL). In this paper, we emphasize four key technologies that have grown in importance since 2022 that have had a significant impact on the creative industries. These are Transformers, Large language models (LLMs), Diffusion Models (DMs), and Implicit Neural Representations (INRs). It is important to note that, while these newer technologies are gaining prominence, those from previous generations remain in widespread use, often in conjunction with the newer ones. For instance, CNNs complement transformers since CNNs effectively capture local features and semantic meaning, while the attention mechanism in transformers captures global dependencies.
One important class within AI that has become dominant since our previous review comprises Foundation models (FMs). These were described by The Stanford Institute for Human-Centered Artificial Intelligence in 2021 Bommasani et al. (2021) as ‘‘any model that is trained on broad data (generally using self-supervision at scale) that can be adapted (e.g., fine-tuned) to a wide range of downstream tasks”. Foundation models have been enabled by rapid advances in AI-oriented computing power and have underpinned the emergence and success of Large Language Models, particularly following the launch of ChatGPT by OpenAI in 2022. ChatGPT has become the fastest-growing consumer software application in history[^7]. These technologies are expanded on below.
2.1 Transformers
In 2017, Google AI introduced the concept of ‘Transformer’ architectures in their publication ‘Attention Is All You Need’ Vaswani et al. (2017). This work has since been instrumental in the development and success of large language models alongside many other applications, including vision understanding Dosovitskiy et al. (2021), and multiple modality learning (e.g., Gato Reed et al. (2022)).
Before the advent of transformers, natural language processing (NLP) was performed using recurrent neural networks (RNNs), processing data sequences sequentially. In contrast, the ability of transformers to capture long-range dependencies through self-attention mechanisms that extend across all words in the sequence meant that the importance of different words could be established globally, understanding relationships regardless of their positions. This context-aware representation enables parallel processing of the entire sequence, making the transformers computationally efficient. A set of several attention layers running in parallel is called Multi-Head Attention.
The Transformer architecture, shown in Fig. 1 (a), comprises Encoder and Decoder sections, similar to many CNN-based generators. However, the encoder is now a stack of identical layers, concatenating a multi-head self-attention mechanism and a fully connected feed-forward network. The decoder is also a stack of identical layers, in which each layer has an additional sub-layer to perform multi-head attention over the output of the encoder stack.
Mathematically, the attention function is computed from inputs: query $Q$, keys $K$, and values $V$. The matrix of outputs of attention function is
$$ \text{Attention}(Q,K,V)=\text{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V, $$
where $d_{k}$ is a dimension of $K$. The term ${QK^{T}}$ is Dot-Product Attention, which yields a high similarity value when the two words are closely related. If $Q$ and $K$ are from the same sentence, Eq. 1 refers to self-attention, but if $Q$ and $K$ are from different sentences, it is referred to as cross-attention. Within the network, multi-head attention is actually employed to concurrently process attention and enable the model to collectively focus on information from distinct representation subspaces at various positions through the learnable parameters $W$s.
$$ \begin{split}\text{MultiHead}(Q,K,V)&=\text{Concat}(\text{head}_{1},...,\text{head}_{h})W^{O},\\ \text{head}_{i}&=\text{Attention}(QW_{i}^{Q},KW_{i}^{K},VW_{i}^{V}).\end{split} $$
It should be noted that attention modules are not solely used in transformers, but have also been successfully integrated into other deep learning architectures such as CNNs, used for image classification Li et al. (2022a), object detection Woo et al. (2018), and other computer vision tasks Guo et al. (2022b).
In 2020, the first successful training of a transformer encoder for image recognition was published Dosovitskiy et al. (2021), referred to as a Vision Transformer (ViT). The ViT decomposes an input image into patches, similar to words in a sentence, and processes them through multi-head attention. Additionally, a Multilayer Perceptron (MLP) is employed as the feedforward network. In later work Microsoft introduced a hierarchical division of image inputs and a shifted window approach in their Swin Transformer Liu et al. (2021). This was reported to outperform ViT by 2.4% in ImageNet-22K classification (21,841 different categories). Its version 2 Liu et al. (2022c) applied a cosine function in the attention module. enabling the scaling of capacity and resolution. More detail on transformer-based object detection is discussed in Section 3.4.2. To date, Swin Transformers have been widely adopted in a range of applications including image restoration Fan et al. (2022). Comprehensive surveys on the use of transformers for image and video processing can be found in Khan et al. (2022) and Selva et al. (2023), respectively.
Transformers have been widely used and offer better performance across many tasks. One reason for this widespread adoption has been the availability of open-source Transformer libraries such as Hugging Face[^8], a platform that assists developers to build applications for tasks including computer vision, NLP, audio, tabular data, multimodal tasks, and reinforcement learning. The platform also provides access to model zoo[^9] pretrained networks and datasets.
In recent years, state space models Gu and Dao (2024); Zhu et al. (2024b), commonly known as ‘Mamba’ have emerged. These are a linear variant of Transformers distinguished by their linear complexity in attention modeling. They are acknowledged to offer an equivalent or better performance than traditional Transformers, while demanding fewer computational resources and less memory.

Figure 1: Generative AI. (a) Transformer architecture Vaswani et al. (2017). (b) The top row represents the diffusion process and the bottom row represents the generation process of the new image Yang et al. (2023f). (c) Latent Diffusion Models (LDM) Rombach et al. (2022).
2.2 Large language models
LLMs are based on transformer models using self-attention mechanisms as their core modules. Training comprises two steps: i) pre-training with large amounts of unlabelled text data in an unsupervised manner to learn word meanings and relationships, and ii) task adaptation through either fine-tuning or prompt-tuning. This pre-training approach is why OpenAI refers to their model as a Generative Pre-trained Transformer (GPT).
Fine-tuning involves training the model on new datasets, which must be large enough to ensure generalization to new tasks. Prompt-tuning and prompt engineering are emerging disciplines focused on developing and optimizing prompts for efficient model use. Prompts guide how AI models interpret and respond to user queries. Prompt engineering structures text or phrasing to steer the model toward desired outputs, relying heavily on experimentation and understanding model behavior. Prompt-tuning instead trains a small set of parameters before using the LLM, requiring relatively little new data. This approach converts text inputs into task-specific virtual tokens while keeping the pre-trained model unchanged Lester et al. (2021). Its main drawback is reduced interpretability, though the paradigm has extended to other domains such as visual prompt tuning Jia et al. (2022). For a comprehensive survey of LLMs, see Zhao et al. (2023b).
To date, there are many LLM platforms as shown in Fig. 2. Fig. 2(a) shows their timeline. Many surveys and evaluations of LLMs are also available Zhao et al. (2023b); Chang et al. (2024); Yao et al. (2024). These include FLASK (Fine-grained Language Model Evaluation based on Alignment Skill Sets) Ye et al. (2024) which evaluates LLMs based on 12 fine-grained skills for comprehensive language model evaluation: logical correctness, logical robustness, logical efficiency, factuality, commonsense understanding, comprehension, insightfulness, completeness, metacognition, conciseness, readability, and harmlessness. Evaluation results from FLASK are shown in Fig. 2 (b).

Figure 2: (a) Timeline of large language models. (b) Performance comparison evaluated by FLASK Ye et al. (2024).
A clear divide exists between open-source and proprietary LLMs. Open-source models such as LLaMA 3, Mistral, and Falcon promote transparency, reproducibility, and creative customization, allowing researchers to fine-tune or deploy them locally with improved data control. Proprietary systems like GPT-4, Gemini 1.5, and Claude 3, meanwhile, deliver stronger multimodal integration, safety alignment, and reliability—benefiting from access to large, private datasets and compute resources. Yet their closed architectures limit interpretability and independent evaluation. Together, these ecosystems form a complementary landscape balancing openness and innovation against performance and safety, shaping accessibility and creative autonomy in AI-driven practice.
2.3 Diffusion Models
A generative model, in the context of AI, exploits machine learning to learn a probability distribution of the training data to generate new data samples. The very first models were based on Autoencoders (AEs) that learn to encode input data into a lower-dimensional representation (latent space) and then decode it back to its original form. A specific type of AE, a variational autoencoders (VAE) Kingma and Welling (2014), learns the latent space as statistical parameters of probabilistic distributions, leading to significant improvement of the generated results. Concurrently, Goodfellow et al. Goodfellow et al. (2014) introduced an alternative architecture known as a Generative Adversarial Network (GAN). GANs comprise two competing AI modules: a generator, which creates a sample, and a discriminator, which determines whether the received sample is real or generated. When comparing VAEs to GANs, VAEs exhibit greater stability during training, whereas GANs excel at producing realistic images. More details about AEs and GANs for creative technologies can be found in our previous review Anantrasirichai and Bull (2022).
An important factor in driving the rapid growth of generative AI has been the development of diffusion probabilistic models (referred to as diffusion models (DMs)). The first DM was introduced in 2015 by Sohl-Dickstein et al. Sohl-Dickstein et al. (2015), using Nonequilibrium Thermodynamics. However, it took a further 5 years for DMs to generate desirable results: the era of DMs began with Denoising Diffusion Probabilistic Models (DDPMs) proposed by Ho et al. Ho et al. (2020) in 2020 and Score-based diffusion models proposed by Song et al. Song et al. (2021) in 2021. These involve a simplified process using a denoising autoencoder to approximate Bayesian inference. In brief, the models leverage a diffusion process to learn a probability distribution of the input data. As the name suggests, the data is diffused by gradually adding noise at each iteration step as shown in Fig. 1 (b). A deep neural network (DNN) is then trained to remove this noise, called the denoising process or reverse process. Consequently, the trained model uses random noise to generate data with characteristics similar to those of the training samples. Comparing to GANs, DMs provide higher diversity samples Dhariwal and Nichol (2021) and a training process that is much more stable and does not suffer from mode collapse. DMs are however computationally intensive and require longer training times compared to GANs. The complexity can be significantly reduced by training the DMs in latent space. Latent Diffusion Models (LDM) Rombach et al. (2022) use pretrained networks to convert images to feature maps, and perform training on a low-dimensional space. The diagram of LDM is shown in Fig. 1 (c).
Generating a synthesized sample at random might not be particularly useful, especially for creative industry applications. Therefore, conditional diffusion models have been proposed, supporting a wide range of applications such as text-to-sound, text-to-images, and image-to-videos. For DMs, the conditional distributions are modeled using a conditional denoising autoencoder. Classifier guidance was introduced in Dhariwal and Nichol (2021) to improve the generation of images of a desired class. For example, when we provide the model with information, such as ‘a flower’, the DM will synthesize a variety of flower images, as the word ‘flower’ guides the model toward the latent distribution that is formed by various images of flowers. The work in Choi et al. (2021) simply refines the latent space of well-trained unconditional DDPM so that the higher-level semantics of the synthetic samples are similar to the reference (conditioning). The LDM Rombach et al. (2022) offers more flexible conditional image generators by adding cross-attention layers (referred to Transformers in Section 2.1) to the denoising autoencoder. A survey on the methods and applications of DMs prior to 2024 can be found in Cao et al. (2024).
2.4 Implicit Neural Representations
Implicit Neural Representations (INR), also called neural fields, neural implicits or coordinate-based neural networks, represent input content implicitly through learned functions $F$, as shown in Eq. 3. They can be considered as fields $x$ (represented by a scalar, vector, or a tensor with a value, such as a magnetic field in physics) that are fully or partially parameterized by a neural network $\Phi$, typically an MLP Xie et al. (2022).
$$ F(x,\Phi,\nabla_{x}\Phi,\nabla_{x}^{2}\Phi,\dots)=0,\quad\Phi:x\mapsto\Phi(x). $$
Although this concept appears complex, the process is actually very straightforward. For example, in the case of an image, the coordinates of each pixel $(x,y)$ contain color information $(r,g,b)$. The INR inputs $(x,y)$ to the MLP and learns to provide the output $(r,g,b)$. The weights and biases of the MLP now represent such an image. Usually, the number of parameters of the MLP is smaller than the total number of pixels multiplied by 3, accounting for the 3 color channels. Hence, one of its emerging applications is in data compression Kwan et al. (2024a). Moreover, the INR can handle complex and high-dimensional data efficiently, attracting attention for visual computing applications such as 3D scene reconstruction.
Traditional MLPs employ ReLU (rectified linear unit) for non-linear activation due to its simplicity. However, Sitzmann et al. Sitzmann et al. (2020b) demonstrated that periodic functions, such as sinusoids, are more suitable for representing complex natural signals, offering a better fit to the first- and second-order derivatives of the signals. However, this activation can cause ringing artifacts. Saragadam et al. instead proposed using complex Gabor wavelets Saragadam et al. (2023), which learn to represent high frequencies better and simultaneously are robust to noise.
One of the fastest-growing areas that exploits INRs is Neural Radiance Fields (NeRF), evidenced by 57 papers presented at CVPR, the largest annual conference in computer vision, in 2022 growing to 175 papers in 2023[^10], before dropping to 71 in 2024, largely due to competition from 3D Gaussian Splatting[^11]. First introduced in 2020 by Mildenhall et al. Mildenhall et al. (2020), NeRF is a form of neural rendering, a subset of generative AI, that generates novel views of a scene based on a partial set of 2D images. It achieves this by learning a mapping from 3D spatial coordinates and view directions $(x,y,z,\theta,\phi)$ to colors and density $(r,g,b,\sigma)$. This implicit representation allows NeRF to handle complex scenes with varying geometry and appearance, resulting in highly realistic renderings that include accurate lighting, shadows, and reflections. More detail can be found in Section 3.5.2.
| Application | Technology | |||
|---|---|---|---|---|
| Trans./Attn.1 | Diffusion model2 | INR | ||
| Creation | Text | Vaswani et al. (2017); Wang et al. (2022c); OpenAI et al. (2023); Wei et al. (2024) | ||
| audio/music | Alayrac et al. (2022); Huang et al. (2022b) | Li et al. (2022c); Yang et al. (2023c); Evans et al. (2025) | ||
| Image | Alayrac et al. (2022); Esser et al. (2024) | Rombach et al. (2022); Brooks et al. (2023); Gal et al. (2023); Gandikota et al. (2024); Lian et al. (2024); Esser et al. (2024); Ren et al. (2024); Feng et al. (2025b); Liu et al. (2025b) | ||
| Animation/video | Hong et al. (2023); Villegas et al. (2023); Azadi et al. (2023); Yu et al. (2023); Liu et al. (2023c); Wang et al. (2024b); Xu et al. (2024a); Corona et al. (2024); Gupta et al. (2024); Hu (2024); Zhu et al. (2024c) | Singer et al. (2023); Molad et al. (2023); Wang et al. (2023a); Wu et al. (2023d); Liu et al. (2023c); Gupta et al. (2024); Zhu et al. (2024c); Wang et al. (2025c); Wu et al. (2025a) | ||
| 3D/AR/VR | Yang et al. (2024d) | Xu et al. (2023a); Melas-Kyriazi et al. (2023); Qian et al. (2024); Tang et al. (2024a) | Tang et al. (2024a); Ren et al. (2023); Zhao et al. (2024d) | |
| Information Analysis | Text categorization | Sun et al. (2023); Shi et al. (2023); Hou et al. (2023); Ai et al. (2025) | ||
| Film analysis | Mao et al. (2023); Krugmann and Hartmann (2024); Hartmann et al. (2023) | |||
| Content retrieval | Metzler et al. (2021); Yan et al. (2023); Lu et al. (2023); Rajput et al. (2023); Li et al. (2024d); Li et al. (2024e) | Jin et al. (2023) | ||
| Intelligent assistants | King et al. (2024) | |||
| Content | Enhancement | Xu et al. (2022); Liang et al. (2022); Wang et al. (2023b); Lin et al. (2024c); Youk et al. (2024); Liang et al. (2024) | HOU et al. (2023); Yi et al. (2023); Jiang et al. (2023a); Lin et al. (2024b) | Yang et al. (2023g) |
| Enhancement | Style transfer | Deng et al. (2022); Moon et al. (2023); Chung et al. (2024) | Zhang et al. (2023d); Chai et al. (2023) | Moon et al. (2023); Kim et al. (2024b) |
| and Post | Super-resolution | Liang et al. (2021); Lu et al. (2022); Liu et al. (2022a); Chen et al. (2023b); Li et al. (2023c); Kang et al. (2023); Liang et al. (2024); Xu et al. (2024b); Wang et al. (2025b) | Saharia et al. (2023); Moliner et al. (2023); Gao et al. (2023); Cao et al. (2025); Wang et al. (2025b) | Chen et al. (2021b); Saharia et al. (2023); Fei et al. (2023); Gao et al. (2023); Yin et al. (2023) |
| Production | Restoration | Wang et al. (2022a); Zamir et al. (2022); Li et al. (2023c); Yang et al. (2023d); Liang et al. (2024); Morris et al. (2025); Liang et al. (2021); Fan et al. (2022); Yu et al. (2022); Wang et al. (2023c); Song et al. (2023); Xu et al. (2023b); Mao et al. (2022); Zhang et al. (2024c); Zou and Anantrasirichai (2024); Fang et al. (2025); Yue et al. (2025); Jin et al. (2025); Yue et al. (2025); Shi et al. (2025) | Jiang et al. (2023a); Fei et al. (2023); Yang et al. (2023a); Nair et al. (2023); Jaiswal et al. (2023); Cao et al. (2025); Feng et al. (2025a) | Jiang et al. (2023b) |
| Inpainting | Li et al. (2022b); Liu et al. (2022b); Ren et al. (2022); Zhou et al. (2023); Huang et al. (2024a) | Moliner et al. (2023); Fei et al. (2023) | ||
| Fusion | Ma et al. (2022); Rao et al. (2023); Liu et al. (2023b); Li and Wu (2024) | Zhao et al. (2023c) | ||
| Editing/VFX | Shi et al. (2024b) | Shi et al. (2024b); Guo et al. (2024b) | ||
| Information | Segmentation | Cheng et al. (2022); Kirillov et al. (2023); Ke et al. (2023); Wang et al. (2023c); Wang et al. (2023d); Zou et al. (2023a); Oquab et al. (2024); Ravi et al. (2024); Zhang et al. (2025b) | Wu et al. (2023g); Xu et al. (2023c); Gu et al. (2024) | Gong et al. (2023); Cen et al. (2023) |
| Extraction | Recognition | Carion et al. (2020); Dosovitskiy et al. (2021); Zhu et al. (2021); Liu et al. (2021); Neimark et al. (2021); Liu et al. (2022c); Huang et al. (2022a); Oquab et al. (2024); Zhao et al. (2024a); Im et al. (2025); Tian et al. (2025) | Li et al. (2023a); Chen et al. (2023a); Zhang et al. (2025a); Wu et al. (2025b) | |
| and | Tracking | Meinhardt et al. (2022); Zeng et al. (2022); Cui et al. (2022); Mayer et al. (2022); Yang et al. (2023e); Chen et al. (2023c); Zhang et al. (2023c); Yi and Anantrasirichai (2024); Kang et al. (2025) | Luo et al. (2024); Xie et al. (2024); Zhang et al. (2024b) | Jung et al. (2023) |
| Understanding | 3D Reconstruction | Wang et al. (2021); Zhang et al. (2023b); Chen et al. (2023e); Yang et al. (2024b); Oquab et al. (2024); Yang et al. (2024c); Liu et al. (2025a) | Barron et al. (2022); Ji et al. (2023); Wynn and Turmukhambetov (2023); Ke et al. (2024) | Mildenhall et al. (2020); Pumarola et al. (2020); Müller et al. (2022); Barron et al. (2022); Mildenhall et al. (2022); Fang et al. (2022); Guo et al. (2022a); Hu et al. (2022); Liu et al. (2024); Azzarelli et al. (2023); Zhan et al. (2024); Tang et al. (2024b); Liu et al. (2025a),Sara Fridovich-Keil and Giacomo Meanti et al. (2023); Kerbl et al. (2023); Wu et al. (2024a); Yu et al. (2024); Huang et al. (2024b); Wang et al. (2025a); Junkawitsch et al. (2025); Kong et al. (2025)† |
| Compression | Image∗ | Zhu et al. (2022); Zou et al. (2022); Liu et al. (2023a) | Careil et al. (2023); Yang and Mandt (2023); Hoogeboom et al. (2023); Ghouse et al. (2023) | Sitzmann et al. (2020a); Dupont et al. (2021); Dupont et al. (2022); Strümpler et al. (2022) |
| Video | Xiang et al. (2022); Mentzer et al. (2022) | Li et al. (2024a) | Chen et al. (2021a); Bai et al. (2023); Kwan et al. (2024a); Kim et al. (2024a); Leguay et al. (2024); Kwan et al. (2024b); Gao et al. (2024); Ruan et al. (2024); Kwan et al. (2024c) | |
| Audio∗ | ||||
| Quality | Image∗ | Cheon et al. (2021); Golestaneh et al. (2022); Shi et al. (2024a) | ||
| Assessment | Video∗ | Wu et al. (2022); Feng et al. (2024a); Wu et al. (2023c); He et al. (2024); Peng et al. (2024a) | ||
| 1 Trans./Attn. include transformers, mamba and CNN-based architectures that use attention module. | ||||
| 2 Some diffusion models employ the transformer in their denoising autoencoders. |
† These methods are based on explicit neural representations. ∗ It is noted that for some compression and quality assessment tasks, there are other dominant network architectures in existing works. For example, LLMs have been used for image and audio compression, and visual quality assessment. Many neural audio codecs are also based on VQ-VAE models.
Table 1: Creative applications and corresponding AI-based methods mentioned in this paper