OER·harvester

← Back to the library
arXiv HTML resource

Advances in Artificial Intelligence: A Review for the Creative Industries

Artificial intelligence (AI) has undergone transformative advances since 2022, particularly through generative AI, large language models (LLMs), and diffusion models, fundamentally reshaping the creative industries. However, existing reviews have not comprehensively addressed these recent breakthroughs and their integrated impact across the creative production pipeline. This paper addresses this gap by providing a s…

Licence
OPEN CC-BY-4.0
Authors
Nantheera Anantrasirichai, Fan Zhang, David Bull
Published
2025-01-06 · arXiv
Language
en
Length
42227 words
Type
narrative text

Cites 155 works

inferred
Open original ↗

3 Advanced AI for the creative industries

Similarly to our previous (2021) review of AI for the creative industries Anantrasirichai and Bull (2022), Table 1 categorizes applications and corresponding AI-based solutions. These areas are explored in more detail below.

3.1 Content creation

Content creation is a fundamental activity of artists and designers and the term ‘AI art’ refers to artforms created with the assistance of an AI algorithm or entirely by an AI system. This can refer to various digital forms including images, texts, audio, and videos. The roots of AI art can be traced back to the 20th century, exemplified by AARON, a computer program initiated in 1972 to autonomously produce paintings and drawings Shapiro and Eckroth (1987). The practicality of AI art has been enhanced with advancements in deep learning, particularly GANs from 2014 and, more recently, transformers, DMs and INRs.

3.1.1 Text generation, script and journalism

In the era of LLMs, AI writing tools have been widely used to assist various writing tasks, including generation of written articles, blog posts, essays, and reports. These tools go beyond mere grammar and spelling checks; they boast advancements enabling them to analyze the style and tone of written material, adding images, videos and tables, offering suggestions to enhance clarity, coherence, and overall readability Ippolito et al. (2022). Moreover, AI tools extend their utility beyond content generation by automating tasks like keyword generation, meta tags, and descriptions, thereby increasing search rankings using search engine optimization (SEO). Additionally, they support the process of publishing across multiple online platforms. Transformers have been used to generate image captions by combining information from the images with a word prefix or questions Wang et al. (2022c).

AI script generators serve as beneficial aids for writers, filmmakers, and game developers, offering inspiration, idea generation, and assistance in crafting entire scripts Jeary and Gajjar (2024); Azzarelli et al. (2025). Human-AI brainstorming is helpful and saves time Guo et al. (2024a). Presently, there are numerous software and websites providing both free and paid script generation services. However, many of these tools are still constrained when it comes to longform creative writing. Dramatron, developed by Google Mirowski et al. (2023), introduces hierarchical language generation, enabling the creation of cohesive scripts and screenplays spanning long ranges. This includes elements such as titles, characters, story beats, location descriptions, and dialogue.

As discussed earlier, chatbots are now powered by LLMs, effectively simulating human conversation. These fundamental LLMs are specialized for specific tasks. For instance, journalist AI and blog AI writers[^12] generate content with layouts suitable for print or online publication. Additionally, AI tools exist that are designed to detect AI-generated content (e.g., for checking for copyright), AI-writing styles, content originality, and to ensure the naturalness and flow of articles. Undoubtedly, generative AI is reshaping the way artists and journalists operate. For an in-depth exploration of the impact and implications of these technological advancements on news organizations, refer to the survey conducted by Beckett et al. Beckett and Yaseen (2023).

Generating text and scripts automatically can also be done through image and video inputs without text prompts (e.g., image captioning Stefanini et al. (2023)) and with text prompts. These approaches are referred to as Vision Language Models (VLMs): multimodal models that learn from images and text. The most common and prominent models often consist of an image encoder, an embedding projector to align image and text representations, often via a dense neural network, and a text decoder stacked in this order. The most well-known technique is Contrastive Language-Image Pre-training (CLIP) Radford et al. (2021). More recent work in Wei et al. (2024) scales up the vision vocabulary by incorporating new image features into the existing CLIP model, resulting in improved content understanding. A comprehensive survey of VLMs for vision tasks can be found in Zhang et al. (2024a).

3.1.2 Audio and music generation

Similar to language models, AI-based music generation has rapidly advanced due to unsupervised learning on large datasets and the use of transformers (see Section 2.2). Examples of such systems include MuseNet[^13], Magenta Studio[^14], and Musicfy[^15]. These tools assist in music composition by learning complex musical patterns, predicting the next word or music note in a sequence, and mixing specified instruments. Moreover, AI tools can convert one type of sound into another, such as from whistling to violin or from flute to saxophone[^16]. This capability is invaluable for artists who may not be proficient in playing all the instruments they wish to incorporate, saving both time and costs. In 2024, Suno has released a model capable of producing radio-quality music that can be created in 2 minutes[^17]. Later, Udio[^18] was launched. This offers a prompt to create lyrics and music with a maximum duration of 90 seconds, and also appears to have, at least some, awareness of copyright.

AI voice software changes vocalizations from one person to another, for example enabling users to train the model to convert other people’s voices into their own, e.g., lalals[^19], Kits[^20], Media.io[^21], etc. Certain software, such as Voice.ai[^22], even offers real-time voice changing capabilities. The technologies behind this use a transformer to learn voice features and patterns in mel-spectrogram form. For example, the framework proposed in Yang et al. (2023c) uses a DM-based method with a transformer backbone to turn text input into a mel-spectrogram using the vector quantized variational autoencoder (VQ-VAE) van den Oord et al. (2017). Next, this mel-spectrogram is transformed into a sound wave. Unlike a regular spectrogram, the mel-spectrogram is based on the mel-frequency scale, which offers higher resolution for lower frequencies. Voice style transfer often uses zero-shot learning (a model is trained to recognize classes or categories that it has never encountered during training) Huang et al. (2022b) or few-shot learning (a model trained with only one or a few examples per class) Wang et al. (2022b). Stable Audio Open Evans et al. (2025) introduces a text-conditioned generative model for non-speech audio, trained on Creative Commons licensed data, capable of producing state-of-the-art 44.1kHz stereo audio.

Another emerging AI technology application is in the field of spatial audio. In 2022, Apple Music revealed that, in just over a year, more than 80% of its worldwide subscribers were enjoying the spatial audio experience, with monthly plays in spatial audio increasing by over 1,000%[^23]. With head tracking, this technology significantly enhances the immersive experience. Masterchannel has launched SpatialAI[^24], claiming it to be the world’s first spatial mastering AI. This processes audio files and returns an optimized track for streaming platforms, along with an individually optimized stereo version for traditional distribution. All these advancements leverage transformer-based technologies.

3.1.3 Image generation

As described in Section 2.3, recent advances in AI technologies for image generation are based on Diffusion Models (DMs). Well-known and highly competitive text-to-image models include Stable Diffusion[^25], Midjourney[^26], DALL·E[^27], and Ideogram[^28]. Released in June 2024, the latest version of Stable Diffusion (SD3), has been reported to outperform state-of-the-art text-to-image generation systems such as DALL·E 3 (released August 2023) Esser et al. (2024), Midjourney v6 (released December 2023), and Ideogram v1 (released February 2024) in terms of typography and prompt adherence, based on human preference evaluations. These open-source tools are built on a Multimodal Diffusion Transformer (MM-DiT) architecture, which integrates attention from both text and images. LLM4GEN Liu et al. (2025b) fuses features from LLM and CLIP models to enhance the semantic understanding in text-to-image diffusion models, enabling them to better handle complex and dense prompts involving multiple objects. Examples of text-to-image generation are shown in Fig. 3 (a) comparing the performance of four models, i.e. Ideogram v1, DALL·E 3, Photoshop 2025, and sdxy-turbo by Nvidia. It is clear that hands are one of the most difficult features to generate, e.g., one hand has six fingers.

Figure 3: Text-to-image generation (generated on 27 November 2024). (a) text-to-image generation by Ideogram v1, DALL·E 3, Photoshop 2025, and sdxy-turbo by Nvidia. (b) The top-row images were generated by DALL·E in ChatGPT 4. The bottom-row images are generated by LLM-grounded Diffusion Lian et al. (2024).

Figure 3: Text-to-image generation (generated on 27 November 2024). (a) text-to-image generation by Ideogram v1, DALL·E 3, Photoshop 2025, and sdxy-turbo by Nvidia. (b) The top-row images were generated by DALL·E in ChatGPT 4. The bottom-row images are generated by LLM-grounded Diffusion Lian et al. (2024).

DALL·E 3, available on ChatGPT 4, also provides an inpainting tool, allowing the user to manually select the area to edit. However, as of April 2024, its performance is still limited. As illustrated in Fig. 3 (b), the selected area is the white car, and with the follow-up request to change the white car to the red car, DALL·E 3 generates correctly. However, if asked to replace it with a bicycle, it does not work. LLM-grounded Diffusion Lian et al. (2024) was the first to introduce a framework that allows multiple rounds of user requests without the need for manual selection on the image. This is achieved by generating layout-grounded images, first using stable diffusion and then masking the latent variables as priors for the next round of generation[^29]. Since then, text-driven image editing has seen significant improvements in quality, with most recent approaches adopting Diffusion Transformer architectures Feng et al. (2025b); Huang et al. (2025b).

Similar to DALL·E 3, Photoshop features a Generative Fill tool[^30] designed to generate new images or assist with photo editing. It accepts a text prompt and provides several generation choices. After defining the editing area, users can remove and add new objects (more inpainting tasks are discussed in Section 3.3.5), transfer to new styles, and expand content within images. Recently, Brooks et al. introduced InstructPix2Pix Brooks et al. (2023), a conditional diffusion model that generates image editing examples without predefined editing areas. By combining GPT-3 and Stable Diffusion, the model effectively captures and matches the semantic meaning of the content in both text and image. Sometimes, style and context are not easy to describe in words. Textual Inversion Gal et al. (2023) personalizes large pre-trained text-to-image diffusion models based on specific objects and styles, using 3-5 images of a user-provided concept. ByteDance announced Hyper-SD Ren et al. (2024) which proposed trajectory segmented consistency distillation and provides real-time high-resolution image generation from drawing with a control text prompt.

3.1.4 Video generation and animation

Despite the success of text-to-image generation, text-to-video generation has not advanced at the same pace, growing more rapidly only in 2024 due to its computational expense and content complexity. Several major companies and private platforms have now released offerings, including Gemini 1.5 by Google, Make-A-Video by Meta, and Sora by OpenAI. Make-A-Video Singer et al. (2023), through a spatiotemporally factorized diffusion model, leverages joint text-image priors and super-resolution in space and time, though some results exhibit flickering artifacts[^31]. Gen-2 by Runway[^32] supports both text- and image-to-video generation, producing smooth 4-second clips. In April 2024, Adobe Premiere Pro announced integration of generative AI tools for video extension with third-party models by OpenAI, Runway, and Pika Labs[^33], including contextual selection, inpainting for object removal, and object addition to videos via text prompts.

Text-to-video technologies, combined with AI voice, have been tested not only by artists and producers but also by a wider audience. Results from these experiments—such as automatically turning scripts into trailers and music videos—have been widely shared online[^34]. However, scene composition and transitions still require further editing to meet production needs[^35]. In April 2024, Microsoft introduced VASA-1 Xu et al. (2024a), which converts a single image and speech audio clip into a realistic video of talking faces mimicking expressions and head movements (Fig. 4, right). The resulting quality surpasses Google’s VLOGGER Corona et al. (2024), which uses a similar diffusion-based approach but additionally generates upper-body and hand motion. Recently, ByteDance proposed an audio-driven interactive head-generation model Zhu et al. (2024c) offering listening and speaking states during multi-turn conversations, based on a conditional diffusion transformer.

The main technologies underpinning text-to-video and image-to-video tasks are based on diffusion models (DMs) combined with 3D convolutions—or separate spatial and temporal convolutions—and attention modules Wang et al. (2023a). Tune-A-Video Wu et al. (2023d) modifies the style of an input video using text prompts, leveraging pretrained text-to-image models and attention tuning for temporal consistency. Early methods often exhibited flickering, as observed in the CVPR2023 text-guided video editing competition. Dreamix Molad et al. (2023) mitigates this but produces blurry videos. CogVideo Hong et al. (2023) employs VQ-VAE to convert frames into tokens fused with text embeddings to generate new videos. Phenaki Villegas et al. (2023) uses transformers for variable-length outputs, though with lower quality than DMs. Comprehensive evaluations appear in Liu et al. (2023c). More recent work applies spatiotemporal layers to model dynamics Gupta et al. (2024), redesigning transformer blocks for latent video diffusion with restricted spatial and spatiotemporal attention. LaVie Wang et al. (2025c) shows that simple temporal self-attention, combined with rotary positional encoding, effectively captures temporal correlations. Image-to-video generation is analogous to text-to-video but conditions diffusion models on images rather than text; hybrid approaches combine textual descriptions (for motion) and images (for scene layout) Wu et al. (2025a). An increasing number of free and commercial tools are emerging, including Veo 3 by Google DeepMind[^36], Kling AI[^37], Pika 2.2[^38], and Hailuo AI[^39]. Though not perfect, their generated videos appear remarkably realistic.

Generating characters with human posture and motion from text prompts has also become popular. Make-An-Animation Azadi et al. (2023) trains on image-text datasets and fine-tunes on motion capture data, adding additional layers to model the temporal dimension. Animate Anyone by Alibaba Group Hu (2024) inputs a real photo or anime of a person with a sequence of guided poses. The results are significantly better than existing techniques, including Disco Wang et al. (2024b) and Bidirectionally Deformable Motion Modulation (BDMM) Yu et al. (2023). They also suggest using Animate Anyone with Outfit Anyone[^40] to produce a character with a reference outfit.

Viggle[^41] claims to be the first video-3D foundation model embodying an actual understanding of physics. It combines a character and a text prompt about motion to generate character animation. Available AI tools for 3D on the market include DeepMotion[^42] that offers text-to-3D post animation and video-to-3D post animation, shown in Fig. 4 (left). The later function can track multiple people from real video and generate replicated characters with the same motions.

Figure 4: (Left) Video-to-3D post animation by DeepMotion. (Right) Image and audio to video by VASA-1 Xu et al. (2024a)

Figure 4: (Left) Video-to-3D post animation by DeepMotion. (Right) Image and audio to video by VASA-1 Xu et al. (2024a)

3.1.5 Augmented, virtual and mixed reality, and 3D content

While the benefits of LLMs in Augmented Reality (AR) directly target educational purposes, enhance cognitive support, and facilitate communication Xu et al. (2025), mixed reality (MR) has once again become exciting since the release of the Apple Vision Pro in February 2024. This demonstrated the potential of MR experiences by merging real-world environments with computer-generated ones. Thanks to the rapid growth of AI-based 3D representation (see Section 3.5), the generation of AR/VR/MR content has advanced significantly. Real-time rendering with immersive interaction has improved, and real scenes can now be generated avoiding uncanny valley effects. There has also been an attempt to use autoregressive and generative models to estimate lighting, achieving a visually coherent environment between virtual and physical spaces in AR Zhao et al. (2024d).

Similar to other content generation tools, LLMs have been influenced on immersive technologies, including text-to-3D and image-to-3D. Exciting examples include Holodeck Yang et al. (2024d), which automatically generates 3D embodied environments via text-prompt interactions with a large language model (GPT-4). 3D objects are gathered from Objaverse Deitke et al. (2023), a dataset with 800K+ annotated 3D objects. RealFusion Melas-Kyriazi et al. (2023), a single image to 3D object generator, merges 2D diffusion models with NeRF, improving Instant-NGP Müller et al. (2022), which provides an API for VR controls. NeuralLift-360 Xu et al. (2023a) also uses diffusion models to generate priors for novel view synthesis. Magic123 Qian et al. (2024) is the latest image-to-3D tool that uses 2D and 3D priors simultaneously to produce high-quality high-resolution 3D geometry and textures. DreamGaussian Tang et al. (2024a) offers text-to-3D and image-to-3D by adapting 3D Gaussian splatting (more in Section 3.5.3) into generative settings using a diffusion prior. This generates photo-realistic 3D assets with explicit mesh and texture maps within only 2 minutes. DreamGaussian4D Ren et al. (2023) employs image-to-video diffusion and a 4D Gaussian Splatting representation to generate an image-to-4D model. The results are not very sharp, but they can be further edited with Blender.

In July 2024, Shutterstock launched its Generative 3D service in commercial beta, powered by NVIDIA Edify, a multimodal generative AI architecture. This service enables creators to rapidly prototype 3D assets and generate 360-degree HDRi backgrounds to light scenes using text or image prompts. In conjunction with OpenUSD, the created scenes can be rendered into 2D images and used as input for AI-powered image generators, allowing for the production of precise, brand-accurate visuals.

3.2 Information analysis

3.2.1 Text categorization

Applications of text categorization include detecting spam emails, automating customer support, monitoring social media for harmful content, etc. At its core, text categorization involves assigning predefined labels to text documents, which can be anything from a tweet to a lengthy article. LLMs are particularly well-suited for this task due to their ability to comprehend complex and nuanced language. One of the main advantages of using LLMs in text categorization is their transfer learning capability. Models can be pre-trained on a large amount of text and then fine-tuned on a smaller, task-specific dataset, with or without further post-processing techniques. For example, CARP Sun et al. (2023) applies kNN to integrate a diagnostic reasoning process for final decision. ChatGraph, proposed by Shi et al. Shi et al. (2023), utilizes ChatGPT to refine text documents. It uses a knowledge graph, extracted using another specifically defined prompt, and finally, a linear model is trained on the text graph for classification. Multiple learners are also used to enhance the performance Hou et al. (2023); Ai et al. (2025).

3.2.2 Advertisements and film analysis

Not only does AI assist in generating ideas and content, but it can also aid creators in effectively matching content to their audiences, particularly on an individual level Feizi et al. (2023). This effectively helps in advertising personalization—eMarketer[^43] reported that nearly nine out of ten consumers are comfortable with their browsing history being utilized to create personalized ads. In contrast to outdated syntax-style searches, advanced LLM tools can comprehensively grasp user intent behind each search through conversation prompts, providing advertisers with a high level of granularity.

Current advances in generative AI would greatly benefit sentiment analysis, also known as opinion mining, where opinions are gathered from social media, articles, customer feedback, and corporate communication and are analyzed to understand the emotion of the owners. This is a potential tool for filmmakers and studios, enabling the creation of effective and targeted marketing campaigns. By analyzing viewer emotions and opinions, AI can provide valuable insights into audience preferences, aiding in the optimization of film marketing strategies. Sentiment analysis with modern generative AI produces more accurate results. Technically, LLMs learn complex patterns and relationships in text data for sentiment classification Mao et al. (2023); Krugmann and Hartmann (2024). SiEBERT Hartmann et al. (2023) provides a pre-trained model with open-source scripts to be fine-tuned to further improve accuracy for novel applications. Cinema Multiverse Lounge Ryu et al. (2025), a multi-agent conversational system, allows users to interact with LLM-driven agents, each embodying a distinct film-related target user.

3.2.3 Content retrieval and recommendation services

Generative retrieval (GR) was pioneered by Metzler et al. Metzler et al. (2021). Unlike traditional retrieval, which adheres to the “index-retrieve-then-rank” paradigm, the GR paradigm employs a single model to obtain results from query input. The model generally involves deep-learning based transformers, generating output token-by-token. More recent work in Li et al. (2024e) introduces learning-to-rank training to enhance the performance system up to 30%. GR has several advantages including substituting the bulky external index with an internal index (i.e., model parameters), significantly reducing memory usage, and enabling optimization during end-to-end model training towards a universal objective for information retrieval tasks. Conversational question answering techniques have been integrated to enhance the document retrieval Li et al. (2024d).

When retrieving visual content, recent work exploits generative models to enhance content-based model search Lu et al. (2023). These models decode the text, image, or video query into samples of possible outputs, which are then used to learn statistics for better matching between the query and output candidates. DMs are also employed for visual retrieval tasks, where they learn joint data distributions between text queries and video candidates Jin et al. (2023). A comprehensive survey on Generative Information Retrieval is available in Li et al. (2025).

While the retrieval task involves users directly defining a specific query input, recommendation services operate by retrieving content based on previous usage patterns. Essentially, a recommendation engine is a system that suggests products, services, or information to users through data analysis. Research in Chua et al. (2023) has reported a positive association between buyers’ attitudes toward AI and their behavioral intention to accept AI-based recommendations, with potential for further growth. Notable examples include the recommendation framework developed by Google Rajput et al. (2023), which utilizes GR. This framework assigns Semantic IDs to each item and trains a retrieval model to predict the Semantic ID of an item that a given user may engage with. A report by Aggarwal et al. Aggarwal (2025) states that the recommendation accuracy of recommendation services has increased from 45.0% to 91.5% with the integration of generative AI.

3.2.4 Intelligent assistants

Intelligent assistants refer to software programs or applications that use AI and NLP to interact with users and provide helpful responses or perform tasks. These assistants can range from simple chatbots to sophisticated virtual agents capable of understanding and responding to complex queries. They’re designed to assist users in various tasks, from answering questions and providing information to scheduling appointments and controlling smart home devices.

Current LLMs obviously enhance the performance of intelligent assistants, designed to understand complex inquiries and generate more natural conversational responses, such as Sasha King et al. (2024). Generative AI can also be used to enhance the performance of human customer support agents, aiding in search and summarization, as discussed in the previous section. Brynjolfsson et al. Brynjolfsson et al. (2023) examined the implementation of a generative AI tool designed to offer conversational guidance to customer support agents. Their research revealed that AI assistance significantly enhances problem resolution and customer satisfaction. Furthermore, they observed that AI recommendations prompt low-skill workers to adopt communication styles akin to those of high-skill workers. AI-based intelligent assistants may currently be more focused on educational purposes, but they can clearly help artists write more efficiently Lee et al. (2024) or assist in customizing personal requirements Sajja et al. (2024). The performance of personalized assistants can be enhanced with domain-specific knowledge to provide more in-depth responses to users Jiang et al. (2025).

3.3 Content enhancement and post production workflows

3.3.1 Enhancement

In our previous review paper Anantrasirichai and Bull (2022), we discussed AI technologies for contrast enhancement and colorization as separate topics, as methods were developed specifically for each task. However, in recent years, there has been a shift towards addressing more complex issues, such as those encountered in low-light environments and underwater scenarios. These real-world situations often involve a combination of challenges, including low contrast, color imbalance, and noise.

In low-light conditions, scenes often exhibit low contrast, leading to focusing difficulties or the need for long exposures, which can result in blurred images and videos. To address this, LEDNet Zhou et al. (2022) has introduced a synthetic dataset for such scenarios and incorporated a learnable non-linear activation function within the network to enhance feature intensities. Meanwhile, SNR-Aware Xu et al. (2022) estimates spatial-varying Signal-to-Noise Ratio (SNR) maps and proposes local and global learning branches using ResNet and transformer architectures, respectively. NeRCo Yang et al. (2023g) addresses the low-light problem with INR, which unifies the diverse degradation factors of real-world scenes with a controllable fitting function. Diffusion models (DMs) have also become popular choices for low-light image enhancement HOU et al. (2023); Yi et al. (2023); Jiang et al. (2023a). Diff-Retinex Yi et al. (2023) formulates the low-light image enhancement problem into Retinex decomposition, and employs multi-path generative diffusion networks to reconstruct the normal-light Retinex probability distribution. A recent state-of-the-art approach presented in Jiang et al. (2023a) decomposes images into high and low frequencies using wavelet transform. High frequencies are enhanced using a transformer-based pipeline, while the low frequencies undergo a diffusion process. This method achieves nearly 2.8dB improvement over the state-of-the-art transformer-based approach, e.g. LLFormer Wang et al. (2023b), and significantly better than INR-based method, NeRCo Yang et al. (2023g), on a real low-light image benchmarking dataset. The technique has been extended for video enhancement in Lin et al. (2024b). The output of the enhancement typically depends on user preferences. This has been viewed as a one-to-many inverse problem, with attempts to solve it using Bayesian approaches. For example, a Bayesian Enhancement Model (BEM) Huang et al. (2025a) incorporates Bayesian Neural Networks (BNNs) to capture data uncertainty and produce diverse outputs. The method can be used with Transformers or Mamba as the architecture backbone.

Regarding video enhancement, transformer and DMs are still in their early stages. STA-SUNet Lin et al. (2024c) has demonstrated that using transformers for low-light video enhancement outperforms CNN-based methods Anantrasirichai et al. (2024). The recent Mamba-based network huang2025bvi also demonstrates promising results, outperforming STA-SUNet by more than 2 dB in PSNR. It is important to note that low-light enhancement is subjective. While most training datasets use normal lighting conditions as ground truth Lin et al. (2024a), the enhanced images and videos may alter the mood and tone of the content. Therefore, the tools for creative industries should be adjustable, not only for entire images and videos but also adaptive to specific areas and content. For instance, CLE Diffusion Yin et al. (2023) enables user-friendly editing of lighting with fine-grained regional controllability.

Recent efforts have focused on enhancing User-Generated Content (UGC) videos—authentic recordings created by individuals rather than brands, often showcasing real experiences with products or services. The winning solution of the NTIRE 2025 Challenge on UGC Video Enhancement Safonov et al. (2025) implemented a pipeline of four sequential modules: color enhancement, denoising, BasicVSR++ restoration Chan et al. (2022), and SwinIR Liang et al. (2021). This method achieved a 17% higher subjective score than the second-place entry, which used a two-stage framework, highlighting a notable improvement in perceived visual quality.

3.3.2 Style transfer

Style transfer in AI art refers to a technique where the artistic style of one image (or video) is applied to another image (or video) while preserving the content of the latter. Style transfer has numerous applications in art, design, and image editing, allowing artists and designers to create unique and visually appealing compositions by blending different artistic styles with existing images (or videos). The applications also include image-to-image and sequence-to-sequence translations.

StyTr2 Deng et al. (2022) is the first transformer-based method for style transfer, applying content as a query and style as a key of attention. InST Zhang et al. (2023d) utilizes Stable Diffusion Models as the generative backbone and introduces an attention-based textual inversion module to learn the description of the content. StableVideo Chai et al. (2023) uses a text prompt to describe the desired appearance of the output, transforming the input video to have a new look based on a diffusion model. For instance, a video of a white car driving in summer can be altered to show a red car driving in winter. A large pre-trained DM is employed in Chung et al. (2024), where the style is injected to manipulate the self-attention of the decoder. To deal with the disharmonious color, they propose an adaptive instance normalization. A survey of style transfer using transformers and diffusion models can be found in Zhou et al. (2025). Implicit Neural Representations (INRs) are less commonly used in style transfer tasks due to the difficulty of modeling the cross-representation between style and content. Moon et al. Moon et al. (2023) combined INRs with vision transformers for generalizable style transfer; however, the results remain limited in quality. In contrast, the method proposed by Kim et al. Kim et al. (2024b) uses multilayer perceptrons (MLPs) to map image coordinates to the colors of the stylized output, guided by features extracted from both the content and style inputs to allow controllability.

3.3.3 Upscaling imagery: super-resolution (SR)

Impressive super-resolution (SR) results from transformer and diffusion models have been published extensively in the past few years. Originally, SR methods were developed using multiple low-resolution (LS) images, as different features in each image are combined to construct an enhanced one. However, these methods are not practical, as in most cases only one LS image is available. Hence, more methods have been developed for single image super-resolution (SISR).

The first use of a transformer, called ESRT, was for capturing long-term dependencies, such as repeating patterns in buildings. This was done in the feature domain extracted by a lightweight CNN module Lu et al. (2022), outperforming those that use only CNNs. Since then, most SISR methods have been based on transformers. The Hybrid Attention Transformer (HAT) Chen et al. (2023b) was introduced, which improves the SR quality over ESRT by more than 2dB when upscaling 2$\times$-4$\times$. However, the NTIRE 2023 Real-Time Super-Resolution Challenge Conde et al. (2023) showed that the winner, Bicubic++ Bilecen and Ayazoglu (2023), uses only convolutional layers and achieves the fastest speed at 1.17ms in upscaling 720p to 4K images. This method is significantly faster than any of the participants in the NTIRE 2025 Challenge Chen et al. (2025b), where Transformer-based architectures continue to dominate as the mainstream approach.

For DMs, SR3 by Google Saharia et al. (2023) has produced truly impressive results. It operates by learning to transform a standard normal distribution into an empirical data distribution through a sequence of refinement steps, interpolating in a cascaded manner—upscaling 4$\times$ at a time. Later, IDM Gao et al. (2023) combines INR with a U-Net denoising model in the reverse process of the DM. It is crucial to emphasize again that DMs are generative models. The SR results are generated based on the statistics we provide to the model during training (LR training samples). This is not for a restoration task, but rather for synthetic generation. A survey in SISR using DMs can be found in Moser et al. (2024).

For video SR, numerous methods have emerged as part of a unified enhancement framework, as discussed in the previous section. One of the pioneering works to incorporate transformers specifically for video SR tasks is the Trajectory-aware Transformer for Video Super-Resolution (TTVSR) Liu et al. (2022a). Although the results are slightly inferior to those of BasicVSR++ Chan et al. (2022), which employs CNN and was introduced around the same time, both methods significantly enhance detail and sharpness compared to previous approaches, albeit not in real time. To address this limitation, the Deformable Attention Pyramid Fuoli et al. (2023) has been introduced, offering slightly lower quality but a speed-up of over 3$\times$. Recently, Adobe announced their VideoGigaGAN Xu et al. (2024b), which can perform 8$\times$ upsampling. This is achieved by adding flow estimation and temporal self-attention to the GigaGAN upsampler Kang et al. (2023), which is primarily used for image SR, and text-to-image synthesis. Cao et al. Cao et al. (2025) introduce a zero-shot video super-resolution framework that leverages a pre-trained image diffusion model, and replaces the spatial self-attention layer with a novel short-long-range (SLR) temporal attention layer. Recently, SeedVR integrated text information (captions) into a Diffusion Transformer (DiT) model, achieving state-of-the-art performance in video super-resolution.

Compared to traditional upscaling methods, generative AI can add details that did not exist in the original input image. These methods excel at generating high-quality natural images and structures, such as buildings, which are commonly included in training datasets. However, the process can be slow and may produce unpredictable results if the input image has very low resolution or contains content rarely seen in natural images. As shown in Fig. 5 (left), generative AI fails to upscale the knitting texture areas, instead generating lines more commonly found in typical images. While AI methods produce sharper edges, they perform less effectively on text.

Figure 5: (Left) Examples of SR ($\times$4) using generative model. (Right) Real-time portrait editing with FacePoke.

Figure 5: (Left) Examples of SR ($\times$4) using generative model. (Right) Real-time portrait editing with FacePoke.

3.3.4 Restoration

In our previous review paper Anantrasirichai and Bull (2022), we categorized the work on restoration into several different types of distortions, including deblurring, denoising, dehazing, and mitigating atmospheric turbulence. Recent work, however, uses a unified network architecture to address these as inverse problems $y=hx+n$, where $x$ and $y$ are the ideal and observed data, respectively. $h$ is a degradation function, such as blur, and $n$ is additive noise. Often, the super-resolution task is also considered as an inverse problem, meaning $h$ includes the downsampling process. Note that although designed as a single network, the model is trained with each distorted dataset separately.

The pioneering transformer-based method for image restoration, SwinIR Liang et al. (2021), employs several concatenated Swin Transformer blocks Liu et al. (2021). SwinIR surpasses state-of-the-art CNN-based methods proposed up to the year 2021 in super-resolution and denoising tasks. The model is smaller and reconstructs fine details more effectively. Other two popular approaches that emerged in the same timeframe are Uformer Wang et al. (2022a) and Restormer Zamir et al. (2022). Both incorporate Transformer blocks into hierarchical encoder-decoder networks, employing skip connections similar to those in U-Net. Their objective was to restore noisy images, sharpen blurry images, and remove rain. The networks focused on predicting the residual $R$ and obtaining the restored image $\hat{x}$ through $\hat{x}=y+R$. While their performance is very similar, Restormer has half the parameters of Uformer. More recently, GRL by Li et al. Li et al. (2023c) exploits a hierarchy of features in a global, regional, and local range using different ways to compute self-attentions as an image often show similarity within itself in different scales and areas, which outperforms SwinIR and Restormer. Additionally, Fei et al. introduced the Generative Diffusion Prior Fei et al. (2023) for unsupervised learning, aiming to model posterior distributions for image restoration and enhancement. VmambaIR Shi et al. (2025) incorporates Mamba blocks into the U-Net architecture, achieving superior performance compared to SwinIR and Restormer in both visual quality and model size.

For video restoration, the general framework comprises frame alignment, feature fusion and reconstruction. The process could be similar to image restoration but input multiple frames and run through the sequences in a sliding window manner to exploit the temporal information within a number of consecutive frames. Recently, Video Restoration Transformer (VRT) Liang et al. (2024) and its improved version with a recurrent process (RVRT) Liang et al. (2022), have emerged as the state of the arts for video super-resolution, deblurring, denoising, and frame interpolation. This method introduces temporal reciprocal self-attention in the transformer architecture and parallel warping using MLP. These innovations enable parallel computation and outperform the previous state-of-the-art methods by up to 2.16dB on benchmark datasets. FMA-Net Youk et al. (2024) proposed multi-attention for joint video super-resolution and deblurring, achieving fast runtime with nearly 40% improvement over RVRT, and the restored quality was reported better by up to 3%.

For audio restoration, most software discussed in Section 3.1.2 offers tools for enhancing audio quality, such as eliminating background noise, echo, microphone rumble, and occasionally room reverberation, which have been well-established even before the advent of deep learning. There have been efforts to utilize AI for learning global contextual information to aid in the removal of unwanted sounds, leading to better final quality Yu et al. (2022). The latest advancements in this domain are primarily focused on addressing issues where significant portions of the audio data are missing. For instance, Moliner et al. Moliner et al. (2023) tackle problems such as audio bandwidth extension, inpainting, and declipping by treating them as inverse problems using a diffusion model. For a comprehensive survey on the use of diffusion models in restoration tasks, refer to He et al. (2025a).

The following methods have been proposed for specific problems, but ideally, they should be adaptable for other tasks, even though they may not perform as well as they do for the original task.

i) Deblurring: A lightweight deep CNN model was recently proposed in Pan et al. (2023a), where a new discriminative temporal feature fusion has been introduced to select the most useful spatial and temporal features from adjacent frames. Feature propagation along the video is done in the wavelet domain. The deblurring performance is comparable to RVRT Liang et al. (2022), but it is 5 times faster. DaBiT Morris et al. (2025) mitigates focal blur content with depth information and applies SR for further enhancing fine details. Note that not only in software, but AI technologies have also been integrated into hardware. This includes autofocus, which is crucial for capturing sharp images of subjects, especially in dynamic environments where manual adjustments are impractical due to rapid movement. AI-driven autofocus methods have emerged, often tailored for specific camera hardware. For instance, Choi et al. proposed an autofocus model optimized for dual-pixel Canon cameras Choi et al. (2023). Additionally, Yang et al. investigated the correlation between language input and blur map estimation, utilizing semantic cues to enhance autofocus performance Yang et al. (2023d). Remarkably, their model achieves comparable results to previous state-of-the-art methods while being more lightweight Yang et al. (2023h). Autofocus could be used in conjunction with real-time object tracking (see Section 3.4.3) to produce desirable sharpness for moving objects in the video. Recently, Feng et al. Feng et al. (2025a) proposed a novel residual diffusion deblurring framework that integrates a conditional diffusion model guided by a defocus map and incorporates residual learning into the single-image defocus deblurring process.

ii) Denoising: SUNet Fan et al. (2022) applies Swin transformer blocks combined in a UNet-like architecture. Denoising with diffusion models (DMs) Yang et al. (2023a) has been proposed by diffusing with estimated noise that is closer to real-world noise rather than Gaussian noise, achieving better performance than SwinIR Liang et al. (2021) and Uformer Wang et al. (2022a). INR with complex Gabor wavelets as activation functions show promising denoising results Saragadam et al. (2023). The NTIRE 2025 Image Denoising Challenge Sun et al. (2025) revealed that the top-performing methods combined transformer-based and convolutional network architectures. Similarly, recent advances in video denoising also adopt a hybrid approach that integrates both architectures Jin et al. (2025); Yue et al. (2025).

iii) Dehazing: Vision transformers for single image dehazing were proposed in DehazeFormer Song et al. (2023). Similar to SUNet, it is a UNet-like architecture, but introduces Rescale Layer Normalization for better suit on improving contrast. The Fast Fourier Transform (FFT) has been employed in Fang et al. (2025) due to the phase spectrum conveying more structural detail than the amplitude spectrum and demonstrating greater robustness to contrast distortion and noise. Then cross-attention between the RGB and YCbCr color spaces is applied. This approach achieves nearly 5 dB higher PSNR than DehazeFormer on a real-world smoke dataset. For video dehazing, Xu et al. Xu et al. (2023b) introduced a recurrent multi-range scene radiance recovery module with the space-time deformable attention. They also employ physics prior to inform haze attenuation. This method outperforms DehazeFormer by approximately 1dB.

iv) Mitigating atmospheric turbulence: Similar to dehazing, physics-inspired models have been widely developed to remove turbulence distortion Jaiswal et al. (2023); Jiang et al. (2023b), while complex-valued CNNs have been proposed to exploit phase information Anantrasirichai (2023). There was also an attempt to use instance normalization (INR) to solve this problem, offering tile and blur correction Jiang et al. (2023b). However, diffusion models outperform on a single image Nair et al. (2023), and transformer-based methods remain state-of-the-art for restoring videos Zhang et al. (2024c); Zou and Anantrasirichai (2024). Mamba architecture, employed in Hill2025MAMAT, outperforms Transformers and improves object detection performance. A recent review can be found in Hill et al. (2025).

3.3.5 Inpainting

Visual inpainting is the process of filling in lost or damaged parts of an image or video. CNNs and GANs have already achieved impressive results (see our previous review paper Anantrasirichai and Bull (2022)). Recent work has focused more on editing rather than simply filling in the missing areas. This means users can now mask large areas of an image, and AI tools generate multiple results for users to choose from, a technique known as pluralistic inpainting Zheng et al. (2019). Some notable methods include the following: Mask-Aware Transformer (MAT) Li et al. (2022b) offers several outputs to fill a large missing area, consisting of a convolutional head, a transformer body, and a convolutional tail for reconstruction, along with a Conv-U-Net for refinement. PUT Liu et al. (2022b) proposes a patch-based vector VQ-VAE and unquantized Transformer to minimize information loss. Spa-former Huang et al. (2024a) employs a UNet-like architecture, where each level performs transformer with sparse self-attention to remove coefficients with low or no correlation, leading to memory reduction, while improving result quality by up to 5% compared to PUT.

Video inpainting presents greater complexity compared to image inpainting, despite the abundance of information available in an image sequence. The process typically involves tracking masks across frames, estimating optical flow, and ensuring temporal consistency. The current state-of-the-art methods include DLFormer Ren et al. (2022) and ProPainter Zhou et al. (2023). DLFormer conducts inpainting in latent space and utilizes discrete codes for video representation. On the other hand, ProPainter employs flow-based deformable alignment to enhance robustness to occlusion and inaccurate flow completion. The method excels in filling complete and rich textures, achieving a speed of 12 fps for full HD video. Video inpainting is also used for dubbing. DINet Zhang et al. (2023e) replaces the mouth area to synchronize with a new language being spoken.

A comprehensive survey of learning-based image and video inpainting, covering approaches such as CNNs, VAEs, GANs, transformers, and diffusion models, can be found in Quan et al. (2024). Additionally, Elharrouss et al. Elharrouss et al. (2025) provide an in-depth review of the current challenges and future directions specific to transformer-based inpainting techniques.

3.3.6 Image Fusion

Image fusion is the process of merging multiple images from either the same source (such as varying focal points or exposures) or different modalities (e.g. visible and infrared cameras) into a single image. This process integrates complementary information from the various images to enhance overall quality, improve interpretation, and increase the usability of the final image.

Transformers and CNNs have been combined to extract global and local information, respectively. Most methods use CNNs for feature extraction, with transformers operating in the latent space Ma et al. (2022); Rao et al. (2023). Notable methods include SwinFusion Ma et al. (2022), which utilizes a self-attention-based intra-domain fusion unit and a cross-attention-based inter-domain fusion unit to achieve multi-modal and digital photography image fusion. Transformer-based image fusion has also been applied to downstream tasks like segmentation Liu et al. (2023b), achieving superior results by leveraging the additional information. Self-attention blocks are employed to enhance intra-feature representations, while the cross-attention mechanism integrates inter-feature information to improve the quality of the fused output Li and Wu (2024).

DDFM, the first diffusion model-based image fusion method, estimates noise in the reverse process by combining multiple inputs Zhao et al. (2023c). The expectation-maximization (EM) algorithm is integrated to estimate the noise distribution parameters, resulting in sharper images compared to traditional DDPM. For an in-depth review, the reader is referred to recent work in Karim et al. (2023); Zhang and Demiris (2023).

3.3.7 Editing and Visual Special Effects (VFX)

Editing or modifying specific areas of an image is much easier with DM technologies, particularly for headshot photos, such as targeting the eyes and mouth on the face Guo et al. (2024b). This capability has been extended to video generation (see Section 3.1.4). Fig. 5 shows an example of the online tool, FacePoke[^44], which allows users to move the head and modify the shapes of the eyes and mouth in real time. Motion-I2V Shi et al. (2024b) provides motion blur and motion drag tools to control specific areas of an image to add motion. The method is based on a diffusion-based motion field predictor and motion-augmented temporal attention.

VFX aims to create and/or manipulate imagery outside the context of a live-action shot in filmmaking and video production. When adding objects, scenes, and effects into traditional photographic videos, generative AI has obviously become an important tool, but some manual operations are still required. For example, in After Effects (EA)[^45], the user selects the area where the object will be added and uses text prompts to describe such object. Subsequently, with the current EA version, the user will need to apply motion tracking so the generated object is moved accordingly.

AI technologies can upscale, enhance, and restore low-quality or old footage. For example, standard definition videos can be converted to high definition or even 4K quality without traditional manual remastering processes. This is particularly useful for remastering old movies or enhancing visual details in scenes. Generative AI has also simplified and accelerated automated processes, such as rotoscoping Tous (2024), an animation technique where animators trace over motion picture footage frame by frame to create realistic action. AI models can accurately detect and segment objects and characters in video frames, significantly speeding up the post-production process. Additionally, AI can assist the rapid creation of 3D models from 2D images generating realistic animations with minimal input data, facilitating complex human motions and synchronized facial expressions to voiceovers. One restriction is that current technologies still cannot yet generate full 4K accurate visual effects.

3.4 Information Extraction and Understanding

AI plays a crucial role in automating and optimizing the process of information extraction and understanding, enabling organizations to derive actionable insights from large and diverse data. Yan et al. Yan et al. (2023) have categorized information extraction tasks based on the Format-Time-Reference space, as illustrated in Fig. 6 (a), where object detection and video object segmentation (VOS) are considered to be the simplest and the most complex tasks, respectively. Recent advancements in this field draw significant inspiration from LLMs. These advancements include the utilization of prompts as conditional inputs for acquiring information. Moreover, following the pipeline approach used in LLMs, there is a growing trend towards leveraging very large datasets to pre-train models before fine-tuning them for specific downstream tasks. For instance, Meta AI Oquab et al. (2024) has introduced DINOv2, aimed at enriching information about visual content through self-supervised learning. This model was trained with 142 million carefully selected images, employing the ViT architecture. Google have introduced VideoPrism Zhao et al. (2024a), a tool for scene understanding including classification, localization, retrieval, captioning, and question answering (QA). The model was trained on an extensive and diverse dataset consisting of 36 million high-quality video-text pairs and 582 million video clips accompanied by noisy or machine-generated parallel text.

Figure 6: (a) Tasks in Object-centric understanding defined by Yan et al. Yan et al. (2023) (REC:, Referring Expression Comprehension, RES: Referring Expression Segmentation, VOS: Video Object Segmentation, RVOS: Referring Video Object Segmentation, MOT: Multiple Object Tracking, MOTS: Multi-Object Tracking and Segmentation, VIS: Video Instance Segmentation, SOT: Single Object Tracking. (b) Current high-quality segmentation Ke et al. (2023).

Figure 6: (a) Tasks in Object-centric understanding defined by Yan et al. Yan et al. (2023) (REC:, Referring Expression Comprehension, RES: Referring Expression Segmentation, VOS: Video Object Segmentation, RVOS: Referring Video Object Segmentation, MOT: Multiple Object Tracking, MOTS: Multi-Object Tracking and Segmentation, VIS: Video Instance Segmentation, SOT: Single Object Tracking. (b) Current high-quality segmentation Ke et al. (2023).

3.4.1 Segmentation