OER·harvester

← Back to the library
arXiv HTML resource

Advances in Artificial Intelligence: A Review for the Creative Industries

Artificial intelligence (AI) has undergone transformative advances since 2022, particularly through generative AI, large language models (LLMs), and diffusion models, fundamentally reshaping the creative industries. However, existing reviews have not comprehensively addressed these recent breakthroughs and their integrated impact across the creative production pipeline. This paper addresses this gap by providing a s…

Licence
OPEN CC-BY-4.0
Authors
Nantheera Anantrasirichai, Fan Zhang, David Bull
Published
2025-01-06 · arXiv
Language
en
Length
42227 words
Type
narrative text

Cites 155 works

inferred
Open original ↗

1 Introduction

The influence of artificial intelligence (AI) has grown dramatically over the past few years, particularly due to the rise of generative AI and large language models (LLMs). These advancements are widely regarded as beneficial by many countries, creating significant opportunities for growth (e.g. as outlined in the UK, by the Authority of the House of Lords Communications and Digital Committee (2024)). These advances have also had significant direct and indirect impacts on the creative industries, influencing the direction of their growth. Generative AI, for instance, primarily focuses on generating new data that is not identical to the training data yet shares similarities with it. However, the cardinality of the training data can be huge, larger than what any individual human has ever encountered. The resulting output may therefore act as a new source of inspiration.

AI tools also provide opportunities for a wider range of users to work more efficiently and effectively, with even greater creativity. Moreover, these new technologies not only influence creators, but they also enable new ways for audiences to experience art and culture Jeary and Gajjar (2024).

A major breakthrough in generative AI has been led by OpenAI[^1], an AI research and deployment company, with their introduction of Generative Pre-trained Transformer (GPT) models for LLMs. LLMs are specifically designed to understand and generate human language. They are characterized by their vast size in terms of parameters and the amount of training data used to create them. This breakthrough was particularly impactful when the company released ChatGPT in 2022, which was fine-tuned from a model in the GPT-3.5 series. ChatGPT is a conversational model that includes advanced safety features that mitigate the generation of inappropriate content. Several other LLM platforms were also developed contemporaneously, such as LaMDA and PaLM by Google AI, Ernie Bot by Baidu, and BLOOM by BigScience. Additionally, Anthropic launched Claude, the LLM trained specifically to be harmless and honest, leveraging reinforcement learning from human feedback (RLHF) - a technique used to train AI systems to appear more human Bai et al. (2021). Nonetheless, ChatGPT stands out as the most renowned, thanks to its quick and efficient responses, and notably its public accessibility, being available for free.

Another breakthrough in 2022 was in the area of text-to-image models. OpenAI achieved a significant milestone with DALL·E 2, producing impressive artworks and photorealistic images despite its limited language understanding. Midjourney by Midjourney, Inc., another well-known text-to-image generator, supports image resolutions of 1024$\times$1024 pixels and also provides upscaling tools. Stable Diffusion by Stability AI, for which the code and model weights are publicly available[^2], allows developers and artists to further adapt AI to suit their own specific applications.

The next breakthrough occurred in 2023 when OpenAI unveiled GPT-4, a significantly larger model with third-party estimates of 1.8 trillion parameters; OpenAI has not disclosed an official figure. It also demonstrated improved performance compared to its predecessors OpenAI et al. (2023). However, this still many orders of magnitude smaller than the human brain’s estimated 100–1,000 trillion synaptic connections[^3]. GPT-4 is a multimodal large language model that can generate responses to both text and images. In the ChatGPT interface, GPT-4 is integrated with DALL·E 3, enabling it to comprehend a much broader range of nuances and details than earlier versions. In March 2024, Claude 3 Opus by Anthropic was released, boasting multimodal capabilities in generating images, tables, graphs, and diagrams. Moreover, Anthropic claims that Claude 3 Opus outperforms GPT-4 in generating human-like dialog and contextually aware responses. These rapid advances have, in turn, led the creative industries to face significant challenges. For example, DMG Media, the Financial Times, and Guardian Media Group have highlighted concerns about the potential impact on print journalism, particularly if AI tools reduce the need for users to click through to news websites, affecting advertising and subscription revenues Communications and Digital Committee (2024). There is also concern about ‘AI-generated slop’—low-quality, mass-produced content created by AI that often lacks coherence or originality[^4]. It is typically used for spam, search engines, or clickbait, and is criticized for cluttering the internet and undermining genuine human-created content.

The generation of videos is significantly more challenging for AI than generating images. In February 2024, Google announced Gemini 1.5, which had the capability to process approximately 8 times more data than GPT-4 ; this comparison refers to a 1M-token context window versus 128k tokens. This opening opportunities for video and audio processing[^5]. In the same month, OpenAI provided its first preview of Sora, a model capable of generating impressive realistic videos up to 1 minute long. Based on the videos released by OpenAI, Sora appears to outperform other text-to-video models. Sora is currently available to ChatGPT subscribers , but new Plus accounts may face a wait-list or monthly credit limits. A month later, Gemini 1.5 announced its support for native audio understanding in 180$+$ countries. With the emergence of these tools, together with the prospect of further advances, it is clear that video content creation will be a major beneficiary. This will further open up the media landscape for creativity and provide more opportunities for diverse storytellers, while also reducing production time. A recent example is the AI-generated Christmas commercial by Coca-Cola[^6] , which elicited mixed responses in the press Yang (2025), thereby pre-empting questions regarding potential hype. Such advertisements overcome the limitations of current technologies by using very short videos with rapid scene transitions, ensuring that any artifacts, such as unnatural fingers, are less apparent.

For the case of post-production workflows, generative AI may not have a direct impact, but the neural networks originally proposed for generative AI have been widely adapted to serve this purpose. This has led to significant improvements in both output quality and computational speed. Moreover, there is a noticeable trend towards adopting a unified framework rather than addressing individual tasks, as it better reflects real-world scenarios. For instance, natural history filmmaking involves challenging acquisition environments and high production standards. Filming often takes place in low light conditions, in the presence of heat haze, underwater or in adverse weather conditions. This often results in increased noise levels, focus issues, low contrast, color balance problems, and blurriness in the footage. In such cases, unified models can offer advantages in generalizing to diverse tasks and providing flexibility. Take Painter by BAAI Vision Wang et al. (2023c) as an example, which employs an image pair as a task prompt (similar to a text prompt in LLMs), their model transfers the input image to produce a similar output as the task prompt, enabling it to undertake various tasks such as segmentation, low-light enhancement or rain removal.

While generative AI can facilitate and accelerate the creation and post-processing of digital media, there is an equivalent need to transmit or stream it efficiently to users. Although AI-based solutions have been proposed both for the enhancement of conventional video coding tools and for new compression frameworks, they are yet to be adopted in practical applications due to hardware constraints, complexity issues and a lack of standardization. Despite this, the latest learning-based video codecs have already demonstrated their potential to compete with conventional standardized video codecs and are being actively investigated in various standards bodies such as MPEG and AOM.

Furthermore, in recent years, AI has also impacted our ability to assess and monitor the perceptual quality of visual media. Advances have included new model architectures based on different attention mechanisms and the application of LLMs, which evidently improve model generalization. New training methodologies have also been proposed based on weakly/unsupervised learning, which address issues associated with the limited availability of labeled training content.

One of the exciting aspects of using LLMs in the creative sector is that ‘The human in the loop’ Chung (2021) is simplified through text prompts, with sophisticated, multilingual language capabilities enabling artists to convey complex emotions and narratives. This is important because generative AI does produce mistakes, known as hallucinations. Human oversight is thus essential to correct this through reinforcement learning with feedback Wu et al. (2023e).

In this paper, the objective is to present the major technological advancements that have emerged since our previous review on AI in the creative industries (2022) Anantrasirichai and Bull (2022). Whereas the earlier paper was written at a time when AI tools primarily served supportive roles, this updated review captures the disruptive shifts driven by generative AI and related technologies over the past few years. Adopting an integrative narrative rather than a systematic approach, we aim to synthesize major developments and identify cross-domain themes, rather than conduct an exhaustive quantitative survey. Sources published between 2022 and mid-2025 were selected from leading AI conferences (CVPR, ICCV, ECCV, NeurIPS, ICLR, IEEE TPAMI, IEEE TIP), verified arXiv preprints highlighting emerging developments, and industry releases, based on their technical significance and impact (e.g., OpenAI, Google DeepMind, Meta, Adobe). Selection was guided by recency, originality, and relevance to creative-industry applications, ensuring representative coverage of key models and trends. In summary, we initially screened more than 1,000 documents published between 2022 and mid-2025, including major AI conference papers ($\approx$60%), journal articles ($\approx$30%), and industry or policy reports ($\approx$10%). From these, around 450 sources were retained for detailed analysis across creation, enhancement, immersive media, and AI-driven production workflows. The review first outlines recent advances in AI technologies (Section 2), then highlights their applications in creative domains (Section 3), and finally discusses emerging challenges and future directions (Section 4).