The need for segmentation has grown dramatically in the past few years, given its central role in visual perception. Many segmentation methods now integrate an input prompt for users to define their preferred output appearances, such as pixel-wise segmentation, bounding boxes around objects, or segmented areas of interest. Most of these methods utilize transformer architectures Cheng et al. (2022). Among them, Segment Anything (SAM) by Meta AI Kirillov et al. (2023) stands out as a pioneer in promptable segmentation approaches. This method computes masks in real-time and has been trained with over 1 billion masks across 11 million images, facilitating transferability from zero-shot to new image distributions and tasks. HQ-SAM Ke et al. (2023) enhances SAM by incorporating global-local feature fusion, leading to high-quality mask predictions. SegGPT Wang et al. (2023d) proposed context ensemble strategies and allows users to tune a prompt for a specific dataset, scene, or even a person, while SEEM Zou et al. (2023a) provides a completely promptable and interactive segmentation interface. More recently, SAM2 Ravi et al. (2024) introduced support for real-time video segmentation. It is a unified model trained on a larger dataset than SAM. Interactive tools enable users to mark areas of interest and specify regions to exclude from the segmentation map. Zhou et al. propose an audio-visual segmentation (AVS) to generate pixel-level segmentation masks for sounding objects in audible videos. DVIS++ by Zhang et al. (2025b) introduces a universal video segmentation framework capable of producing instance, semantic, and panoptic segmentation outputs. This transformer-based architecture comprises a segmentor, tracker, and refinement module, achieving state-of-the-art performance across several video segmentation benchmarks.
With DMs tehchnologies, Baranchuk et al. Baranchuk et al. (2022) have investigated semantic representation, and found DMs outperform other few-shot learning approaches. DiffuMask Wu et al. (2023g) automatically generate image and pixel-level semantic annotation using pre-trained Stable Diffusion with input as a text prompt. It has been proven that using these synthetic data improve segmentation accuracy. Currently, the state-of-the-art panoptic segmentation is the method developed by Nvdia, which is based on text-to-image DMs Xu et al. (2023c), outperforming the previous methods by up to 7.6%.
Applying INRs to segmentation is more popular in the medical domain, as the specific signals used, such as computed tomography (CT) and magnetic resonance imaging (MRI), can be formulated as continuous functions. In creative technologies, unsupervised domain adaptation (UDA) and INRs are used for continuous rectification function modeling in Gong et al. (2023), achieving superior segmentation results in night vision. Recently, this work has been integrated with a non-local means block in Lin et al. (2025) showning a significant improvement for instant segmentation in low-light scenes.
3D segmentation is also crucial for scene manipulation. In radiance fields, earlier segmentation methods required additional modules such as using k-means clustering to separate objects from the background Goel et al. (2023). However, the recent SA3D approach Cen et al. (2023) segments 3D objects using NeRFs as the structural prior. SA3D operates by taking a trained NeRF and a set of prompts from a single view, then performing an iterative procedure. This involves rendering novel 2D views, self-prompting SAM for 2D segmentation, and projecting the segmentation back onto 3D mask grids. A comprehensive survey of 3D segmentation in computer vision can be found in He et al. (2025b).
3.4.2 Detection and recognition
Introduced in 2020, DETR by Facebook AI Carion et al. (2020) was one of the first to adopt a transformer architecture for object detection. The approach achieves comparable results to an optimized Faster R-CNN Ren et al. (2017), introduced in 2015. Deformable convolution has alson been used, (Deformable DETR Zhu et al. (2021)), resulting in training faster with approximately 5% accuracy improvement. A survey until 2022 Zou et al. (2023b) reported that Deformable DETR and Swin Transformers Liu et al. (2021) outperform pure CNN-based YOLOv4 Bochkovskiy et al. (2020). SwinV2 improves the first version by replacing original dot product attention with scaled cosine attention, improving accuracy by approximately 5%. Later, RT-DETR Zhao et al. (2024c) improved inference speed by decoupling the intra-scale interaction and cross-scale fusion of features with different scales. RT-DETR is 25% faster than YOLOv8[^46] with 6% improvement on MS COCO Object Detection dataset. Recently, YOLOv10 Wang et al. (2024a) has been released. YOLOv10 further improves the speed of detection approximately by 30% over RT-DETR with the same accuracy. A review of transformer-based methods for object detection can be found in Li et al. (2023d); Kheddar (2025). Recently, YOLO12 Tian et al. (2025) introduced an attention-centric architecture, achieving a 2.1% and 1.2% mAP improvement over YOLOv10-N and YOLOv11-N respectively, with only a slight decrease in speed.
To detect 3D objects, the transformer-based method MonoDTR Huang et al. (2022a) incorporates depth estimation from a single 2D image Yang et al. (2024b) to predict 3D bounding boxes. More 3D object detection methods have been developed for autonomous driving Song et al. (2024); however, these approaches can also be adapted for AR and VR applications Im et al. (2025).
While DMs are primarily used to generate synthetic datasets Wu et al. (2023f); Fang et al. (2024), they have also been demonstrated to function as zero-shot classifiers by Li et al. Li et al. (2023a). DMs are also of interest for detection tasks, Although feature extractors are still predominantly based on CNNs, such as ResNet, or Transformers (like Swin). DiffusionDet Chen et al. (2023a) formulates object detection as a denoising diffusion process from noisy boxes to object boxes, reporting performance that surpasses DETR. DMs have also been employed for anomaly detection Zhang et al. (2025a); Wu et al. (2025b), functioning similarly to zero-shot classifiers.
3.4.3 Tracking
Object tracking stands out as one of the tasks that greatly benefits from transformers since attention is needed in both space and time. An experimental survey cited in Kugarajeevan et al. (2023) reveals that transformer-based methods consistently rank at the top of the leaderboard across various datasets. In the Visual Object Tracking (VOT) challenges of 2023[^47], all of the top-10 employed transformer-based methodologies. The highest-performing approach achieved a 10% improvement in tracking quality compared to the winner in 2020. The current state-of-the-art for single-object tracking[^48], however, is based on cross-attention and Mamba Kang et al. (2025).
The first three tracking-by-attention approaches are TrackFormer Meinhardt et al. (2022), MixFormer Cui et al. (2022), and ToMP Mayer et al. (2022). TrackFormer extracts visual features using a CNN-based encoder, which are then tracked using a vanilla transformer Vaswani et al. (2017) in a frame sequence, while MixFormer introduces cross-attention between the target and search regions. ToMP tracks the objects using prediction aspects. Many more methods have been proposed, including SeqTrack Chen et al. (2023c) and Track Anything Model (TAM) Yang et al. (2023e). SeqTrack extracts visual features with a bidirectional transformer, while the decoder generates a sequence of bounding box values autoregressively with a causal transformer. TAM combines SAM Kirillov et al. (2023) and XMem Cheng and Schwing (2022), offering tracking and segmentation performance on the human-selected target. However, the masked area is still not very sharp, and there is a subtle degree of temporal inconsistency. MOTRv2 Zhang et al. (2023c) combines YOLOX Ge et al. (2021) for object recognition and MOTR Zeng et al. (2022) for tracking, outperforming TrackFormer by 20%. Additionally, some methods have been specifically proposed for challenging environments, such as low light Yi and Anantrasirichai (2024) and small objects, as seen in AnyFlow Jung et al. (2023). The latter exploits INR to upsample a continuous coordinate-based flow map, similar to SISR technique proposed in Chen et al. (2021b).
Similarly to detection tasks, DMs for tracking tasks are used as downstream processes by concatenating the diffusion head to the feature extraction backbone. However, a spatial-temporal fusion module has been added to the diffusion head to exploit temporal video featuresLuo et al. (2024). DiffusionTrack Xie et al. (2024) localizes the target in a progressive diffusion manner, which is claimed to better handle challenging scenarios. The method in Zhang et al. (2024b) exploits spatial-temporal weighting to suppress the probability of the tracker changing the target to the distractors. It, however, reports underperformance compared to MixFormer.
3.5 3D Reconstruction and Rendering
Bridging the gap between digital and physical realms, 3D reconstruction and rendering are integral to various creative technologies. In film and animation, they enable the creation of detailed digital models that blend seamlessly with live-action footage. Video games and digital twins leverage these technologies for dynamic environmental rendering. VR and AR use 3D reconstruction to create immersive and interactive experiences, with AR integrating digital content into real-world contexts. With recent AI technologies, 3D reconstruction and rendering have become faster and closer to reality. In particular, neural radiance fields and Gaussian Splatting enable artists and film producers to create shots that cannot be done in the real shooting environments.
3.5.1 Depth Estimation
Accurate depth information (alongside texture data) is typically required to construct 3D models. Depth sensors, such as lidar (Light Detection and Ranging) and structured-light 3D scanners, can be used for this purpose, but their applications are often limited by distance and cost. Consequently, vision-based sensors have become widely used. These sensors utilize two or more cameras to simulate human binocular vision or employ a single camera to capture images from different locations.
As deep learning can capture monocular cues such as object size, texture gradients, and perspective, depth estimation from a single image can produce accurate results. There have been attempts to use transformers, such as Zhang et al. (2023b) and Chen et al. (2023e), and diffusion models, such as Ji et al. (2023) and Ke et al. (2024). Amongst these, Depth Anything v2 Yang et al. (2024c) has become a state-of-the-art monocular depth estimation method. It is built on the previous version Yang et al. (2024b), jointly trained on large-scale labeled and unlabeled images and uses semantic priors from pretrained encoders. Depth Anything v2 significantly outperforms V1 in fine-grained details and robustness by using synthetic images and pseudo-labeled real images, as well as by extracting intermediate features from DINOv2 Oquab et al. (2024), which is trained with vision transformers. One of the notable capabilities of Depth Anything v2 is its ability to predict the depth of transparent and reflective surfaces.
3.5.2 Neural Radiance Fields
Neural Radiance Fields (NeRFs), introduced in Mildenhall et al. (2020), have demonstrated the ability to learn a 3D scene from a smaller number of images captured from various viewpoints, as opposed to photogrammetry. They excel in neural rendering, particularly in view-dependent novel view synthesis, and have effectively tackled several challenges associated with automated 3D capture Xie et al. (2022), such as accurately representing the reflectance properties of the scene. NeRFs offer high-resolution photo-realistic novel views and flexibility in postprocessing. They have hence gained significant attention in cinematography Azzarelli et al. (2024), as they offer reduced time and cost, particularly for outdoor shooting.
In the NeRF process (see Fig. 7 (a)), the camera positions and orientations are typically estimated from a series of 2D images using techniques like feature-mapping and Structure-from-Motion (SfM), as demonstrated in Schönberger and Frahm (2016). Leveraging INR, each image (or camera pose) is mapped into camera rays that traverse the scene, generating 3D points with directional radiance (towards the camera). These points are then processed by an MLP to predict volume density and emitted radiance. Subsequently, volume rendering techniques are employed to generate an image, which is compared with the original via loss calculation. The MLP iteratively refines the model by minimizing this loss.
Since their introduction, there have been many variants of NeRFs aimed at improving their performance. Mip-NeRF360 Barron et al. (2022) proposed an unbounded anti-aliased technique achieving full 360 degree content. Google Research Mildenhall et al. (2022) trains NeRF from noisy RAW images captured in the dark scene, allowing changing viewpoint, focus, exposure, and tone mapping simultaneously. With segmentation techniques significantly advanced (see Section 3.4.1), there have been integrations utilizing semantic segmentation to enhance 3D representation Guo et al. (2022a). DSEM-NeR Liu et al. (2025a) integrates the pretrained CLIP model to extract multimodal features—including color, depth, and semantics—from multi-view 2D images, thereby enhancing the reconstruction quality of complex scenes.
While the rendering quality of NeRF is very good, training and rendering times remain extremely high. The Instant-NGP tool developed by Nvidia Müller et al. (2022) enables real-time training of NeRFs by bypassing sampling in empty spaces and dense areas, and by incorporating multi-resolution hash encoding techniques. These advancements substantially reduce the computational burden associated with representing high-resolution image features -- training times have been reduced from hours to just a few seconds. Moreover, it offers VR controls for immersive 3D rendering experiences using OpenXR[^49]. This allows users to navigate scenes, manipulate objects, and interact with the environment directly through VR headsets. Diffusion models are integrated to regularize NeRF reconstructions Wynn and Turmukhambetov (2023), resulting in smoother depth continuity and clearer edges where depth discontinuities occur.
The initial application of NeRFs to dynamic scenes was undertaken by Pumarola et al. Pumarola et al. (2020), known as D-NeRF. However, the current leading method for generating high-quality novel views of real dynamic scenes is TiNeuVox Fang et al. (2022). It enhances temporal information by interpolating voxel features before feeding them into the radiance network to estimate density and color, similar to ordinary NeRF. DynVideo-E Liu et al. (2024) adds an MLP to predict motion fields but focuses on human-centric content. PaReNeRF Tang et al. (2024b) addresses large-scale dynamic scenes using patch-based sampling. The main drawback of these methods is the large model size and/or long training time. Therefore, $K$-planes Sara Fridovich-Keil and Giacomo Meanti et al. (2023) propose a simple planar factorization for volumetric rendering, achieving low memory usage (1000$\times$ compression over a full 4D grid). Wavelet transforms are employed in Azzarelli et al. (2023) to further reduce model size. KFD-NeRF Zhan et al. (2024) incorporates a Kalman filter-guided deformation field for more accurate motion estimation.

Figure 7: 3D representation. (a) Neural Radiance Fields (NeRFs) Mildenhall et al. (2020). (b) Gaussian Splatting Kerbl et al. (2023). (c) Example scenes of VR-GS system for 3D content interaction in VR Jiang et al. (2024).
3.5.3 3D Gaussian Splatting
The main issue with NeRFs as a method to generate high-quality novel views is training time, which can exceed a day for high-resolution content on a single RTX 3090 GPU Hu et al. (2022). 3D Gaussian Splatting (3D-GS) Kerbl et al. (2023) has been introduced to address this, using anisotropic 3D Gaussians to form a high-quality, unstructured representation of radiance fields. The process estimates a sparse point cloud through SfM. Each point possesses 3D Gaussian properties, such as position, covariance matrix, opacity, and spherical harmonics coefficients representing colors. The optimization of these parameters is interleaved with steps that control the density of the Gaussians to better represent the scene, as shown in Fig. 7 (b). A survey of 3D-GS can be found in Fei et al. (2024).
In contrast to traditional NeRFs based on implicit scene representations, 3D-GS provides an explicit representation that can be seamlessly integrated with post-processing manipulations, such as animating and editing. VR-GS Jiang et al. (2024) offers intuitive and interactive physics-based game-play with deformable virtual objects and realistic environments represented with 3D-GS. The example scenes are shown in Fig. 7 (c). Physics-inspired approaches are also integrated to improve 3D modeling in different media, such as 3D underwater scenes Wang et al. (2025a).
For dynamic scenes, 4D Gaussian Splatting (4D-GS) Wu et al. (2024a) introduces a Gaussian deformation field for motion and shape. It exploits a multi-resolution encoding method, achieving real-time rendering of up to 82 fps at a resolution of 800$\times$800 pixels on an RTX 3090 GPU. Instead of developing in 4D, CoGS Yu et al. (2024) exploits 3D-GS by integrating control mechanisms in separate regions to learn individual temporal dimensions. SC-GS Huang et al. (2024b) extracts sparse control points and uses an MLP to predict time-varying 6 DoF transformations. While the results show better visual quality than 4D-GS and CoGS, the performance heavily relies on camera pose estimation. Kong et al. Kong et al. (2025) represent dynamic scenes using sparse, time-variant attribute modeling with a deformable MLP, while efficiently filtering out anchors corresponding to static regions. Their model achieves fast rendering speeds of over 110 FPS at a resolution of 960$\times$540—nearly 10 times faster than SC-GS—and delivers a 1 dB improvement in PSNR.
LUMA AI[^50] and Polycam[^51] offer free tools for Gaussian splatting and photogrammetry creation for non-commercial use. The 3D objects created can be experienced with VR headsets for more immersive 3D and further used or developed in other applications. However, these tools have limitations in handling dynamic scenes due to occlusions, sparse observations per timestamp, and object reappearances over time. Rendering dynamic avatars can produce higher quality results by incorporating additional information. For example, EVA Junkawitsch et al. (2025) disentangles the 3D Gaussian appearance into skeletal motion, facial expressions, body movements, and skin. These components are then splatted to render the final photorealistic image.
3.5.4 Digital Twins
A digital twin is a virtual replica of a physical object, system, or process, continuously updated with real-time data for purposes such as simulation, testing, monitoring, and maintenance. This technology is increasingly adopted across various applications within the creative industries. For example, in product design and branding, it enables immediate observation of how a design performs in various contexts, facilitating the development of user-friendly products. Unilever reported that integrating digital product twins with 3D technologies, such as NVIDIA Omniverse[^52], enabled the creation of product imagery twice as fast and 50% more cost-effective[^53]. Digital twins also allow consumers to explore products or spaces virtually, simulating real-world interactions.
Accenture plc, a global professional services company, collaborated with Walt Disney Studios to develop digital twin technologies aimed at transforming the filmmaking process[^54]. Their goal is to generate remotely accessible 3D models, enabling virtual exploration of potential shooting locations without requiring physical visits. The Virtual StudioLAB provides a digital replica created using 360-degree imagery and 3D modeling. These innovations have streamlined pre-production workflows for major productions from Marvel Studios and 20th Century Studios.
Digital representations such as avatars, proxies, and digital twins are increasingly being explored in artistic contexts, particularly in relation to identity, presence, and embodiment in virtual environments. The Tate Modern’s film programme Avatars, Proxies and Digital Twins (Feb–May 2025) investigated these themes through curated audiovisual works, offering critical reflections on digital personhood. By engaging with diverse narrative forms, the programme highlighted the sociocultural implications of digital self-representation, prompting discourse on authenticity, agency, and the role of immersive media in shaping future human–machine interaction.
3.6 Data Compression
Data compression plays an important role in the delivery of creative content to audiences, effectively reducing memory and bandwidth requirements during signal storage and transmission Bull and Zhang (2021). Although coding methods based on conventional signal processing theories are still widely employed in most standards and application scenarios, learning-based solutions have emerged in research, showing great potential to achieve competitive performance in recent years. This subsection provides a brief overview of the recent advances in image, video, and audio compression, in particular focusing on the approaches proposed after 2021.
3.6.1 Image Compression
Since the first neural image codec Ballé et al. (2016) was proposed in 2016, numerous learning-based image compression methods have been developed, with significant performance improvements reported Ballé et al. (2018); Cheng et al. (2020). Driven by the latest advances in neural network architectures, neural image codecs now outperform standard image codecs. Instead of using CNNs as the basic network structure, transformer-based architectures have become popular, offering the potential for better compression efficiency. Notable examples include SwinT-ChARM Zhu et al. (2022), STF Zou et al. (2022) and LIC-TCM Liu et al. (2023a). SwinT-ChARM Zhu et al. (2022) employs Swin transformers for non-linear transforms and outperforms the latest standard image codec, the Versatile Video Coding (VVC) Test Model (VTM, All Intra). STF Zou et al. (2022) is based on a symmetrical transformer framework containing absolute transformer blocks in both the down-sampling encoder and the up-sampling decoder, which also shows improved rate-quality performance over VTM. LIC-TCM Liu et al. (2023a) exploits the local modeling ability of CNN and the non-local modeling performance of transformers, and proposes a parallel transformer-CNN mixture block. This new network structure, together with a channel-wise entropy model based on attention modules using Swin transformers, contributes to the superior performance of STF, with more than 10% bitrate savings over VTM.
An alternative approach to learned image coding is based on advanced generative models. Early works Agustsson et al. (2019); Mentzer et al. (2020) employed GANs to generate more photo-realistic results with improved visual quality. Although these models fail to outperform conventional, CNN-based or transformer-based approaches, when distortion-based quality metrics, e.g., PSNR, are used for performance evaluation, they have been reported to perform well when perceptual quality models, such as MS-SSIM Wang et al. (2003a) and VMAF Li et al. (), or subjective tests are employed to measure perceived video quality. More recently, diffusion models have been applied in image compression to allow realistic reconstruction at ultra-low bitrates Careil et al. (2023) achieving competitive performance compared to GAN-based models Yang and Mandt (2023). However, it should be noted that some of these generative models aim to generate (or synthesize) images with “perfect realism” rather than reconstruct results that are most similar to the original content. Notable work in this category includes image codecs using score-based generative models Hoogeboom et al. (2023) and the diffusion-based residual augmentation codec (DIRAC) Ghouse et al. (2023). Moreover, another type of generative model based on INR has been employed for image compression; this learns a mapping between the spatial coordinates and the respective pixel values for the input image. The learned INR model is then compressed through parameter quantization and model compression to minimize the required bitrate. Notable INR-based image codecs include COIN/COIN++ Dupont et al. (2021); Dupont et al. (2022) and Strümpler et al. (2022) that combine SIREN networks Sitzmann et al. (2020a) with positional encoding.
In order to evaluate and compare neural image codecs under fair test conditions, public grand challenges have been increasingly run, typically associated with international conferences. One of the most well-known of these is the Challenge on Learned Image Compression (CLIC) [1]. In its latest competition, the best performing learned image codec Li et al. (2024b), which is based on a GAN-enhanced Vector Quantized Variational AutoEncoder (VQ-VAE) framework, offered up to 0.6dB PSNR gain over VTM (version 22.2, All Intra) at similar bitrates; this codec is based on an autoencoder architecture with latent refinement and perceptual losses.
To support the deployment of neural image codecs, the International Organization for Standardization (ISO)/International Electrotechnical Commission(IEC) has developed a royalty-free learned image coding standard, denoted as JPEG AI Ascenso et al. (2023), which aims to offer significant performance improvement over existing standards for both human and machine vision tasks. The Call for Proposals of JPEG AI was published in 2022, while the Working Draft and the Committee Draft outlining its core coding system were released in 2023 Alshina et al. (2024), with its first version published in October 2024 Alshina et al. (2024). JPEG AI follows the same framework (the auto-encoder structure) as most existing neural image codecs, and its test model JPEG AI VM (version 4.3) has been reported to achieve up to 28.5% coding gains over VVC VTM (All Intra mode) Karabutov et al. (2024).
3.6.2 Video Compression
Compared to image coding, the compression of video content is a much more challenging task, particularly for immersive video formats and diverse content types. Although video coding standards including H.264/AVC (Advanced Video Coding), H.265/HEVC (High Efficiency Video Coding) and H.266/VVC (Versatile Video Coding) are still predominant in real-world applications, learning-based video coding has advanced dramatically in the past five years, with new deep learning enhanced conventional coding tools and end-to-end optimized neural video coding frameworks proposed.
i) The enhancement of conventional coding tools focuses on employing deep learning techniques to improve the performance of one or more coding modules in a standard-compliant codec. These modules include intra prediction Li et al. (2021b), inter prediction Jin et al. (2021), in-loop filtering Feng et al. (2024c), post-decode filtering Zhang et al. (2023a) and resolution re-sampling Wang et al. (2023e). To facilitate efficient integration, the MPEG Joint Video Experts Team (JVET) built a test model in 2022 based on VTM 11, named Neural Network-based Video Coding (NNVC) Li et al. (2023e), with its latest version NNVC-7.1 containing two major learning-based coding tools, neural-network based intra prediction and in-loop filtering, which has achieved an up to 13% coding gain over VTM 11 (Random Access mode) Galpin et al. (2024). However, this learning-based codec requires much higher computational complexity (up to 477 kMACs/pixel) and high-spec GPU support compared to conventional codecs. Meanwhile, members of the Alliance of Open Media (AOM) have also developed multiple CNN-based coding tools for the next generation of video coding standard beyond AV1. The latest proposals focus on the trade-off between performance and complexity, with one of them based on in-loop filtering and super-resolution, which achieves an average BD-rate saving of 3.9% (in PSNR) over AVM, the test model of AV2, but only requires a much lower computational complexity (below 1.5kMACs/pixel) Joshi et al. (2023). More recently, research has been conducted to further improve the performance of these learning-based coding tools utilizing more advanced network architectures, including ViTs Kathariya et al. (2023), and diffusion models Li et al. (2024a). There are also investigations on applying preprocessing before compression Chadha and Andreopoulos (2021); Tan et al. (2024), where the training of the deep preprocessors is based on proxy video codecs and/or rate-distortion loss functions to simulate the behavior of conventional video coding algorithms.
ii) End-to-end optimized neural video codecs. Alongside the enhancement of coding tools in conventional video codecs, more recent research activities have focused on using neural networks to implement the whole coding workflow, enabling data-driven end-to-end optimization. The performance of these neural video codecs has advanced significantly in the last five years, since the first attempt, DVC Lu et al. (2019), was published. DVC matched the performance of a fast implementation of H.264 (x264). However currently, learned video coding algorithms (e.g., DCVC-FM Li et al. (2024c) and DCVC-LCG Qi et al. (2024)) are able to compete or even outperform the state-of-the-art standard codecs, such as VVC VTM under certain coding configurations. These learning-based methods often focus on enhancement from different perspectives, including feature space conditional coding (e.g., FVC Hu et al. (2021) and DCVC Li et al. (2021a)), instance adaptation Khani et al. (2021); Oh et al. (2024), and motion estimation (e.g., DCVC-DC Li et al. (2023b)). New architectures have also been proposed such as CANF-VC Ho et al. (2022) based on a video generative model, MTMT Xiang et al. (2022) using a masked image modeling transformer-based entropy model and VCT Mentzer et al. (2022) based on a video compression transformer. It is noted that, although promising coding performance has been achieved by the aforementioned contributions, these neural video codecs (in particular those based on autoencoder backbones) are typically associated with high computational complexity (especially in the decoder), which constrains their deployment for practical applications. To address this issue, researchers are now focused on complexity reduction while maintaining coding performance through model pruning and knowledge distillation Guo-Hua et al. (2023); Peng et al. (2024b).
It should be noted that the neural codecs mentioned above are typically trained offline with diverse video content Nawała et al. (2024), and deployed online for inference. In this case, model generalization becomes important, and this is why these codecs often have a large model capacity, resulting in large model sizes and slow inference runtime. Inspired by recent advances in implicit neural representations (INR), a new type of video codec has emerged that employs INR models to “represent” the video by learning a coordinate-based mapping and compressing the network parameters for transmission. This approach converts a video coding problem into a model compression task, which allows the use of a much smaller network to “overfit” the input video, with the real potential for fast decoding. Existing implicit neural video representation (NeRV) models can be classified into index-based and content-based methods. The former takes frame Chen et al. (2021a), patch Bai et al. (2023) or disentangled spatial/grid coordinates Li et al. (2022d) as model input, while content-based approaches Kwan et al. (2024a); Kim et al. (2024a); Leguay et al. (2024) have content-specific embedding as inputs. Currently, one of the best INR-based video codecs Kwan et al. (2024b) has already achieved a performance similar to that of VVC VTM (RA), but with a much lower decoding complexity compared to autoencoder-based neural codecs. Some of these models have also been applied to volumetric video content Ruan et al. (2024); Kwan et al. (2024c), demonstrating their potential to compete with standard and other learning-based methods. However, it should be noted that the training of most NeRV models is based on an entire video sequence or even datasets; this results in a high system delay and does not meet the requirement of many low latency video streaming or real-time applications. To address this limitation, significant advances have been made Gao et al. (2024) towards more practical INR-based video compression (such as the Low Delay and Random Access modes in VVC VTM Bossen et al. (2023)) by combining pre-training and online model overfitting.
Similarly to image compression, international grand challenges are used to compare neural video compression methods, with notable venues including the NN-based Video Coding Grand Challenge associated with The IEEE International Symposium on Circuits and Systems (ISCAS) and the Challenge on Learned Image Compression (CLIC, video coding track) with IEEE/CVF CVPR and Data Compression Conference (in 2024). The best performer in ISCAS 2024 NN-based Video Coding Grand Challenge offers an overall 55% BD-rate saving over HEVC Test Model HM [129], while the winner of the CLIC (video coding track) in 2024, a neural-network enhanced ECM codec Zhao et al. (2024b) with a CNN-based in-loop filter, shows a more than 2dB (in PSNR) gain compared to VTM (RA) at the same bitrates.
3.6.3 Audio Compression
Similarly to images and videos, learning-based solutions have also been researched to compress audio signals, and most neural audio codecs are based on VQ-VAE van den Oord et al. (2017). SoundStream Zeghidour et al. (2021) is one such model, which can encode audio content at various bitrates. It is based on a residual vector quantizer (RVQ) which trades off rate, distortion and complexity. This work has been further enhanced with a multi-scale spectrogram adversary and a loss balancer mechanism, resulting in improved rate-distortion performance. A more advanced universal model has been further developed Kumar et al. (2023) based on improved adversarial and reconstruction losses, which can compress different types of audio. RVQ has also been extended from a single scale to multiple scales Siuzdak et al. (2024), which performs hierarchical quantization at variable frame rates.
More recently, researchers have started to exploit the use of LLMs for audio compression, leveraging the audio generation/synthesis abilities of generative models. UniAudio 1.5 Yang et al. (2024a) is one of such attempts, which converts an audio into the textural space, which can be represented by a pre-trained LLM that shares a similar backbone of UniAudio Yang et al. (2023b), a universal audio foundation model. LFSC is another neural audio codec based on LLMs, which achieved fast LLM training and inference through finite scalar quantization and adversarial training.
3.7 Visual Quality Assessment
Assessing the quality of visual signals remains an important and challenging task for many image and video processing applications. While subjective tests involving human participants remain the gold standard, objective quality models are frequently used because of their time and cost efficiency. These quality assessment methods are typically used to evaluate the performance of different visual processing approaches, and they can also be converted to loss functions, which are employed for optimizing learning-based processing models.
In recent years, quality assessment methods have been enhanced using deep learning techniques. The resulting learning-based quality models can quickly adapt to a specific type of content, leading to better performance compared to conventional, hand-crafted quality metrics. This section provides a brief summary of existing work in this research area, and highlights the main challenges which should be addressed in the near future. A more comprehensive overview of the image and video quality assessment literature can be found in Zhai and Min (2020); Zheng et al. (2024); Zhang et al. (2024d).
3.7.1 Quality assessment models
Image and video quality assessment methods can be classified into two primary categories according to the availability of the corresponding reference (un processed) content. These are referred to as full-reference and no-reference models[^55]. Prior to the AI era, conventional visual quality methods often exploited characteristics of the human vision system capturing information related to structural similarity (such as in SSIM and its variants Wang et al. (2004); Wang et al. (2003b); Rehman et al. (2015)), distortion Chandler and Hemami (2007); Larson and Chandler (2010); Vu et al. (2011), and artifacts Ou et al. (2010); Zhu et al. (2014); Zhang and Bull (2015). In many cases, the extracted features are further processed by models that simulate texture masking von Helmholtz (1896), contrast sensitivity Kelly (1977), and saliency Itti and Koch (2001). These hand-crafted quality models have also been combined with features within a regression-based framework in order to achieve more accurate prediction performance - VMAF is one such example Li et al. (). When neural networks are used for feature extraction, they are trained to capture information which can directly contribute to quality prediction through an end-to-end optimization strategy. Initially, convolutional neural networks were used for this, with notable examples such as DeepQA Kim and Lee (2017), LPIPS Zhang et al. (2018) and CONTRIQUE Madhusudana et al. (2022) for image quality assessment, and TLVQA Korhonen (2019), C3DVQA Xu et al. (2020) and DeepVQA Kim et al. (2018) for video quality assessment. Recent works have been reported to achieve better performance when Vision Transformers (ViTs) (or similar variants) are employed due to the effectiveness of their self-attention mechanism. Important works in this class include IQT Cheon et al. (2021), TRes Golestaneh et al. (2022), SaTQA Shi et al. (2024a), FastVQA Wu et al. (2022) and RankDVQA Feng et al. (2024a). The former has been further extended as DOVER Wu et al. (2023c) and COVER He et al. (2024) when aesthetic and/or semantic aspects in the content are taken into account.
More recently, inspired by the success of large language models (LLMs) OpenAI et al. (2023); Touvron et al. (2023) in other machine learning tasks, these have been utilized in image and video quality assessment, demonstrating significant potential to achieve better model generalization. Q-Bench Wu et al. (2024b) is one of the first attempts that employs multimodal large language models to predict the perceptual quality of images based on prompt-driven evaluation. It queries the LLMs to provide information related to the final quality rating of the input image and the quality description. This has been further extended for video quality assessment tasks in Q-Align Wu et al. (2024c). Other notable works include X-iqe Chen et al. (2023d) that performs the quality prompt in a multi-iteration manner focusing on both image fidelity and aesthetics. Prompt-based approaches have also been proposed for differentiating the quality difference between multiple images, such as 2AFC-LMMs Zhu et al. (2024a) based on a two-alternative forced choice prompt and MAP (maximum a posteriori) estimation. Moreover, recent research works also focus on using pre-trained vision-language models, such as CLIP Radford et al. (2021), which align better image and text modalities. Important examples in this class for image quality assessment include ZEN-IQA Miyata (2024), QA-CLIP Pan et al. (2023b) and PromptIQA Chen et al. (2025a). Similar works have also been proposed for video quality assessment, such as BVQI Wu et al. (2023a); Wu et al. (2023b) and COVER He et al. (2024).
To support the training and validation of learning-based quality assessment models, image or video databases containing ground-truth subjective quality scores are typically employed. Commonly used image quality databases include LIVE Sheikh et al. (2006), CSIQ Larson and Chandler (2010), TID2013 Ponomarenko et al. (2013), PieAPP and PIPAL, while video quality databases such as LIVE-VQA Seshadrinathan et al. (2010), KoNViD-1K Hosu et al. (2017), YouTube UGC Wang et al. (2019) and LIVE-VQC Sinno and Bovik (2018) are typically employed for benchmarking in the literature. There are also databases developed that investigate the impact of specific video formats and/or artifacts, such as LIVE-YT-HFR Madhusudana et al. (2021) focusing on frame rates, VSR-QAD Zhou et al. (2024) on spatial resolution (or super-resolution artifacts), BAND-2k Chen et al. (2024) on banding artifacts and Maxwell Wu et al. (2023b)/BVI-Artifact Feng et al. (2024b) containing multiple artifacts commonly produced in video streaming. Based on these databases, many learning-based quality assessment models are trained to minimize the difference (L1 or L2 norm) between predicted quality indices and subjective scores. However, due to the limited number of ground-truth quality labels associated with these databases and the resource requirements associated with collecting subjective data using human participants in psychophysical experiments, this type of training methodology cannot offer satisfactory performance, in particular when the model capacity is large. Moreover, since the experimental settings and conditions used for quality labeling are different in these databases, intra-database cross-validation is always required due to the limited model generalization and potential overfitting problems.
To address these issues, various proxy quality metrics have been used to label images and videos, which avoid expensive subjective tests and enable the generation of a large amounts of training material with pseudo-ground-truth quality annotations. To further improve the reliability of quality labels, instead of learning the absolute values of the quality labels, ranking-inspired training strategies have been developed, which focus on improving the monotonicity characteristics of quality. Important examples based on these weakly supervised training methodologies include RankIQA Liu et al. (2017) and UNIQUE Zhang et al. (2021) for the image quality assessment task, and VFIPS Hou et al. (2022) and RankDVQA Feng et al. (2024a) for video quality assessment. Moreover, different self-supervised learning approaches have also been employed, which transform quality labeling to an auxiliary task. For example, CONTRIQUE Madhusudana et al. (2022) learns relevant features from an unannotated image database based on the prediction of distortion types and degrees through contrastive learning. This method has been further applied to video quality assessment, resulting in a contrastive video quality estimator, CONVIQT Madhusudana et al. (2023). More recently, quality-aware contrastive loss has been designed in Zhao et al. (2023a); Peng et al. (2024a) to stabilize the learning process.
3.7.2 Performance and main challenges
Due to the lack of standard test conditions and limited model generalization within many existing image and video quality assessment models, deep compression methods are typically trained and benchmarked using different databases in conjunction with intra-database cross-validation. This can result in inconsistent evaluation results and conclusions. To enable a fair and meaningful comparison, various challenges and contests have been held for visual quality assessment. The Sixth Challenge on Learned Image Compression (CLIC) [1] associated with the Data Compression Conference 2024 is one of the latest examples which includes two quality assessment tracks for image and video compression. The best performer in the video quality assessment track achieves a Spearman Ranking Correlation Coefficient value of 0.825 Feng et al. (2024a), which is based on a ranking-inspired training methodology. Other notable challenges include the IEEE/CVF WACV 2023 HDR VQA Grand Challenge and the Video Super-Resolution Quality Assessment Challenge in ECCV 2024, which focus on high dynamic range and super-resolved content, respectively.
Although significant progress has been made in the past few years in visual quality assessment, including new models and training methodologies, challenges remain, including limited model generalization and high computational complexity. Another important use of quality metrics is as embedded loss functions for image and video processing optimization. This requires further capability and robustness, alongside complexity reduction, all topics to be addressed in future work.
3.8 Summary of AI technologies for creative industries
This section consolidates the preceding discussion by providing a comparative overview of the main classes of AI models shaping contemporary creative practice. While earlier sections examined individual technologies in detail, the following summary highlights how these models—ranging from large language models (LLMs) and diffusion models (DMs) to Neural Radiance Fields (NeRFs) and Implicit Neural Representations (INRs)—differ in application domains, key advantages, and persistent limitations. This synthesis enables readers from both creative and technical backgrounds to discern where each model type contributes most effectively, where challenges remain, and how these approaches collectively reshape workflows across the creative industries. Table 2 summarizes these relationships and serves as a conceptual reference for future research, mapping core model classes to their applications, strengths, and constraints to guide the evaluation and development of emerging AI methods in the creative industries.
| Technology | Core Mechanism | Creative Applications | Key Strengths | Main Limitations | Accessibility |
|---|---|---|---|---|---|
| Transformers/ Attention | Self-attention mechanisms capture global dependencies | Image/video restoration, super-resolution, segmentation, object detection, tracking | Parallel processing, long-range context, scalable, state-of-the-art performance | High memory usage, quadratic complexity, requires large training datasets | High: SwinIR, SAM, Restormer (open-source) |
| Large Language Models (LLMs) | Language understanding and generation | Text generation, dialogue systems, screenwriting, storyboarding, code assistance, script analysis | Highly flexible across domains; capable of reasoning and context understanding; supports creative ideation and natural language interaction | Hallucinations and factual errors; limited interpretability; potential bias from training data; copyright and authorship ambiguity | Moderate: GPT-4, Claude (APIs); LLaMA, Qwen (open weights) |
| Diffusion Models (DMs) | Iterative denoising from Gaussian noise | Image synthesis, Text-to-image/video generation, animation, VFX, design prototyping, restoration, style transfer, inpainting, restoration | High-quality diverse outputs; stable training; handles complex distributions; photorealistic; fine-grained control through prompts or conditioning inputs | Computationally expensive; slow sampling; temporal flicker in videos; difficulty ensuring semantic or style consistency | High: Stable Diffusion (open); DALL·E 3, Sora (APIs) |
| Neural Radiance Fields (NeRFs) | Volumetric scene representation via MLPs | 3D scene reconstruction, virtual production, spatial storytelling, AR/VR content creation | Photorealistic rendering from sparse views; compact scene representation; supports dynamic viewpoint changes | Limited to static or semi-static scenes; slow training; poor performance in textureless regions; large memory footprint | Moderate: Instant-NGP, Mip-NeRF (open); requires CV expertise |
| Implicit Neural Representations (INRs) | Coordinate-based continuous functions | Video/image compression; super-resolution; 3D reconstruction; dynamic scene modeling | Continuous scene representation; parameter-efficient encoding; smooth interpolation and compact storage | Challenging to generalize across scenes; training instability; limited semantic control | Moderate: NeRV, COIN (open); technical expertise required |
| Gaussian Splatting | Explicit 3D Gaussians for scene representation | Real-time novel view synthesis; 3D scene reconstruction; VR/AR;interactive rendering | Fast training (minutes) and rendering (real-time); explicit representation; easy manipulation | Lower quality than NeRF for complex scenes, requires SfM initialization, memory intensive | High: 3D-GS, 4D-GS (open); LUMA AI, Polycam (free tools) |
- Product and platform data (e.g., Sora, Gemini, Stable Diffusion 3) are accurate as of mid-2025 and may evolve rapidly.
Table 2: Comparative summary of key AI technologies for creative industries