- Challenge on Learned Image Compression. Note: https://www.compression.cc/Accessed: 2024-09-23 Cited by: §3.6.1, §3.7.2.
- A. Aggarwal Evolution of recommendation systems in the age of generative ai. International Journal of Science and Research Archive 14 (01), pp. 485–492. Note: https://doi.org/10.30574/ijsra.2025.14.1.0061 External Links: Document Cited by: §3.2.3.
- E. Agustsson, M. Tschannen, F. Mentzer, R. Timofte, and L. V. Gool Generative adversarial networks for extreme learned image compression. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 221–231. Cited by: §3.6.1.
- W. Ai, J. Li, Z. Wang, Y. Wei, T. Meng, and K. Li Contrastive multi-graph learning with neighbor hierarchical sifting for semi-supervised text classification. Expert Systems with Applications 266, pp. 125952. External Links: ISSN 0957-4174, Document Cited by: Table 1, §3.2.1.
- J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Bińkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems, Vol. 35, pp. 23716–23736. Cited by: Table 1, Table 1.
- E. Alshina, J. Ascenso, S. Esenlik, A. Karabutov, Y. Wu, and T. Solovyev JPEG AI: Future Plans and Timeline v2. JPEG AI ISO/IEC JTC 1/SC29/WG1 N1100634. Cited by: §3.6.1.
- N. Anantrasirichai and D. Bull Artificial intelligence in the creative industries: a review. Artificial Intelligence Review 55, pp. 589–656. External Links: Document Cited by: §1, §2.3, §2, §3.3.1, §3.3.4, §3.3.5, §3.
- N. Anantrasirichai, R. Lin, A. Malyugina, and D. Bull BVI-Lowlight: Fully registered benchmark dataset for low-light video enhancement. arXiv preprint arXiv:2402.01970. Cited by: §3.3.1.
- N. Anantrasirichai Atmospheric turbulence removal with complex-valued convolutional neural network. Pattern Recognition Letters 171, pp. 69–75. External Links: Document Cited by: §3.3.4.
- J. Ascenso, E. Alshina, and T. Ebrahimi The JPEG AI standard: providing efficient human and machine visual data consumption. IEEE Multimedia 30 (1), pp. 100–111. Cited by: §3.6.1.
- S. Azadi, A. Shah, T. Hayes, D. Parikh, and S. Gupta Make-An-Animation: Large-scale text-conditional 3D human motion generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 15039–15048. Cited by: Table 1, §3.1.4.
- A. Azzarelli, N. Anantrasirichai, and D. Bull Intelligent cinematography: a review of ai research for cinematographic production. Artificial Intelligence Review 58 (108). Cited by: §3.1.1.
- A. Azzarelli, N. Anantrasirichai, and D. R. Bull Exploring dynamic novel view synthesis technologies for cinematography. arXiv: 2412.17532. Cited by: §3.5.2.
- A. Azzarelli, N. Anantrasirichai, and D. R. Bull WavePlanes: A compact wavelet representation for dynamic neural radiance fields. arXiv preprint arXiv:2312.02218. Cited by: Table 1, §3.5.2.
- Y. Bai, C. Dong, C. Wang, and C. Yuan PS-NeRV: patch-wise stylized neural representations for videos. In IEEE International Conference on Image Processing, pp. 41–45. Cited by: Table 1, §3.6.2.
- Y. Bai, A. Jones, K. Ndousse, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862. Cited by: §1.
- J. Ballé, V. Laparra, and E. P. Simoncelli Density modeling of images using a generalized normalization transformation. In International Conference on Learning Representations (ICLR), Cited by: §3.6.1.
- J. Ballé, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston Variational image compression with a scale hyperprior. In International Conference on Learning Representations, Cited by: §3.6.1.
- D. Baranchuk, A. Voynov, I. Rubachev, V. Khrulkov, and A. Babenko Label-efficient semantic segmentation with diffusion models. In International Conference on Learning Representations, Cited by: §3.4.1.
- J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman Mip-NeRF 360: Unbounded anti-aliased neural radiance fields. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 5460–5469. External Links: Document Cited by: Table 1, Table 1, §3.5.2.
- C. Beckett and M. Yaseen Generating change: A global survey of what news organisations are doing with AI. Technical report JournalismAI. Cited by: §3.1.1.
- B. Bilecen and M. Ayazoglu Bicubic++: Slim, slimmer, slimmest designing an industry-grade super-resolution network. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vol. , pp. 1623–1632. External Links: ISSN , Document Cited by: §3.3.3.
- A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934. Cited by: §3.4.2.
- R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, et al. On the opportunities and risks of foundation models. arXiv:2108.07258. Note: https://crfm.stanford.edu/assets/report.pdf Cited by: §2.
- F. Bossen, J. Boyce, K. Suehring, X. Li, and V. Seregin VTM common test conditions and software reference configurations for SDR video. In the JVET meeting, Cited by: §3.6.2.
- T. Brooks, A. Holynski, and A. A. Efros InstructPix2Pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18392–18402. Cited by: Table 1, §3.1.3.
- E. Brynjolfsson, D. Li, and L. Raymond Generative AI at work. Technical report National Bureau of Economic Research. Cited by: §3.2.4.
- D. Bull and F. Zhang Intelligent image and video compression: communicating pictures. Academic Press. Cited by: §3.6.
- C. Cao, H. Yue, X. Liu, and J. Yang Zero-shot video restoration and enhancement using pre-trained image diffusion model. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 1935–1943. External Links: Document Cited by: Table 1, Table 1, §3.3.3.
- H. Cao, C. Tan, Z. Gao, Y. Xu, G. Chen, P. Heng, and S. Z. Li A survey on generative diffusion models. IEEE Transactions on Knowledge and Data Engineering (), pp. 1–20. External Links: Document Cited by: §2.3.
- M. Careil, M. J. Muckley, J. Verbeek, and S. Lathuilière Towards image compression with perfect realism at ultra-low bitrates. In The Twelfth International Conference on Learning Representations, Cited by: Table 1, §3.6.1.
- N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko End-to-end object detection with transformers. In Eur. Conf. Comput. Vis., pp. 213–229. Cited by: Table 1, §3.4.2.
- J. Cen, Z. Zhou, J. Fang, c. yang, W. Shen, L. Xie, D. Jiang, X. ZHANG, and Q. Tian Segment anything in 3D with NeRFs. In Advances in Neural Information Processing Systems, Vol. 36, pp. 25971–25990. Cited by: Table 1, §3.4.1.
- A. Chadha and Y. Andreopoulos Deep perceptual preprocessing for video coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14852–14861. Cited by: §3.6.2.
- W. Chai, X. Guo, G. Wang, and Y. Lu StableVideo: Text-driven consistency-aware diffusion video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 23040–23050. Cited by: Table 1, §3.3.2.
- K. C.K. Chan, S. Zhou, X. Xu, and C. C. Loy BasicVSR++: improving video super-resolution with enhanced propagation and alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5972–5981. Cited by: §3.3.1, §3.3.3.
- D. M. Chandler and S. S. Hemami VSNR: a wavelet-based visual signal-to-noise ratio for natural images. IEEE transactions on image processing 16 (9), pp. 2284–2298. Cited by: §3.7.1.
- Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, et al. A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol. 15 (3). External Links: Document Cited by: §2.2.
- H. Chen, B. He, H. Wang, Y. Ren, S. N. Lim, and A. Shrivastava NeRV: neural representations for videos. Advances in Neural Information Processing Systems 34, pp. 21557–21568. Cited by: Table 1, §3.6.2.
- S. Chen, P. Sun, Y. Song, and P. Luo DiffusionDet: diffusion model for object detection. In IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 19773–19786. External Links: Document Cited by: Table 1, §3.4.2.
- X. Chen, X. Wang, J. Zhou, Y. Qiao, and C. Dong Activating more pixels in image super-resolution transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22367–22377. Cited by: Table 1, §3.3.3.
- X. Chen, H. Peng, D. Wang, H. Lu, and H. Hu SeqTrack: sequence to sequence learning for visual object tracking. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 14572–14581. External Links: Document Cited by: Table 1, §3.4.3.
- Y. Chen, S. Liu, and X. Wang Learning continuous image representation with local implicit image function. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8624–8634. External Links: Document Cited by: Table 1, §3.4.3.
- Y. Chen, L. Liu, and C. Ding X-iqe: explainable image quality evaluation for text-to-image generation with visual large language models. arXiv preprint arXiv:2305.10843. Cited by: §3.7.1.
- Z. Chen, H. Qin, J. Wang, C. Yuan, B. Li, W. Hu, and L. Wang Promptiqa: boosting the performance and generalization for no-reference image quality assessment via prompts. In European Conference on Computer Vision, pp. 247–264. Cited by: §3.7.1.
- Z. Chen, Y. Duan, W. Wang, J. He, T. Lu, J. Dai, and Y. Qiao Vision transformer adapter for dense predictions. In Proceedings of the International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.5.1.
- Z. Chen, K. Liu, J. Gong, J. Wang, et al. NTIRE 2025 challenge on image super-resolution (×4): methods and results. arXiv preprint arXiv:2504.14582. Cited by: §3.3.3.
- Z. Chen, W. Sun, J. Jia, F. Lu, Z. Zhang, J. Liu, R. Huang, X. Min, and G. Zhai BAND-2k: banding artifact noticeable database for banding detection and quality assessment. IEEE Transactions on Circuits and Systems for Video Technology 34 (7), pp. 6347–6362. External Links: Document Cited by: §3.7.1.
- B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.4.1.
- H. K. Cheng and A. G. Schwing Xmem: long-term video object segmentation with an atkinson-shiffrin memory model. In European Conference on Computer Vision (ECCV), Vol. 13688, pp. 640–658. Cited by: §3.4.3.
- Z. Cheng, H. Sun, M. Takeuchi, and J. Katto Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 7939–7948. Cited by: §3.6.1.
- M. Cheon, S. Yoon, B. Kang, and J. Lee Perceptual image quality assessment with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 433–442. Cited by: Table 1, §3.7.1.
- J. Choi, S. Kim, Y. Jeong, Y. Gwon, and S. Yoon ILVR: Conditioning method for denoising diffusion probabilistic models. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 14347–14356. External Links: Document Cited by: §2.3.
- M. Choi, H. Lee, and H. Lee Exploring positional characteristics of dual-pixel data for camera autofocus. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 13112–13122. External Links: Document Cited by: §3.3.4.
- A. Y.K. Chua, A. Pal, and S. Banerjee AI-enabled investment advice: will users buy it?. Computers in Human Behavior 138, pp. 107481. External Links: Document Cited by: §3.2.3.
- J. Chung, S. Hyun, and J. Heo Style injection in diffusion: a training-free approach for adapting large-scale diffusion models for style transfer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8795–8805. Cited by: Table 1, §3.3.2.
- N. C. Chung Human in the loop for machine creativity. In AAAI Conference on Human Computation and Crowdsourcing, Cited by: §1.
- Communications and Digital Committee Large language models and generative AI. 1st Report of Session 2023–24. Cited by: §1, §1, §4.2.
- M. V. Conde, E. Zamfir, R. Timofte, D. Motilla, et al. Efficient deep models for real-time 4k image super-resolution. ntire 2023 benchmark and report. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vol. , pp. 1495–1521. External Links: Document Cited by: §3.3.3.
- E. Corona, A. Zanfir, E. G. Bazavan, N. Kolotouros, T. Alldieck, and C. Sminchisescu VLOGGER: multimodal diffusion for embodied avatar synthesis. arXiv:2403.08764. Cited by: Table 1, §3.1.4.
- C. I. Council How createch added 1+1 to make £981m. Note: https://www.thecreativeindustries.co.uk/site-content/how-createch-added-one-plus-one-to-make-ps981m (Accessed: 2025-01-10) Cited by: §4.3.
- Y. Cui, C. Jiang, L. Wang, and G. Wu Mixformer: end-to-end tracking with iterative mixed attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13608–13618. Cited by: Table 1, §3.4.3.
- DeepSeek-AI, A. Liu, B. Feng, B. Xue, et al. DeepSeek-v3 technical report. arXiv preprint arXiv:2412.19437. Note: https://arxiv.org/abs/2412.19437 Cited by: §4.3.
- M. Deitke, R. Liu, M. Wallingford, H. Ngo, O. Michel, A. Kusupati, A. Fan, C. Laforte, V. Voleti, S. Y. Gadre, E. VanderBilt, A. Kembhavi, C. Vondrick, G. Gkioxari, K. Ehsani, L. Schmidt, and A. Farhadi Objaverse-XL: A universe of 10m+ 3D objects. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §3.1.5.
- Y. Deng, F. Tang, W. Dong, C. Ma, X. Pan, L. Wang, and C. Xu StyTr2: Image style transfer with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11326–11336. Cited by: Table 1, §3.3.2.
- P. Dhariwal and A. Q. Nichol Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Eds.), Cited by: §2.3, §2.3.
- A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby An image is worth 16$\times$16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, Cited by: §2.1, §2.1, Table 1.
- E. Dupont, A. Golinski, M. Alizadeh, Y. W. Teh, and A. Doucet COIN: compression with implicit neural representations. In ICLR Workshop in Neural Compression, Cited by: Table 1, §3.6.1.
- E. Dupont, H. Loya, M. Alizadeh, A. Golinski, Y. W. Teh, and A. Doucet COIN++: neural compression across modalities. Transactions on Machine Learning Research. Cited by: Table 1, §3.6.1.
- O. Elharrouss, R. Damseh, A. N. Belkacem, S. Al-Maadeed, A. Saleous, and S. Bourennane Transformer-based image and video inpainting: current challenges and future directions. Artificial Intelligence Review 58, pp. 124. External Links: Document Cited by: §3.3.5.
- P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Proceedings of the 41 st International Conference on Machine Learning, Cited by: Table 1, Table 1, §3.1.3.
- Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Taylor, and J. Pons Stable audio open. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , pp. 1–5. External Links: Document Cited by: Table 1, §3.1.2.
- C. Fan, T. Liu, and K. Liu SUNet: swin transformer unet for image denoising. In 2022 IEEE International Symposium on Circuits and Systems (ISCAS), pp. 2333–2337. Cited by: §2.1, Table 1, §3.3.4.
- H. Fang, B. Han, S. Zhang, S. Zhou, C. Hu, and W. Ye Data augmentation for object detection via controllable diffusion models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 1257–1266. Cited by: §3.4.2.
- J. Fang, T. Yi, X. Wang, L. Xie, X. Zhang, W. Liu, M. Nießner, and Q. Tian Fast dynamic radiance fields with time-aware neural voxels. In SIGGRAPH Asia 2022 Conference Papers, Note: https://doi.org/10.1145/3550469.3555383 External Links: Document Cited by: Table 1, §3.5.2.
- W. Fang, J. Fan, Y. Zheng, J. Weng, Y. Tai, and J. Li Guided real image dehazing using ycbcr color space. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 2906–2914. External Links: Document Cited by: Table 1, §3.3.4.
- B. Fei, Z. Lyu, L. Pan, J. Zhang, W. Yang, T. Luo, B. Zhang, and B. Dai Generative diffusion prior for unified image restoration and enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9935–9946. Cited by: Table 1, Table 1, Table 1, §3.3.4.
- B. Fei, J. Xu, R. Zhang, Q. Zhou, W. Yang, and Y. He 3D gaussian splatting as new era: a survey. IEEE Transactions on Visualization and Computer Graphics (), pp. 1–20. External Links: Document Cited by: §3.5.3.
- S. Feizi, M. Hajiaghayi, K. Rezaei, and S. Shin Online advertisements with llms: opportunities and challenges. arXiv preprint arXiv:2311.07601. Cited by: §3.2.2.
- C. Feng, D. Danier, F. Zhang, and D. Bull Rankdvqa: deep vqa based on ranking-inspired hybrid training. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1648–1658. Cited by: Table 1, §3.7.1, §3.7.1, §3.7.2.
- C. Feng, D. Danier, F. Zhang, A. Mackin, A. Collins, and D. Bull BVI-artefact: an artefact detection benchmark dataset for streamed videos. In 2024 Picture Coding Symposium (PCS), pp. 1–5. Cited by: §3.7.1.
- H. Feng, H. Zhou, T. Ye, S. Chen, and L. Zhu Residual diffusion deblurring model for single image defocus deblurring. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 2960–2968. External Links: Document Cited by: Table 1, §3.3.4.
- K. Feng, Y. Ma, B. Wang, C. Qi, H. Chen, Q. Chen, and Z. Wang DiT4Edit: Diffusion transformer for image editing. Proceedings of the AAAI Conference on Artificial Intelligence 39 (3), pp. 2969–2977. Note: https://ojs.aaai.org/index.php/AAAI/article/view/32304 External Links: Document Cited by: Table 1, §3.1.3.
- Z. Feng, C. Jung, H. Zhang, Y. Liu, and M. Li Low complexity in-loop filter for VVC based on convolution and transformer. IEEE Access. Cited by: §3.6.2.
- D. Fuoli, M. Danelljan, R. Timofte, and L. Van Gool Fast online video super-resolution with deformable attention pyramid. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 1735–1744. Cited by: §3.3.3.
- R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-or An image is worth one word: personalizing text-to-image generation using textual inversion. In The Eleventh International Conference on Learning Representations, Cited by: Table 1, §3.1.3.
- F. Galpin, Y. Li, Y. Li, J. N. Shingala, L. Wang, and Z. Xie NNVC software development AhG14. Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AG0014. Cited by: §3.6.2.
- R. Gandikota, H. Orgad, Y. Belinkov, J. Materzyńska, and D. Bau Unified concept editing in diffusion models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 5111–5120. Cited by: Table 1, §4.2.
- G. Gao, H. M. Kwan, F. Zhang, and D. Bull PNVC: towards practical INR-based video compression. arXiv preprint arXiv:2409.00953. Cited by: Table 1, §3.6.2.
- S. Gao, X. Liu, B. Zeng, S. Xu, Y. Li, X. Luo, J. Liu, X. Zhen, and B. Zhang Implicit diffusion models for continuous super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10021–10030. Cited by: Table 1, Table 1, §3.3.3.
- Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun YOLOX: exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430. Cited by: §3.4.3.
- N. F. Ghouse, J. Petersen, A. Wiggers, T. Xu, and G. Sautiere A residual diffusion model for high perceptual quality codec augmentation. arXiv preprint arXiv:2301.05489. Cited by: Table 1, §3.6.1.
- R. Goel, D. Sirikonda, S. Saini, and P. J. Narayanan Interactive segmentation of radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4201–4211. Cited by: §3.4.1.
- S. A. Golestaneh, S. Dadsetan, and K. M. Kitani No-reference image quality assessment via transformers, relative ranking, and self-consistency. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 1220–1230. Cited by: Table 1, §3.7.1.
- R. Gong, Q. Wang, M. Danelljan, D. Dai, and L. Van Gool Continuous pseudo-label rectified domain adaptive semantic segmentation with implicit neural representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7225–7235. Cited by: Table 1, §3.4.1.
- I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio Generative adversarial nets. In Advances in Neural Information Processing Systems 27, Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger (Eds.), pp. 2672–2680. Note: http://papers.nips.cc/paper/5423-generative-adversarial-nets.pdf Cited by: §2.3.
- A. Gu and T. Dao Mamba: Linear-time sequence modeling with selective state spaces. In Conference on Language Modeling, Cited by: §2.1.
- Z. Gu, H. Chen, and Z. Xu Diffusioninst: Diffusion model for instance segmentation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , pp. 2730–2734. External Links: Document Cited by: Table 1.
- A. Guo, P. Pataranutaporn, and P. Maes Exploring the interaction of creative writers with AI-powered writing tools. arXiv:2402.12814. Cited by: §3.1.1.
- H. Guo, S. Peng, H. Lin, Q. Wang, G. Zhang, H. Bao, and X. Zhou Neural 3D scene reconstruction with the manhattan-world assumption. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5511–5520. Cited by: Table 1, §3.5.2.
- J. Guo, D. Zhang, X. Liu, Z. Zhong, Y. Zhang, P. Wan, and D. Zhang LivePortrait: Efficient portrait animation with stitching and retargeting control. arXiv preprint arXiv:2407.03168. Cited by: Table 1, §3.3.7.
- M. Guo, T. Xu, J. Liu, and et al. Attention mechanisms in computer vision: a survey. Computational Visual Media 8, pp. 331–368. Cited by: §2.1.
- W. Guo-Hua, J. Li, B. Li, and Y. Lu EVC: towards real-time neural image compression with mask decay. In International Conference on Learning Representations, Cited by: §3.6.2.
- A. Gupta, L. Yu, K. Sohn, X. Gu, M. Hahn, F. Li, I. Essa, L. Jiang, and J. Lezama Photorealistic video generation with diffusion models. In European Conference Computer Vision, pp. 393–411. External Links: Document Cited by: Table 1, Table 1, §3.1.4.
- J. Hartmann, M. Heitmann, C. Siebert, and C. Schamp More than a feeling: accuracy and application of sentiment analysis. International Journal of Research in Marketing 40 (1), pp. 75–87. External Links: Document Cited by: Table 1, §3.2.2.
- C. He, Q. Zheng, R. Zhu, X. Zeng, Y. Fan, and Z. Tu COVER: a comprehensive video quality evaluator. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5799–5809. Cited by: Table 1, §3.7.1, §3.7.1.
- C. He, Y. Shen, C. Fang, F. Xiao, L. Tang, Y. Zhang, W. Zuo, Z. Guo, and X. Li Diffusion models in low-level vision: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 (6), pp. 4630–4651. Cited by: §3.3.4.
- Y. He, H. Yu, X. Liu, Z. Yang, W. Sun, S. Anwar, and A. Mian Deep learning based 3d segmentation in computer vision: a survey. Information Fusion 115, pp. 102722. External Links: ISSN 1566-2535, Document Cited by: §3.4.1.
- P. Hill, N. Anantrasirichai, A. Achim, and D. Bull Deep learning techniques for atmospheric turbulence removal: a review. Artificial Intelligence Review 58 (101). External Links: Document Cited by: §3.3.4.
- J. Ho, A. Jain, and P. Abbeel Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, pp. 6840–6851.. Cited by: §2.3.
- Y. Ho, C. Chang, P. Chen, A. Gnutti, and W. Peng CANF-VC: conditional augmented normalizing flows for video compression. In European Conference on Computer Vision, pp. 207–223. Cited by: §3.6.2.
- W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang CogVideo: large-scale pretraining for text-to-video generation via transformers. In The Eleventh International Conference on Learning Representations, Cited by: Table 1, §3.1.4.
- E. Hoogeboom, E. Agustsson, F. Mentzer, L. Versari, G. Toderici, and L. Theis High-fidelity image compression with score-based generative models. arXiv preprint arXiv:2305.18231. Cited by: Table 1, §3.6.1.
- V. Hosu, F. Hahn, M. Jenadeleh, H. Lin, H. Men, T. Szirányi, S. Li, and D. Saupe The konstanz natural video database (konvid-1k). In 2017 Ninth International Conf. on Quality of Multimedia Experience (QoMEX), pp. 1–6. Cited by: §3.7.1.
- B. Hou, J. O’Connor, J. Andreas, S. Chang, and Y. Zhang PromptBoosting: black-box text classification with ten forward passes. In Proceedings of the 40th International Conference on Machine Learning, Vol. 202, pp. 13309–13324. Cited by: Table 1, §3.2.1.
- J. HOU, Z. Zhu, J. Hou, H. LIU, H. Zeng, and H. Yuan Global structure-aware diffusion process for low-light image enhancement. In Advances in Neural Information Processing Systems, Vol. 36, pp. 79734–79747. Cited by: Table 1, §3.3.1.
- Q. Hou, A. Ghildyal, and F. Liu A perceptual quality metric for video frame interpolation. In European Conf. on Computer Vision, pp. 234–253. Cited by: §3.7.1.
- L. Hu Animate Anyone: Consistent and controllable image-to-video synthesis for character animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8153–8163. Cited by: Table 1, §3.1.4.
- T. Hu, S. Liu, Y. Chen, T. Shen, and J. Jia EfficientNeRF - Efficient neural radiance fields. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 12892–12901. External Links: Document Cited by: Table 1, §3.5.3.
- Z. Hu, G. Lu, and D. Xu FVC: a new framework towards deep video compression in feature space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1502–1511. Cited by: §3.6.2.
- G. Huang, N. Anantrasirichai, F. Ye, Z. Qi, R. Lin, Q. Yang, and D. Bull Bayesian neural networks for one-to-many mapping in image enhancement. arXiv preprint arXiv:2501.14265. Cited by: §3.3.1.
- K. Huang, T. Wu, H. Su, and W. H. Hsu MonoDTR: Monocular 3D object detection with depth-aware transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.4.2.
- R. Huang, Y. Ren, J. Liu, C. Cui, and Z. Zhao GenerSpeech: Towards style transfer for generalizable out-of-domain text-to-speech. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.1.2.
- W. Huang, Y. Deng, S. Hui, Y. Wu, S. Zhou, and J. Wang Sparse self-attention transformer for image inpainting. Pattern Recognition 145, pp. 109897. External Links: ISSN 0031-3203 Cited by: Table 1, §3.3.5.
- Y. Huang, J. Huang, Y. Liu, M. Yan, J. Lv, J. Liu, W. Xiong, H. Zhang, L. Cao, and S. Chen Diffusion model-based image editing: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 (6), pp. 4409–4437. External Links: Document Cited by: §3.1.3.
- Y. Huang, Y. Sun, Z. Yang, X. Lyu, Y. Cao, and X. Qi SC-GS: Sparse-controlled gaussian splatting for editable dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4220–4230. Cited by: Table 1, §3.5.3.
- E. Im, C. Jee, and J. K. Lee GATE3D: Generalized Attention-based Task-synergized Estimation in 3D. arXiv preprint arXiv:2504.11014. Cited by: Table 1, §3.4.2.
- D. Ippolito, A. Yuan, A. Coenen, and S. Burnam Creative writing with an AI-powered writing assistant: perspectives from professional writers. arXiv:2211.05030. Cited by: §3.1.1.
- ISCAS 2024 grand challenge on neural network-based video coding. Note: https://iscasnnvcgc.github.io/Accessed: 2025-05-21 Cited by: §3.6.2.
- L. Itti and C. Koch Computational modelling of visual attention. Nature reviews neuroscience 2 (3), pp. 194–203. Cited by: §3.7.1.
- A. Jaiswal, X. Zhang, S. H. Chan, and Z. Wang Physics-driven turbulence image restoration with stochastic refinement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12170–12181. Cited by: Table 1, §3.3.4.
- L. Jeary and D. Gajjar Artificial intelligence and new technology in creative industries. UK Parliament Horizon Scanning Post. Note: https://post.parliament.uk/artificial-intelligence-and-new-technology-in-creative-industries/ (Accessed: 2025-01-10) Cited by: §1, §3.1.1, §4.2, §4.3.
- Y. Ji, Z. Chen, E. Xie, L. Hong, X. Liu, Z. Liu, T. Lu, Z. Li, and P. Luo DDP: Diffusion model for dense visual prediction. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 21684–21695. External Links: Document Cited by: Table 1, §3.5.1.
- M. Jia, L. Tang, B. Chen, C. Cardie, S. Belongie, B. Hariharan, and S. Lim Visual prompt tuning. In European Conference on Computer Vision (ECCV), Cited by: §2.2.
- H. Jiang, A. Luo, H. Fan, S. Han, and S. Liu Low-light image enhancement with wavelet-based diffusion models. ACM Transactions on Graphics (TOG) 42 (6), pp. 1–14. Cited by: Table 1, Table 1, §3.3.1.
- N. Jiang, S. Hasanzadeh, and V. G. Duffy Domain-tailored generative ai for personalized assistant. In HCI International 2024 – Late Breaking Papers, V. G. Duffy (Ed.), Lecture Notes in Computer Science, Vol. 15376, pp. 227–237. External Links: Document Cited by: §3.2.4.
- W. Jiang, V. Boominathan, and A. Veeraraghavan NeRT: Implicit neural representations for unsupervised atmospheric turbulence mitigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 4236–4243. Cited by: Table 1, §3.3.4.
- Y. Jiang, C. Yu, T. Xie, X. Li, Y. Feng, H. Wang, M. Li, H. Lau, F. Gao, Y. Yang, and C. Jiang VR-GS: A physical dynamics-aware interactive gaussian splatting system in virtual reality. In ACM SIGGRAPH 2024 Conference Papers, External Links: Document Cited by: Figure 7, §3.5.3.
- D. Jin, J. Lei, B. Peng, W. Li, N. Ling, and Q. Huang Deep affine motion compensation network for inter prediction in VVC. IEEE Transactions on Circuits and Systems for Video Technology 32 (6), pp. 3923–3933. Cited by: §3.6.2.
- P. Jin, H. Li, Z. Cheng, K. Li, X. Ji, C. Liu, L. Yuan, and J. Chen DiffusionRet: Generative text-video retrieval with diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 2470–2481. Cited by: Table 1, §3.2.3.
- Y. Jin, X. Ma, R. Zhang, H. Chen, Y. Gu, P. Ling, and E. Chen Masked video pretraining advances real-world video denoising. IEEE Transactions on Multimedia 27 (), pp. 622–636. External Links: Document Cited by: Table 1, §3.3.4.
- U. Joshi, Y. Chen, I. Yoo, S. Li, F. Yang, and D. Mukherjee Switchable cnns for in-loop restoration and super-resolution for av2. In Applications of Digital Image Processing XLVI, Vol. 12674, pp. 121–130. Cited by: §3.6.2.
- H. Jung, Z. Hui, L. Luo, H. Yang, F. Liu, S. Yoo, R. Ranjan, and D. Demandolx AnyFlow: Arbitrary scale optical flow with implicit neural representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5455–5465. Cited by: Table 1, §3.4.3.
- H. Junkawitsch, G. Sun, H. Zhu, C. Theobalt, and M. Habermann EVA: Expressive Virtual Avatars from Multi-view Videos. In ACM SIGGRAPH 2025 Conference Proceedings, Cited by: Table 1, §3.5.3.
- B. Kang, X. Chen, S. Lai, Y. Liu, Y. Liu, and D. Wang Exploring enhanced contextual information for video-level object tracking. In Proceedings of the AAAI Conference on Artificial Intelligence, Cited by: Table 1, §3.4.3.
- M. Kang, J. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park Scaling up gans for text-to-image synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.3.3.
- A. Karabutov, Y. Wu, E. Alshina, and J. Ascenso JPEG AI sw v4.x status. JPEG AI ISO/IEC JTC 1/SC29/WG1 M101081. Cited by: §3.6.1.
- S. Karim, G. Tong, J. Li, A. Qadir, U. Farooq, and Y. Yu Current advances and future perspectives of image fusion: a comprehensive review. Information Fusion 90, pp. 185–217. External Links: ISSN 1566-2535, Document Cited by: §3.3.6.
- B. Kathariya, Z. Li, and G. Van der Auwera Joint pixel and frequency feature learning and fusion via channel-wise transformer for high-efficiency learned in-loop filter in vvc. IEEE Transactions on Circuits and Systems for Video Technology. Cited by: §3.6.2.
- B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler Repurposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9492–9502. Cited by: Table 1, §3.5.1.
- L. Ke, M. Ye, M. Danelljan, Y. liu, Y. Tai, C. Tang, and F. Yu Segment anything in high quality. In Advances in Neural Information Processing Systems, Vol. 36, pp. 29914–29934. Cited by: Table 1, Figure 6, §3.4.1.
- D. Kelly Visual contrast sensitivity. Optica Acta: International Journal of Optics 24 (2), pp. 107–129. Cited by: §3.7.1.
- B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis 3D gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics 42 (4). Note: https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/ Cited by: Table 1, Figure 7, §3.5.3.
- S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah Transformers in vision: a survey. ACM computing surveys (CSUR) 54 (10s). Note: https://doi.org/10.1145/3505244 External Links: ISSN 0360-0300, Document Cited by: §2.1.
- M. Khani, V. Sivaraman, and M. Alizadeh Efficient video compression via content-adaptive super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4521–4530. Cited by: §3.6.2.
- H. Kheddar Transformers and large language models for efficient intrusion detection systems: a comprehensive survey. Information Fusion 124, pp. 103347. Note: https://www.sciencedirect.com/science/article/pii/S1566253525004208 External Links: ISSN 1566-2535, Document Cited by: §3.4.2.
- H. Kim, M. Bauer, L. Theis, J. R. Schwarz, and E. Dupont C3: high-performance and low-complexity neural compression from a single image or video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9347–9358. Cited by: Table 1, §3.6.2.
- J. Kim and S. Lee Deep learning of human visual sensitivity in image quality assessment framework. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1676–1684. Cited by: §3.7.1.
- S. Kim, Y. Min, Y. Jung, and S. Kim Controllable style transfer via test-time training of implicit neural representation. Pattern Recognition 146, pp. 109988. External Links: Document Cited by: Table 1, §3.3.2.
- W. Kim, J. Kim, S. Ahn, J. Kim, and S. Lee Deep video quality assessor: from spatio-temporal visual sensitivity to a convolutional neural aggregation network. In Proceedings of the European conference on computer vision (ECCV), pp. 219–234. Cited by: §3.7.1.
- E. King, H. Yu, S. Lee, and C. Julien Sasha: Creative goal-oriented reasoning in smart homes with large language models. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 8 (1). External Links: Document Cited by: Table 1, §3.2.4.
- D.P. Kingma and M. Welling Auto-encoding variational bayes. In International Conference on Learning Representations, Cited by: §2.3.
- A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, P. Dollar, and R. Girshick Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4015–4026. Cited by: Table 1, §3.4.1, §3.4.3.
- H. Kong, X. Yang, and X. Wang Efficient gaussian splatting for monocular dynamic scene rendering via sparse time-variant attribute modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 4374–4382. Cited by: Table 1, §3.5.3.
- J. Korhonen Two-level approach for no-reference consumer video quality assessment. IEEE Transactions on Image Processing 28 (12), pp. 5923–5938. Cited by: §3.7.1.
- J. O. Krugmann and J. Hartmann Sentiment analysis in the age of generative AI. Customer Needs and Solutions 11 (3). Cited by: Table 1, §3.2.2.
- J. Kugarajeevan, T. Kokul, A. Ramanan, and S. Fernando Transformers in single object tracking: an experimental survey. IEEE Access 11 (), pp. 80297–80326. External Links: Document Cited by: §3.4.3.
- R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar High-fidelity audio compression with improved rvqgan. Advances in Neural Information Processing Systems 36. Cited by: §3.6.3.
- H. M. Kwan, G. Gao, F. Zhang, A. Gower, and D. Bull HiNeRV: video compression with hierarchical encoding-based neural representation. Advances in Neural Information Processing Systems 36, pp. 72692–72704. Cited by: §2.4, Table 1, §3.6.2.
- H. M. Kwan, G. Gao, F. Zhang, A. Gower, and D. Bull NVRC: neural video representation compression. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.6.2.
- H. M. Kwan, F. Zhang, A. Gower, and D. Bull Immersive video compression using implicit neural representations. In Picture Coding Symposium, Cited by: Table 1, §3.6.2.
- E. C. Larson and D. M. Chandler Most apparent distortion: full-reference image quality assessment and the role of strategy. Journal of electronic imaging 19 (1), pp. 011006–011006. Cited by: §3.7.1, §3.7.1.
- M. Lee, K. I. Gero, J. J. Y. Chung, S. B. Shum, V. Raheja, H. Shen, S. Venugopalan, T. Wambsganss, D. Zhou, E. A. Alghamdi, T. August, A. Bhat, M. Z. Choksi, S. Dutta, J. L.C. Guo, M. N. Hoque, Y. Kim, S. Knight, S. P. Neshaei, A. Shibani, D. Shrivastava, L. Shroff, A. Sergeyuk, J. Stark, S. Sterman, S. Wang, A. Bosselut, D. Buschek, J. C. Chang, S. Chen, M. Kreminski, J. Park, R. Pea, E. H. R. Rho, Z. Shen, and P. Siangliulue A design space for intelligent and interactive writing assistants. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, External Links: Document Cited by: §3.2.4.
- T. Leguay, T. Ladune, P. Philippe, and O. Déforges Cool-chic video: Learned video coding with 800 parameters. In 2024 Data Compression Conference (DCC), Vol. , pp. 23–32. External Links: Document Cited by: Table 1, §3.6.2.
- B. Lester, R. Al-Rfou, and N. Constant The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3045–3059. External Links: Document Cited by: §2.2.
- A. C. Li, M. Prabhudesai, S. Duggal, E. Brown, and D. Pathak Your diffusion model is secretly a zero-shot classifier. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 2206–2217. External Links: Document Cited by: Table 1, §3.4.2.
- B. Li, Y. Liu, X. Niu, B. Bai, L. Deng, and D. Gündüz Extreme video compression with pre-trained diffusion models. arXiv preprint arXiv:2402.08934. Cited by: Table 1, §3.6.2.
- D. Li, Y. Bai, K. Wang, J. Jiang, and X. Liu Semantic ensemble loss and latent refinement for high-fidelity neural image compression. In 2024 IEEE International Conference on Visual Communications and Image Processing (VCIP), pp. 1–5. Cited by: §3.6.1.
- G. Li, Q. Fang, L. Zha, X. Gao, and N. Zheng HAM: Hybrid attention module in deep convolutional neural networks for image classification. Pattern Recognition 129, pp. 108785. Cited by: §2.1.
- H. Li and X. Wu CrossFuse: A novel cross attention mechanism based infrared and visible image fusion approach. Information Fusion 103, pp. 102147. External Links: ISSN 1566-2535, Document Cited by: Table 1, §3.3.6.
- J. Li, B. Li, and Y. Lu Deep contextual video compression. Advances in Neural Information Processing Systems 34, pp. 18114–18125. Cited by: §3.6.2.
- J. Li, B. Li, and Y. Lu Neural video compression with diverse contexts. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22616–22626. Cited by: §3.6.2.
- J. Li, B. Li, and Y. Lu Neural video compression with feature modulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26099–26108. Cited by: §3.6.2.
- T. Li, M. Xu, R. Tang, Y. Chen, and Q. Xing DeepQTMT: a deep learning approach for fast QTMT-based CU partition of intra-mode VVC. IEEE Transactions on Image Processing 30, pp. 5377–5390. Cited by: §3.6.2.
- W. Li, Z. Lin, K. Zhou, L. Qi, Y. Wang, and J. Jia MAT: Mask-aware transformer for large hole image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10758–10768. Cited by: Table 1, §3.3.5.
- X. Li, Y. Zhou, and Z. Dou UniGen: A unified generative framework for retrieval and question answering with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 8688–8696. Cited by: Table 1, §3.2.3.
- X. Li, J. Thickstun, I. Gulrajani, P. S. Liang, and T. B. Hashimoto Diffusion-lm improves controllable text generation. In Advances in Neural Information Processing Systems, Vol. 35, pp. 4328–4343. Cited by: Table 1.
- X. Li, J. Jin, Y. Zhou, Y. Zhang, P. Zhang, Y. Zhu, and Z. Dou From matching to generation: a survey on generative information retrieval. ACM Trans. Inf. Syst. 43 (3). Note: https://doi.org/10.1145/3722552 External Links: ISSN 1046-8188, Document Cited by: §3.2.3.
- Y. Li, Y. Fan, X. Xiang, R. R. Denis Demandolx, R. Timofte, and L. V. Gool Efficient and explicit modelling of image hierarchies for image restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Cited by: Table 1, Table 1, §3.3.4.
- Y. Li, N. Miao, L. Ma, F. Shuang, and X. Huang Transformer for object detection: review and benchmark. Engineering Applications of Artificial Intelligence 126, pp. 107021. Note: https://www.sciencedirect.com/science/article/pii/S0952197623012058 External Links: Document Cited by: §3.4.2.
- Y. Li, N. Yang, L. Wang, F. Wei, and W. Li Learning to rank in generative retrieval. In AAAI 2024, Cited by: Table 1, §3.2.3.
- Y. Li, J. Li, C. Lin, K. Zhang, L. Zhang, F. Galpin, T. Dumas, H. Wang, M. Coban, J. Ström, et al. Designs and implementations in neural network-based video coding. arXiv preprint arXiv:2309.05846. Cited by: §3.6.2.
- Z. Li, A. Aaron, I. Katsavounidis, A. Moorthy, and M. Manohara The NETFLIX tech blog: Toward a practical perceptual video quality metric. Note: http://techblog.netflix.com/2016/06/toward-practical-perceptual-video.html, note = [Online; accessed 2018-08-04], Cited by: §3.6.1, §3.7.1.
- Z. Li, M. Wang, H. Pi, K. Xu, J. Mei, and Y. Liu E-NeRV: expedite neural video representation with disentangled spatial-temporal context. In European Conference on Computer Vision, pp. 267–284. Cited by: §3.6.2.
- L. Lian, B. Li, A. Yala, and T. Darrell LLM-grounded diffusion: enhancing prompt understanding of text-to-image diffusion models with large language models. Transactions on Machine Learning Research. Note: Featured Certification External Links: ISSN 2835-8856 Cited by: Table 1, Figure 3, §3.1.3.
- J. Liang, Y. Fan, X. Xiang, R. Ranjan, E. Ilg, S. Green, J. Cao, K. Zhang, R. Timofte, and L. Gool Recurrent video restoration transformer with guided deformable attention. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.3.4, §3.3.4.
- J. Liang, J. Cao, Y. Fan, K. Zhang, R. Ranjan, Y. Li, R. Timofte, and L. Van Gool VRT: a video restoration transformer. IEEE Transactions on Image Processing 33 (), pp. 2171–2182. External Links: Document Cited by: Table 1, Table 1, Table 1, §3.3.4.
- J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte SwinIR: image restoration using swin transformer. In 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Vol. , pp. 1833–1844. External Links: Document Cited by: Table 1, Table 1, §3.3.1, §3.3.4, §3.3.4.
- J. Lin, N. Anantrasirichai, and D. Bull Feature denoising for low-light instance segmentation using weighted non-local blocks. In IEEE International Conference on Acoustics, Speech, and Signal Processing, Note: https://arxiv.org/pdf/2402.18307.pdf Cited by: §3.4.1.
- R. Lin, N. Anantrasirichai, G. Huang, J. Lin, Q. Sun, A. Malyugina, and D. Bull BVI-RLV: A fully registered dataset and benchmarks for low-light video enhancement. arXiv preprint arXiv:2407.03535. Cited by: §3.3.1.
- R. Lin, Q. Sun, and N. Anantrasirichai Low-light video enhancement with conditional diffusion models and wavelet interscale attentions. In Proceedings of the 21st ACM SIGGRAPH Conference on Visual Media Production, pp. 1–10. Cited by: Table 1, §3.3.1.
- R. Lin, N. Anantrasirichai, A. Malyugina, and D. Bull A spatio-temporal aligned sunet model for low-light video enhancement. In IEEE International Conference on Image Processing, Cited by: Table 1, §3.3.1.
- C. Liu, H. Yang, J. Fu, and X. Qian Learning trajectory-aware transformer for video super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5687–5696. Cited by: Table 1, §3.3.3.
- D. Liu, Z. Wang, and P. Chen DSEM-NeRF: Multimodal feature fusion and global–local attention for enhanced 3d scene reconstruction. Information Fusion 115, pp. 102752. Note: https://www.sciencedirect.com/science/article/pii/S156625352400530X External Links: ISSN 1566-2535, Document Cited by: Table 1, Table 1, §3.5.2.
- J. Liu, Y. Cao, J. Z. Wu, W. Mao, Y. Gu, R. Zhao, J. Keppo, Y. Shan, and M. Z. Shou DynVideo-e: harnessing dynamic nerf for large-scale motion- and view-change human-centric video editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7664–7674. Cited by: Table 1, §3.5.2.
- J. Liu, H. Sun, and J. Katto Learned image compression with mixed transformer-cnn architectures. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14388–14397. Cited by: Table 1, §3.6.1.
- J. Liu, Z. Liu, G. Wu, L. Ma, R. Liu, W. Zhong, Z. Luo, and X. Fan Multi-interactive feature learning and a full-time multi-modality benchmark for image fusion and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8115–8124. Cited by: Table 1, §3.3.6.
- M. Liu, Y. Ma, Z. Yang, J. Dan, Y. Yu, Z. Zhao, Z. Hu, B. Liu, and C. Fan LLM4GEN: Leveraging semantic representation of llms for text-to-image generation. Proceedings of the AAAI Conference on Artificial Intelligence 39 (5), pp. 5523–5531. Note: https://ojs.aaai.org/index.php/AAAI/article/view/32588 External Links: Document Cited by: Table 1, §3.1.3.
- Q. Liu, Z. Tan, D. Chen, Q. Chu, X. Dai, Y. Chen, M. Liu, L. Yuan, and N. Yu Reduce information loss in transformers for pluralistic image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11347–11357. Cited by: Table 1, §3.3.5.
- X. Liu, J. Van De Weijer, and A. D. Bagdanov RankIQA: learning from rankings for no-reference image quality assessment. In 2017 IEEE International Conference on Computer Vision (ICCV), Vol. , pp. 1040–1049. External Links: Document Cited by: §3.7.1.
- Y. Liu, L. Li, S. Ren, R. Gao, S. Li, S. Chen, X. Sun, and L. Hou FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation. In Advances in Neural Information Processing Systems, Vol. 36, pp. 62352–62387. Cited by: Table 1, Table 1, §3.1.4.
- Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo Swin transformer v2: scaling up capacity and resolution. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 11999–12009. External Links: Document Cited by: §2.1, Table 1.
- Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Cited by: §2.1, Table 1, §3.3.4, §3.4.2.
- D. Lu, S. Wang, N. Kumari, R. Agarwal, M. Tang, D. Bau, and J. Zhu Content-based search for deep generative models. In SIGGRAPH Asia 2023 Conference Papers, External Links: Document Cited by: Table 1, §3.2.3.
- G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao DVC: an end-to-end deep video compression framework. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11006–11015. Cited by: §3.6.2.
- Z. Lu, J. Li, H. Liu, C. Huang, L. Zhang, and T. Zeng Transformer for single image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 457–466. Cited by: Table 1, §3.3.3.
- R. Luo, Z. Song, L. Ma, J. Wei, W. Yang, and M. Yang DiffusionTrack: Diffusion model for multi-object tracking. Proceedings of the AAAI Conference on Artificial Intelligence 38 (5), pp. 3991–3999. Note: https://ojs.aaai.org/index.php/AAAI/article/view/28192 External Links: Document Cited by: Table 1, §3.4.3.
- J. Ma, L. Tang, F. Fan, J. Huang, X. Mei, and Y. Ma SwinFusion: Cross-domain long-range learning for general image fusion via swin transformer. IEEE/CAA Journal of Automatica Sinica 9 (7), pp. 1200–1217. External Links: Document Cited by: Table 1, §3.3.6.
- P. C. Madhusudana, N. Birkbeck, Y. Wang, B. Adsumilli, and A. C. Bovik Image quality assessment using contrastive learning. IEEE Transactions on Image Processing 31, pp. 4149–4161. Cited by: §3.7.1, §3.7.1.
- P. C. Madhusudana, N. Birkbeck, Y. Wang, B. Adsumilli, and A. C. Bovik Conviqt: contrastive video quality estimator. IEEE Transactions on Image Processing. Cited by: §3.7.1.
- P. C. Madhusudana, X. Yu, N. Birkbeck, Y. Wang, B. Adsumilli, and A. C. Bovik Subjective and objective quality assessment of high frame rate videos. IEEE Access 9, pp. 108069–108082. Cited by: §3.7.1.
- R. Mao, Q. Liu, K. He, W. Li, and E. Cambria The biases of pre-trained language models: an empirical study on prompt-based sentiment analysis and emotion detection. IEEE Transactions on Affective Computing 14 (3), pp. 1743–1753. External Links: Document Cited by: Table 1, §3.2.2.
- Z. Mao, A. Jaiswal, Z. Wang, and S. H. Chan Single frame atmospheric turbulence mitigation: a benchmark study and a new physics-inspired transformer model. In European Conference on Computer Vision (ECCV), Cited by: Table 1.
- C. Mayer, M. Danelljan, G. Bhat, M. Paul, D. P. Paudel,
Advances in Artificial Intelligence: A Review for the Creative Industries
Artificial intelligence (AI) has undergone transformative advances since 2022, particularly through generative AI, large language models (LLMs), and diffusion models, fundamentally reshaping the creative industries. However, existing reviews have not comprehensively addressed these recent breakthroughs and their integrated impact across the creative production pipeline. This paper addresses this gap by providing a s…
- Licence
- OPEN CC-BY-4.0
- Authors
- Nantheera Anantrasirichai, Fan Zhang, David Bull
- Published
- 2025-01-06 · arXiv
- Language
- en
- Length
- 42227 words
- Type
- narrative text
Cites 155 works
inferredOne description, two standard projections
Built from what the sources declared and what the gates observed. Nothing absent has been filled in here. 1 value read out of the text by the enrichment rules and 155 links to or from other resources — a lab's files, the pages it links, the works it cites stand under their elements, marked inferred, and are kept apart in the exports.
- Profile
aimpro-oer-profile/1built 2026-10-09- Content language
- en declared English
Identity 3/3
Responsibility 1/3
Content 3/5
Rights 3/4
Form and dates 4/10
Relations and custody 1/9
1 General 5/8
2 Life cycle 3/3
3 Meta-metadata 4/4
4 Technical 3/7
5 Educational 1/11
6 Rights 3/3
7 Relation 0/2
8 Annotation 0/3
9 Classification 0/4
What could not be established3
| Status | Field | Why |
|---|---|---|
| Not available | rights_holder |
no rights holder is named at source; the licence is recorded without one rather than attributed to the platform that served it |
| Not available | publisher |
the source named no publisher of the work; where it was collected from is recorded as collection provenance instead, which is a different claim |
| Not available | educational |
the source declared no educational metadata — no resource type, audience, context, difficulty or learning time. Nothing here estimates them |