OER·harvester

← Back to the library
arXiv HTML resource

Advances in Artificial Intelligence: A Review for the Creative Industries

Artificial intelligence (AI) has undergone transformative advances since 2022, particularly through generative AI, large language models (LLMs), and diffusion models, fundamentally reshaping the creative industries. However, existing reviews have not comprehensively addressed these recent breakthroughs and their integrated impact across the creative production pipeline. This paper addresses this gap by providing a s…

Licence
OPEN CC-BY-4.0
Authors
Nantheera Anantrasirichai, Fan Zhang, David Bull
Published
2025-01-06 · arXiv
Language
en
Length
42227 words
Type
narrative text

Cites 155 works

inferred
Open original ↗

F. Yu, and L. Van Gool Transforming model prediction for tracking. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8721–8730. External Links: Document Cited by: Table 1, §3.4.3.

  • T. Meinhardt, A. Kirillov, L. Leal-Taixé, and C. Feichtenhofer TrackFormer: multi-object tracking with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8844–8854. Cited by: Table 1, §3.4.3.
  • L. Melas-Kyriazi, I. Laina, C. Rupprecht, and A. Vedaldi RealFusion 360${}^{\circ}$ reconstruction of any object from a single image. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8446–8455. External Links: Document Cited by: Table 1, §3.1.5.
  • F. Mentzer, G. D. Toderici, D. Minnen, S. Caelles, S. J. Hwang, M. Lucic, and E. Agustsson VCT: a video compression transformer. Advances in Neural Information Processing Systems 35, pp. 13091–13103. Cited by: Table 1, §3.6.2.
  • F. Mentzer, G. D. Toderici, M. Tschannen, and E. Agustsson High-fidelity generative image compression. Advances in Neural Information Processing Systems 33, pp. 11913–11924. Cited by: §3.6.1.
  • D. Metzler, Y. Tay, D. Bahri, and M. Najork Rethinking search: making domain experts out of dilettantes. SIGIR Forum 55 (1). External Links: Document Cited by: Table 1, §3.2.3.
  • B. Mildenhall, P. Hedman, R. Martin-Brualla, P. P. Srinivasan, and J. T. Barron NeRF in the Dark: High dynamic range view synthesis from noisy raw images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16190–16199. Cited by: Table 1, §3.5.2.
  • B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng NeRF: Representing scenes as neural radiance fields for view synthesis. In Computer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J. Frahm (Eds.), pp. 405–421. Cited by: §2.4, Table 1, Figure 7, §3.5.2.
  • P. Mirowski, K. W. Mathewson, J. Pittman, and R. Evans Co-writing screenplays and theatre scripts with language models: evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, External Links: ISBN 9781450394215, Document Cited by: §3.1.1.
  • T. Miyata ZEN-iqa: zero-shot explainable and no-reference image quality assessment with vision language model. IEEE Access 12, pp. 70973–70983. Cited by: §3.7.1.
  • E. Molad, E. Horwitz, D. Valevski, A. R. Acha, Y. Matias, Y. Pritch, Y. Leviathan, and Y. Hoshen Dreamix: video diffusion models are general video editors. arXiv: 2302.01329. Cited by: Table 1, §3.1.4.
  • E. Moliner, J. Lehtinen, and V. Välimäki Solving audio inverse problems with a diffusion model. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , pp. 1–5. External Links: Document Cited by: Table 1, Table 1, §3.3.4.
  • J. Moon, T. Moon, and W. Seo Generalizable style transfer for implicit neural representation. In International Conference on Learning Representations (ICLR), Cited by: Table 1, Table 1, §3.3.2.
  • C. Morris, N. Anantrasirichai, F. Zhang, and D. Bull DaBiT: Depth and blur informed transformer for video deblurring. In IEEE/CVF Winter Conference on Applications of Computer Vision Workshop, Cited by: Table 1, §3.3.4.
  • B. B. Moser, A. S. Shanbhag, F. Raue, S. Frolov, S. Palacio, and A. Dengel Diffusion models, image super-resolution, and everything: a survey. IEEE Transactions on Neural Networks and Learning Systems (), pp. 1–21. External Links: Document Cited by: §3.3.3.
  • T. Müller, A. Evans, C. Schied, and A. Keller Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph. 41 (4), pp. 102:1–102:15. External Links: Document Cited by: Table 1, §3.1.5, §3.5.2.
  • N. G. Nair, K. Mei, and V. M. Patel AT-DDPM: Restoring faces degraded by atmospheric turbulence using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 3434–3443. Cited by: Table 1, §3.3.4.
  • J. Nawała, Y. Jiang, F. Zhang, X. Zhu, J. Sole, and D. Bull BVI-AOM: a new training dataset for deep video compression optimization. In IEEE Visual Communications and Image Processing (VCIP)), Cited by: §3.6.2.
  • D. Neimark, O. Bar, M. Zohar, and D. Asselmann Video transformer network. In 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Vol. , pp. 3156–3165. External Links: Document Cited by: Table 1.
  • S. Oh, H. Yang, and E. Park Parameter-efficient instance-adaptive neural video compression. In Asian Conference on Computer Vision, External Links: Document Cited by: §3.6.2.
  • OpenAI, J. Achiam, S. Adler, S. Agarwal, et al. GPT-4 technical report. arXiv: 2303.08774. Cited by: §1, Table 1, §3.7.1.
  • M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P. Huang, S. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856 Cited by: Table 1, Table 1, Table 1, §3.4, §3.5.1.
  • Y. Ou, Z. Ma, T. Liu, and Y. Wang Perceptual quality assessment of video considering both frame rate and quantization artifacts. IEEE Transactions on Circuits and Systems for Video Technology 21 (3), pp. 286–298. Cited by: §3.7.1.
  • J. Pan, B. Xu, J. Dong, J. Ge, and J. Tang Deep discriminative spatial and temporal network for efficient video deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22191–22200. Cited by: §3.3.4.
  • W. Pan, Z. Yang, D. Liu, C. Fang, Y. Zhang, and P. Dai Quality-aware clip for blind image quality assessment. In Chinese Conference on Pattern Recognition and Computer Vision (PRCV), pp. 396–408. Cited by: §3.7.1.
  • T. Peng, C. Feng, D. Danier, F. Zhang, B. Vallade, A. Mackin, and D. Bull RMT-BVQA: recurrent memory transformer-based blind video quality assessment for enhanced video content. In European Conference on Computer Vision (ECCV) Workshop on Advances in Image Manipulation, Cited by: Table 1, §3.7.1.
  • T. Peng, G. Gao, H. Sun, F. Zhang, and D. Bull Accelerating learnt video codecs with gradient decay and layer-wise distillation. In 2024 Picture Coding Symposium (PCS), pp. 1–5. Cited by: §3.6.2.
  • N. Ponomarenko, O. Ieremeiev, V. Lukin, K. Egiazarian, L. Jin, J. Astola, B. Vozel, K. Chehdi, M. Carli, F. Battisti, and C.-C. J. Kuo Color image database tid2013: peculiarities and preliminary results. In European Workshop on Visual Information Processing (EUVIP), Vol. , pp. 106–111. External Links: Document Cited by: §3.7.1.
  • A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer D-NeRF: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: Table 1, §3.5.2.
  • L. Qi, Z. Jia, J. Li, B. Li, H. Li, and Y. Lu Long-term temporal context gathering for neural video compression. In European Conference on Computer Vision (ECCV), Cited by: §3.6.2.
  • G. Qian, J. Mai, A. Hamdi, J. Ren, A. Siarohin, B. Li, H. Lee, I. Skorokhodov, P. Wonka, S. Tulyakov, and B. Ghanem Magic123: One image to high-quality 3D object generation using both 2d and 3D diffusion priors. In The Twelfth International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.1.5.
  • W. Quan, J. Chen, Y. Liu, and et al. Deep learning-based image and video inpainting: a survey. International Journal of Computer Vision. Cited by: §3.3.5.
  • A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, Cited by: §3.1.1, §3.7.1.
  • S. Rajput, N. Mehta, A. Singh, R. Hulikal Keshavan, T. Vu, et al. Recommender systems with generative retrieval. In Advances in Neural Information Processing Systems, Vol. 36, pp. 10299–10315. Cited by: Table 1, §3.2.3.
  • D. Rao, T. Xu, and X. Wu TGFuse: An infrared and visible image fusion approach based on transformer and generative adversarial network. IEEE Transactions on Image Processing (), pp. 1–1. External Links: Document Cited by: Table 1, §3.3.6.
  • N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer SAM 2: segment anything in images and videos. arXiv preprint arXiv:2408.00714. Cited by: Table 1, §3.4.1.
  • S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas A generalist agent. Transactions on Machine Learning Research. External Links: ISSN 2835-8856 Cited by: §2.1.
  • A. Rehman, K. Zeng, and Z. Wang Display device-adapted video quality-of-experience assessment. In Human vision and electronic imaging XX, Vol. 9394, pp. 27–37. Cited by: §3.7.1.
  • J. Ren, L. Pan, J. Tang, C. Zhang, A. Cao, G. Zeng, and Z. Liu DreamGaussian4D: Generative 4d gaussian splatting. arXiv preprint arXiv:2312.17142. Cited by: Table 1, §3.1.5.
  • J. Ren, Q. Zheng, Y. Zhao, X. Xu, and C. Li DLFormer: Discrete latent transformer for video inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3511–3520. Cited by: Table 1, §3.3.5.
  • S. Ren, K. He, R. Girshick, and J. Sun Faster r-cnn: towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 39 (6), pp. 1137–1149. External Links: Document Cited by: §3.4.2.
  • Y. Ren, X. Xia, Y. Lu, J. Zhang, J. Wu, P. Xie, X. Wang, and X. Xiao Hyper-SD: Trajectory segmented consistency model for efficient image synthesis. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.1.3.
  • R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 10674–10685. External Links: Document Cited by: Figure 1, §2.3, §2.3, Table 1.
  • H. Ruan, Y. Shao, Q. Yang, L. Zhao, and D. Niyato Point cloud compression with implicit neural representations: a unified framework. arXiv preprint arXiv:2405.11493. Cited by: Table 1, §3.6.2.
  • J. Ryu, K. Kim, D. Heo, H. Song, C. Oh, and B. Suh Cinema multiverse lounge: enhancing film appreciation via multi-agent conversations. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25. External Links: Document Cited by: §3.2.2.
  • N. Safonov, A. Bryncev, A. Moskalenko, D. Kulikov, et al. NTIRE 2025 challenge on UGC video enhancement: methods and results. arXiv preprint arXiv:2505.03007. Cited by: §3.3.1.
  • C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (4), pp. 4713–4726. External Links: Document Cited by: Table 1, Table 1, §3.3.3.
  • R. Sajja, Y. Sermet, M. Cikmaz, D. Cwiertny, and I. Demir Artificial intelligence-enabled intelligent assistant for personalized and adaptive learning in higher education. Information 15 (10), pp. 596. External Links: Document Cited by: §3.2.4.
  • Sara Fridovich-Keil and Giacomo Meanti, F. R. Warburg, B. Recht, and A. Kanazawa K-planes: explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.5.2.
  • V. Saragadam, D. LeJeune, J. Tan, G. Balakrishnan, A. Veeraraghavan, and R. G. Baraniuk WIRE: wavelet implicit neural representations. In Conf. Computer Vision and Pattern Recognition, Cited by: §2.4, §3.3.4.
  • J. L. Schönberger and J. Frahm Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §3.5.2.
  • J. Selva, A. S. Johansen, S. Escalera, K. Nasrollahi, T. B. Moeslund, and A. Clapes Video transformers: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), pp. 12922–12943. External Links: ISSN 1939-3539, Document Cited by: §2.1.
  • K. Seshadrinathan, R. Soundararajan, A. C. Bovik, and L. K. Cormack Study of subjective and objective quality assessment of video. IEEE Trans. on Image Processing 19 (6), pp. 1427–1441. Cited by: §3.7.1.
  • S. C. Shapiro and D. Eckroth (Eds.) Encyclopedia of artificial intelligence. Vol. 1, John Wiley & Sons, New York. Cited by: §3.1.
  • H. R. Sheikh, M. F. Sabir, and A. C. Bovik A statistical evaluation of recent full reference image quality assessment algorithms. IEEE Transactions on image processing 15 (11), pp. 3440–3451. Cited by: §3.7.1.
  • J. Shi, P. Gao, and J. Qin Transformer-based no-reference image quality assessment via supervised contrastive learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 4829–4837. Cited by: Table 1, §3.7.1.
  • X. Shi, Z. Huang, F. Wang, W. Bian, D. Li, Y. Zhang, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li Motion-I2V: Consistent and controllable image-to-video generation with explicit motion modeling. In ACM SIGGRAPH Conference Papers, External Links: Document Cited by: Table 1, Table 1, §3.3.7.
  • Y. Shi, B. Xia, X. Jin, X. Wang, T. Zhao, X. Xia, X. Xiao, and W. Yang VmambaIR: Visual state space model for image restoration. IEEE Transactions on Circuits and Systems for Video Technology (). External Links: Document Cited by: Table 1, §3.3.4.
  • Y. Shi, H. Ma, W. Zhong, Q. Tan, G. Mai, X. Li, T. Liu, and J. Huang ChatGraph: Interpretable text classification by converting chatgpt knowledge to graphs. In 2023 IEEE International Conference on Data Mining Workshops (ICDMW), Vol. , pp. 515–520. External Links: Document Cited by: Table 1, §3.2.1.
  • U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman Make-A-Video: Text-to-video generation without text-video data. In The Eleventh International Conference on Learning Representations, Cited by: Table 1, §3.1.4.
  • Z. Sinno and A. C. Bovik Large-scale study of perceptual video quality. IEEE Trans. on Image Processing 28 (2), pp. 612–627. Cited by: §3.7.1.
  • V. Sitzmann, J. Martel, A. Bergman, D. Lindell, and G. Wetzstein Implicit neural representations with periodic activation functions. Advances in neural information processing systems 33, pp. 7462–7473. Cited by: Table 1, §3.6.1.
  • V. Sitzmann, J. N.P. Martel, A. W. Bergman, D. B. Lindell, and G. Wetzstein Implicit neural representations with periodic activation functions. In Proc. NeurIPS, Cited by: §2.4.
  • H. Siuzdak, F. Grötschla, and L. A. Lanzendörfer SNAC: multi-scale neural audio codec. In Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation, Note: https://openreview.net/forum?id=PFBF5ctj4X Cited by: §3.6.3.
  • J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, Vol. 37, pp. 2256–2265. Cited by: §2.3.
  • Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, Cited by: §2.3.
  • Y. Song, Z. He, H. Qian, and X. Du Vision transformers for single image dehazing. IEEE Transactions on Image Processing 32 (), pp. 1927–1941. External Links: Document Cited by: Table 1, §3.3.4.
  • Z. Song, L. Liu, F. Jia, Y. Luo, C. Jia, G. Zhang, L. Yang, and L. Wang Robustness-aware 3d object detection in autonomous driving: a review and outlook. IEEE Transactions on Intelligent Transportation Systems 25 (11), pp. 15407–15436. External Links: Document Cited by: §3.4.2.
  • M. Stefanini, M. Cornia, L. Baraldi, S. Cascianelli, G. Fiameni, and R. Cucchiara From show to tell: a survey on deep learning-based image captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (1), pp. 539–559. External Links: Document Cited by: §3.1.1.
  • Y. Strümpler, J. Postels, R. Yang, L. V. Gool, and F. Tombari Implicit neural representations for image compression. In European Conference on Computer Vision, pp. 74–91. Cited by: Table 1, §3.6.1.
  • L. Sun, H. Guo, B. Ren, et al. The tenth ntire 2025 image denoising challenge report. arXiv preprint arXiv:2504.12276. Cited by: §3.3.4.
  • X. Sun, X. Li, J. Li, F. Wu, S. Guo, T. Zhang, and G. Wang Text classification via large language models. In The 2023 Conference on Empirical Methods in Natural Language Processing, Cited by: Table 1, §3.2.1.
  • H. Tan, G. Xiang, X. Xie, and H. Jia Joint frame-level and block-level rate-perception optimized preprocessing for video coding. In Proceedings of the 6th ACM International Conference on Multimedia in Asia, Cited by: §3.6.2.
  • J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng DreamGaussian: generative gaussian splatting for efficient 3D content creation. In The Twelfth International Conference on Learning Representations, Cited by: Table 1, Table 1, §3.1.5.
  • X. Tang, M. Yang, P. Sun, H. Li, Y. Dai, F. Zhu, and H. Lee PaReNeRF: toward fast large-scale dynamic nerf with patch-based reference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5428–5438. Cited by: Table 1, §3.5.2.
  • Q. Team Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. Cited by: §4.3.
  • Y. Tian, Q. Ye, and D. Doermann YOLOv12: attention-centric real-time object detectors. arXiv preprint arXiv:2502.12524. Cited by: Table 1, §3.4.2.
  • R. Tous Lester: Rotoscope animation through video object segmentation and tracking. Algorithms 17 (8). External Links: Document Cited by: §3.3.7.
  • H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: open and efficient foundation language models. arXiv preprint arXiv:2302.13971. Cited by: §3.7.1.
  • A. van den Oord, O. Vinyals, and K. Kavukcuoglu Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 6309–6318. Cited by: §3.1.2, §3.6.3.
  • A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30, pp. . Cited by: Figure 1, §2.1, Table 1, §3.4.3.
  • R. Villegas, M. Babaeizadeh, P. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, Cited by: Table 1, §3.1.4.
  • H. L. F. von Helmholtz Handbook of physiological optics. 1st edition, Voss, Hamburg and Leipzig, Germany. Cited by: §3.7.1.
  • P. V. Vu, C. T. Vu, and D. M. Chandler A spatiotemporal most-apparent-distortion model for video quality assessment. In IEEE ICIP, pp. 2505–2508. Cited by: §3.7.1.
  • A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding YOLOv10: Real-time end-to-end object detection. In Advances in Neural Information Processing Systems, Cited by: §3.4.2.
  • D. Wang, X. Cui, X. Chen, Z. Zou, T. Shi, S. Salcudean, Z. J. Wang, and R. Ward Multi-view 3D reconstruction with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5722–5731. Cited by: Table 1.
  • H. Wang, N. Anantrasirichai, F. Zhang, and D. Bull UW-GS: Distractor-aware 3d gaussian splatting for enhanced underwater scene reconstruction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Cited by: Table 1, §3.5.3.
  • J. Wang, Z. Lin, M. Wei, Y. Zhao, C. Yang, C. C. Loy, and L. Jiang SeedVR: Seeding infinity in diffusion transformer towards generic video restoration. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, Table 1.
  • J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang ModelScope text-to-video technical report. arXiv: 2308.06571. Cited by: Table 1, §3.1.4.
  • T. Wang, L. Li, K. Lin, Y. Zhai, C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang DisCo: disentangled control for realistic human dance generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9326–9336. Cited by: Table 1, §3.1.4.
  • T. Wang, K. Zhang, T. Shen, W. Luo, B. Stenger, and T. Lu Ultra-high-definition low-light image enhancement: a benchmark and transformer-based method. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 2654–2662. Cited by: Table 1, §3.3.1.
  • X. Wang, W. Wang, Y. Cao, C. Shen, and T. Huang Images speak in images: a generalist painter for in-context visual learning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 6830–6839. External Links: Document Cited by: §1, Table 1, Table 1.
  • X. Wang, X. Zhang, Y. Cao, W. Wang, C. Shen, and T. Huang SegGPT: towards segmenting everything in context. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 1130–1140. External Links: Document Cited by: Table 1, §3.4.1.
  • Y. Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y. Wang, C. Yang, Y. He, J. Yu, P. Yang, Y. Guo, T. Wu, C. Si, Y. Jiang, C. Chen, C. C. Loy, B. Dai, D. Lin, Y. Qiao, and Z. Liu LaVie: High-quality video generation with cascaded latent diffusion models. International Journal of Computer Vision 133 (5), pp. 3059–3078. External Links: Document Cited by: Table 1, §3.1.4.
  • Y. Wang, S. Inguva, and B. Adsumilli YouTube UGC dataset for video compression research. In 2019 IEEE 21st International Workshop on Multimedia Signal Processing (MMSP), pp. 1–5. Cited by: §3.7.1.
  • Y. Wang, T. Isobe, X. Jia, X. Tao, H. Lu, and Y. Tai Compression-aware video super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2012–2021. Cited by: §3.6.2.
  • Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. 13 (4), pp. 600–612. Cited by: §3.7.1.
  • Z. Wang, E. P. Simoncelli, and A. C. Bovik Multi-scale structural similarity for image quality assessment. In Asilomar Conf. Signals Syst. Comput., pp. 1398–1402. Cited by: §3.6.1.
  • Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li Uformer: a general u-shaped transformer for image restoration. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 17662–17672. Cited by: Table 1, §3.3.4, §3.3.4.
  • Z. Wang, Q. Xie, T. Li, H. Du, L. Xie, P. Zhu, and M. Bi One-shot voice conversion for style transfer based on speaker adaptation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6792–6796. External Links: Document Cited by: §3.1.2.
  • Z. Wang, E. P. Simoncelli, and A. C. Bovik Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, Vol. 2, pp. 1398–1402. Cited by: §3.7.1.
  • Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao SIMVLM: Simple visual language model pretraining with weak supervision. In International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.1.1.
  • H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, J. Yang, J. Sun, C. Han, and X. Zhang Vary: Scaling up the vision vocabulary for large vision-language models. In European Conference on Computer Vision, Cited by: Table 1, §3.1.1.
  • S. Woo, J. Park, J. Lee, and I. S. Kweon CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Cited by: §2.1.
  • G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and W. Xinggang 4D gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.5.3.
  • H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin Fast-VQA: efficient end-to-end video quality assessment with fragment sampling. In European conference on computer vision, pp. 538–554. Cited by: Table 1, §3.7.1.
  • H. Wu, L. Liao, J. Hou, C. Chen, E. Zhang, A. Wang, W. Sun, Q. Yan, and W. Lin Exploring opinion-unaware video quality assessment with semantic affinity criterion. In 2023 IEEE International Conference on Multimedia and Expo (ICME), pp. 366–371. Cited by: §3.7.1.
  • H. Wu, L. Liao, A. Wang, C. Chen, J. Hou, W. Sun, Q. Yan, and W. Lin Towards robust text-prompted semantic criterion for in-the-wild video quality assessment. arXiv preprint arXiv:2304.14672. Cited by: §3.7.1, §3.7.1.
  • H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20144–20154. Cited by: Table 1, §3.7.1.
  • H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, et al. Q-bench: a benchmark for general-purpose foundation models on low-level vision. In The Twelfth International Conference on Learning Representations, Cited by: §3.7.1.
  • H. Wu, Z. Zhang, W. Zhang, C. Chen, L. Liao, C. Li, Y. Gao, A. Wang, E. Zhang, W. Sun, et al. Q-Align: teaching LMMs for visual scoring via discrete text-defined levels. In International Conference on Machine Learning, Cited by: §3.7.1.
  • J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7623–7633. Cited by: Table 1, §3.1.4.
  • T. Wu, Y. Zhang, X. Wang, X. Zhou, G. Zheng, Z. Qi, Y. Shan, and X. Li CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 8469–8477. External Links: Document Cited by: Table 1, §3.1.4.
  • T. Wu, S. He, J. Liu, S. Sun, K. Liu, Q. Han, and Y. Tang A brief overview of ChatGPT: the history, status quo and potential future development. IEEE/CAA Journal of Automatica Sinica 10 (5), pp. 1122–1136. External Links: Document Cited by: §1.
  • W. Wu, Y. Zhao, H. Chen, Y. Gu, R. Zhao, Y. He, H. Zhou, M. Z. Shou, and C. Shen DatasetDM: Synthesizing data with perception annotations using diffusion models. In Advances in Neural Information Processing Systems, Vol. 36, pp. 54683–54695. Cited by: §3.4.2.
  • W. Wu, Y. Zhao, M. Z. Shou, H. Zhou, and C. Shen DiffuMask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1206–1217. Cited by: Table 1, §3.4.1.
  • Z. Wu, L. Zhu, Z. Yin, X. Xu, J. Zhu, X. Wei, and X. Yang MAFCD: Multi-level and adaptive conditional diffusion model for anomaly detection. Information Fusion 118, pp. 102965. Note: https://www.sciencedirect.com/science/article/pii/S1566253525000387 External Links: ISSN 1566-2535, Document Cited by: Table 1, §3.4.2.
  • J. Wynn and D. Turmukhambetov DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.5.2.
  • J. Xiang, K. Tian, and J. Zhang Mimt: masked image modeling transformer for video compression. In International Conference on Learning Representations, Cited by: Table 1, §3.6.2.
  • F. Xie, Z. Wang, and C. Ma DiffusionTrack: Point set diffusion model for visual object tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 19113–19124. Cited by: Table 1, §3.4.3.
  • Y. Xie, T. Takikawa, S. Saito, O. Litany, S. Yan, N. Khan, F. Tombari, J. Tompkin, V. Sitzmann, and S. Sridhar Neural fields in visual computing and beyond. In Computer Graphics Forum, Vol. 41(2), pp. 641–676. Cited by: §2.4, §3.5.2.
  • D. Xu, Y. Jiang, P. Wang, Z. Fan, Y. Wang, and Z. Wang NeuralLift-360: Lifting an in-the-wild 2d photo to a 3D object with 360° views. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4479–4489. External Links: Document Cited by: Table 1, §3.1.5.
  • F. Xu, T. Zhou, T. Nguyen, H. Bao, C. Lin, and J. Du Integrating augmented reality and llm for enhanced cognitive support in critical audio communications. International Journal of Human-Computer Studies 194, pp. 103402. External Links: ISSN 1071-5819, Document Cited by: §3.1.5.
  • J. Xu, X. Hu, L. Zhu, Q. Dou, J. Dai, Y. Qiao, and P. Heng Video dehazing via a multi-range temporal alignment network with physical prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18053–18062. Cited by: Table 1, §3.3.4.
  • J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello Open-vocabulary panoptic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2955–2966. Cited by: Table 1, §3.4.1.
  • M. Xu, J. Chen, H. Wang, S. Liu, G. Li, and Z. Bai C3DVQA: full-reference video quality assessment with 3d convolutional neural network. In IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 4447–4451. Cited by: §3.7.1.
  • S. Xu, G. Chen, Y. Guo, J. Yang, C. Li, Z. Zang, Y. Zhang, X. Tong, and B. Guo VASA-1: lifelike audio-driven talking faces generated in real time. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: Table 1, Figure 4, §3.1.4.
  • X. Xu, R. Wang, C. Fu, and J. Jia SNR-aware low-light image enhancement. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 17693–17703. External Links: Document Cited by: Table 1, §3.3.1.
  • Y. Xu, T. Park, R. Zhang, Y. Zhou, E. Shechtman, F. Liu, J. Huang, and D. Liu VideoGigaGAN: Towards detail-rich video super-resolution. arXiv:2404.12388. Cited by: Table 1, §3.3.3.
  • B. Yan, Y. Jiang, J. Wu, D. Wang, P. Luo, Z. Yuan, and H. Lu Universal instance perception as object discovery and retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 15325–15336. Cited by: Table 1, Figure 6, §3.4.
  • C. Yang, L. Liang, and Z. Su Real-world denoising via diffusion model. arXiv preprint arXiv:2305.04457. Cited by: Table 1, §3.3.4.
  • D. Yang, H. Guo, Y. Wang, R. Huang, X. Li, X. Tan, X. Wu, and H. M. Meng UniAudio 1.5: large language model-driven audio codec is a few-shot audio task learner. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Note: https://openreview.net/forum?id=NGrINZyZKk Cited by: §3.6.3.
  • D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, X. Wu, et al. UniAudio: an audio foundation model toward universal audio generation. arXiv preprint arXiv:2310.00704. Cited by: §3.6.3.
  • D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu Diffsound: discrete diffusion model for text-to-sound generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (), pp. 1720–1733. External Links: Document Cited by: Table 1, §3.1.2.
  • H. Yang, L. Pan, Y. Yang, R. Hartley, and M. Liu LDP: language-driven dual-pixel image defocus deblurring network. arXiv: 2307.09815. External Links: Cited by: Table 1, §3.3.4.
  • J. Yang Exploring the productivity of generative ai-powered ad campaigns: a consumer response perspective. Journal of Interactive Advertising 25 (3), pp. 222–239. External Links: Document Cited by: §1.
  • J. Yang, M. Gao, Z. Li, S. Gao, F. Wang, and F. Zheng Track anything: segment anything meets videos. arXiv:2304.11968. Cited by: Table 1, §3.4.3.
  • L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao Depth anything: unleashing the power of large-scale unlabeled data. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.4.2, §3.5.1.
  • L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao Depth anything v2. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.5.1.
  • L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M. Yang Diffusion models: a comprehensive survey of methods and applications. ACM Comput. Surv. 56 (4). External Links: Document Cited by: Figure 1.
  • R. Yang and S. Mandt Lossy image compression with conditional diffusion models. Advances in Neural Information Processing Systems 36. Cited by: Table 1, §3.6.1.
  • S. Yang, M. Ding, Y. Wu, Z. Li, and J. Zhang Implicit neural representation for cooperative low-light image enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12918–12927. Cited by: Table 1, §3.3.1.
  • Y. Yang, L. Pan, L. Liu, and M. Liu K3DN: Disparity-aware kernel estimation for dual-pixel defocus deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13263–13272. Cited by: §3.3.4.
  • Y. Yang, F. Sun, L. Weihs, E. VanderBilt, A. Herrasti, W. Han, J. Wu, N. Haber, R. Krishna, L. Liu, C. Callison-Burch, M. Yatskar, A. Kembhavi, and C. Clark Holodeck: Language guided generation of 3D embodied AI environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.1.5.
  • Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang A survey on large language model (llm) security and privacy: the good, the bad, and the ugly. High-Confidence Computing 4 (2), pp. 100211. External Links: Document Cited by: §2.2.
  • S. Ye, D. Kim, S. Kim, H. Hwang, S. Kim, Y. Jo, J. Thorne, J. Kim, and M. Seo FLASK: fine-grained language model evaluation based on alignment skill sets. In The Twelfth International Conference on Learning Representations, Note: https://openreview.net/forum?id=CYmF38ysDa Cited by: Figure 2, §2.2.
  • A. Yi and N. Anantrasirichai A comprehensive study of object tracking in low-light environments. Sensors 24 (14). Note: https://arxiv.org/pdf/2312.16250.pdf Cited by: Table 1, §3.4.3.
  • X. Yi, H. Xu, H. Zhang, L. Tang, and J. Ma Diff-retinex: rethinking low-light image enhancement with a generative diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12302–12311. Cited by: Table 1, §3.3.1.
  • Y. Yin, D. Xu, C. Tan, P. Liu, Y. Zhao, and Y. Wei CLE Diffusion: Controllable light enhancement diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia, pp. 8145–8156. External Links: Document Cited by: Table 1, §3.3.1.
  • G. Youk, J. Oh, and M. Kim FMA-Net: Flow-guided dynamic filtering and iterative feature refinement with multi-attention for joint video super-resolution and deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.3.4.
  • G. Yu, A. Li, H. Wang, Y. Wang, Y. Ke, and C. Zheng DBT-net: dual-branch federative magnitude and phase estimation with attention-in-attention transformer for monaural speech enhancement. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30 (), pp. 2629–2644. External Links: Document Cited by: Table 1, §3.3.4.
  • H. Yu, J. Julin, Z. A. Milacski, K. Niinuma, and L. A. Jeni CoGS: Controllable gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 21624–21633. Cited by: Table 1, §3.5.3.
  • W. Yu, L. Po, R. C.C. Cheung, Y. Zhao, Y. Xue, and K. Li Bidirectionally deformable motion modulation for video-based human pose transfer. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 7468–7478. External Links: Document Cited by: Table 1, §3.1.4.
  • H. Yue, C. Cao, L. Liao, and J. Yang RViDeformer: Efficient raw video denoising transformer with a larger benchmark dataset. IEEE Transactions on Circuits and Systems for Video Technology (), pp. 1–1. External Links: Document Cited by: Table 1, §3.3.4.
  • S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. Yang Restormer: efficient transformer for high-resolution image restoration. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 5718–5729. Cited by: Table 1, §3.3.4.
  • N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi Soundstream: an end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30, pp. 495–507. Cited by: §3.6.3.
  • F. Zeng, B. Dong, Y. Zhang, T. Wang, X. Zhang, and Y. Wei MOTR: end-to-end multiple-object tracking with transformer. In European Conference on Computer Vision (ECCV), Cited by: Table 1, §3.4.3.
  • G. Zhai and X. Min Perceptual image quality assessment: a survey. Science China Information Sciences 63, pp. 1–52. Cited by: §3.7.
  • Y. Zhan, Z. Li, M. Niu, Z. Zhong, S. Nobuhara, K. Nishino, and Y. Zheng KFD-NeRF: rethinking dynamic nerf with kalman filter. In Proceedings of the European Conference on Computer Vision (ECCV), Cited by: Table 1, §3.5.2.
  • F. Zhang and D. R. Bull A perception-based hybrid model for video quality assessment. IEEE Transactions on Circuits and Systems for Video Technology 26 (6), pp. 1017–1028. Cited by: §3.7.1.
  • H. Zhang, C. Jung, D. Zou, and M. Li WCDANN: a lightweight CNN post-processing filter for VVC-based video compression. IEEE Access. Cited by: §3.6.2.
  • H. Zhang, Z. Wang, D. Zeng, Z. Wu, and Y. Jiang DiffusionAD: Norm-guided one-step denoising diffusion for anomaly detection. IEEE Transactions on Pattern Analysis and Machine Intelligence (), pp. 1–13. External Links: Document Cited by: Table 1, §3.4.2.
  • J. Zhang, J. Huang, S. Jin, and S. Lu Vision-language models for vision tasks: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (), pp. 1–20. External Links: Document Cited by: §3.1.1.
  • N. Zhang, F. Nex, G. Vosselman, and N. Kerle Lite-Mono: A lightweight cnn and transformer architecture for self-supervised monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18537–18546. Cited by: Table 1, §3.5.1.
  • R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Cited by: §3.7.1.
  • R. Zhang, D. Cai, L. Qian, Y. Du, H. Lu, and Y. Zhang DiffusionTracker: Targets denoising based on diffusion model for visual tracking. In Pattern Recognition and Computer Vision, pp. 225–237. Cited by: Table 1, §3.4.3.
  • T. Zhang, X. Tian, Y. Zhou, S. Ji, X. Wang, X. Tao, Y. Zhang, P. Wan, Z. Wang, and Y. Wu DVIS++: Improved decoupled framework for universal video segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence. External Links: Document Cited by: Table 1, §3.4.1.
  • W. Zhang, K. Ma, G. Zhai, and X. Yang Uncertainty-aware blind image quality assessment in the laboratory and wild. IEEE Transactions on Image Processing 30, pp. 3474–3486. Cited by: §3.7.1.
  • X. Zhang and Y. Demiris Visible and infrared image fusion using deep learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (8), pp. 10535–10554. External Links: Document Cited by: §3.3.6.
  • X. Zhang, Z. Mao, N. Chimitt, and S. H. Chan Imaging through the atmosphere using turbulence mitigation transformer. IEEE Transactions on Computational Imaging 10 (), pp. 115–128. External Links: Document Cited by: Table 1, §3.3.4.
  • Y. Zhang, T. Wang, and X. Zhang MOTRv2: Bootstrapping end-to-end multi-object tracking by pretrained object detectors. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 22056–22065. External Links: ISSN , Document Cited by: Table 1, §3.4.3.
  • Y. Zhang, N. Huang, F. Tang, H. Huang, C. Ma, W. Dong, and C. Xu Inversion-based style transfer with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10146–10156. Cited by: Table 1, §3.3.2.
  • Z. Zhang, Z. Hu, W. Deng, C. Fan, T. Lv, and Y. Ding DINet: Deformation inpainting network for realistic face visually dubbing on high resolution video. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 3543–3551. Cited by: §3.3.5.
  • Z. Zhang, Y. Zhou, C. Li, B. Zhao, X. Liu, and G. Zhai Quality assessment in the era of large models: a survey. arXiv preprint arXiv:2409.00031. Cited by: §3.7.
  • K. Zhao, K. Yuan, M. Sun, M. Li, and X. Wen Quality-aware pre-trained models for blind image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 22302–22313. Cited by: §3.7.1.
  • L. Zhao, N. B. Gundavarapu, L. Yuan, H. Zhou, S. Yan, J. J., et al. VideoPrism: A foundational visual encoder for video understanding. In Proceedings of the 41st International Conference on Machine Learning, Cited by: Table 1, §3.4.
  • W. X. Zhao, K. Zhou, J. Li, T. Tang, et al. A survey of large language models. arXiv:2303.18223. Cited by: §2.2, §2.2.
  • Y. Zhao, W. He, C. Jia, Q. Wang, J. Li, Y. Li, C. Lin, K. Zhang, L. Zhang, and S. Ma A neural-network enhanced video coding framework beyond ECM. In 2024 Data Compression Conference (DCC), pp. 605–605. Cited by: §3.6.2.
  • Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and J. Chen DETRs beat yolos on real-time object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §3.4.2.
  • Y. Zhao, M. Dasari, and T. Guo CleAR: robust context-guided generative lighting estimation for mobile augmented reality. arXiv preprint arXiv:2411.02179. Cited by: Table 1, §3.1.5.
  • Z. Zhao, H. Bai, Y. Zhu, J. Zhang, S. Xu, Y. Zhang, K. Zhang, D. Meng, R. Timofte, and L. Van Gool DDFM: Denoising diffusion model for multi-modality image fusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8082–8093. Cited by: Table 1, §3.3.6.
  • C. Zheng, T. Cham, and J. Cai Pluralistic image completion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1438–1447. Cited by: §3.3.5.
  • Q. Zheng, Y. Fan, L. Huang, T. Zhu, J. Liu, Z. Hao, S. Xing, C. Chen, X. Min, A. C. Bovik, et al. Video quality assessment: a comprehensive survey. arXiv preprint arXiv:2412.04508. Cited by: §3.7.
  • L. Zhong, Z. Wang, and J. Shang LDB: A large language model debugger via verifying runtime execution step-by-step. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics: Findings., Cited by: §4.1.
  • F. Zhou, W. Sheng, Z. Lu, and G. Qiu A database and model for the visual quality assessment of super-resolution videos. IEEE Transactions on Broadcasting 70 (2), pp. 516–532. External Links: Document Cited by: §3.7.1.
  • S. Zhou, C. Li, K. C.K. Chan, and C. C. Loy ProPainter: Improving propagation and transformer for video inpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10477–10486. Cited by: Table 1, §3.3.5.
  • S. Zhou, C. Li, and C. Change Loy LEDNet: joint low-light enhancement and deblurring in the dark. In Computer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), pp. 573–589. Cited by: §3.3.1.
  • X. Zhou, Y. Zheng, and J. Yang Bridging the metrics gap in image style transfer: a comprehensive survey of models and criteria. Neurocomputing 624, pp. 129430. External Links: ISSN 0925-2312, Document Cited by: §3.3.2.
  • H. Zhu, X. Sui, B. Chen, X. Liu, Y. Fang, and S. Wang 2AFC prompting of large multimodal models for image quality assessment. IEEE Transactions on Circuits and Systems for Video Technology. Cited by: §3.7.1.
  • K. Zhu, C. Li, V. Asari, and D. Saupe No-reference video quality assessment based on artifact measurement and statistical analysis. IEEE Transactions on Circuits and Systems for Video Technology 25 (4), pp. 533–546. Cited by: §3.7.1.
  • L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang Vision mamba: efficient visual representation learning with bidirectional state space model. Cited by: §2.1.
  • X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai Deformable DETR: Deformable transformers for end-to-end object detection. In International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.4.2.
  • Y. Zhu, Y. Yang, and T. Cohen Transformer-based transform coding. In International Conference on Learning Representations, Cited by: Table 1, §3.6.1.
  • Y. Zhu, L. Zhang, Z. Rong, T. Hu, S. Liang, and Z. Ge INFP: Audio-driven interactive head generation in dyadic conversations. arXiv:2412.04037. Cited by: Table 1, Table 1, §3.1.4.
  • R. Zou, C. Song, and Z. Zhang The devil is in the details: window-based attention for image compression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17492–17501. Cited by: Table 1, §3.6.1.
  • X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y. J. Lee Segment everything everywhere all at once. In Advances in Neural Information Processing Systems, Vol. 36, pp. 19769–19782. Cited by: Table 1, §3.4.1.
  • Z. Zou and N. Anantrasirichai DeTurb: Atmospheric turbulence mitigation with deformable 3d convolutions and 3d swin transformers. In Proceedings of the Asian Conference on Computer Vision (ACCV), Cited by: Table 1, §3.3.4.
  • Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye Object detection in 20 years: a survey. Proceedings of the IEEE 111 (3), pp. 257–276. External Links: Document Cited by: §3.4.2.