F. Yu, and L. Van Gool Transforming model prediction for tracking. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8721–8730. External Links: Document Cited by: Table 1, §3.4.3.
- T. Meinhardt, A. Kirillov, L. Leal-Taixé, and C. Feichtenhofer TrackFormer: multi-object tracking with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8844–8854. Cited by: Table 1, §3.4.3.
- L. Melas-Kyriazi, I. Laina, C. Rupprecht, and A. Vedaldi RealFusion 360${}^{\circ}$ reconstruction of any object from a single image. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 8446–8455. External Links: Document Cited by: Table 1, §3.1.5.
- F. Mentzer, G. D. Toderici, D. Minnen, S. Caelles, S. J. Hwang, M. Lucic, and E. Agustsson VCT: a video compression transformer. Advances in Neural Information Processing Systems 35, pp. 13091–13103. Cited by: Table 1, §3.6.2.
- F. Mentzer, G. D. Toderici, M. Tschannen, and E. Agustsson High-fidelity generative image compression. Advances in Neural Information Processing Systems 33, pp. 11913–11924. Cited by: §3.6.1.
- D. Metzler, Y. Tay, D. Bahri, and M. Najork Rethinking search: making domain experts out of dilettantes. SIGIR Forum 55 (1). External Links: Document Cited by: Table 1, §3.2.3.
- B. Mildenhall, P. Hedman, R. Martin-Brualla, P. P. Srinivasan, and J. T. Barron NeRF in the Dark: High dynamic range view synthesis from noisy raw images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16190–16199. Cited by: Table 1, §3.5.2.
- B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng NeRF: Representing scenes as neural radiance fields for view synthesis. In Computer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J. Frahm (Eds.), pp. 405–421. Cited by: §2.4, Table 1, Figure 7, §3.5.2.
- P. Mirowski, K. W. Mathewson, J. Pittman, and R. Evans Co-writing screenplays and theatre scripts with language models: evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, External Links: ISBN 9781450394215, Document Cited by: §3.1.1.
- T. Miyata ZEN-iqa: zero-shot explainable and no-reference image quality assessment with vision language model. IEEE Access 12, pp. 70973–70983. Cited by: §3.7.1.
- E. Molad, E. Horwitz, D. Valevski, A. R. Acha, Y. Matias, Y. Pritch, Y. Leviathan, and Y. Hoshen Dreamix: video diffusion models are general video editors. arXiv: 2302.01329. Cited by: Table 1, §3.1.4.
- E. Moliner, J. Lehtinen, and V. Välimäki Solving audio inverse problems with a diffusion model. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , pp. 1–5. External Links: Document Cited by: Table 1, Table 1, §3.3.4.
- J. Moon, T. Moon, and W. Seo Generalizable style transfer for implicit neural representation. In International Conference on Learning Representations (ICLR), Cited by: Table 1, Table 1, §3.3.2.
- C. Morris, N. Anantrasirichai, F. Zhang, and D. Bull DaBiT: Depth and blur informed transformer for video deblurring. In IEEE/CVF Winter Conference on Applications of Computer Vision Workshop, Cited by: Table 1, §3.3.4.
- B. B. Moser, A. S. Shanbhag, F. Raue, S. Frolov, S. Palacio, and A. Dengel Diffusion models, image super-resolution, and everything: a survey. IEEE Transactions on Neural Networks and Learning Systems (), pp. 1–21. External Links: Document Cited by: §3.3.3.
- T. Müller, A. Evans, C. Schied, and A. Keller Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph. 41 (4), pp. 102:1–102:15. External Links: Document Cited by: Table 1, §3.1.5, §3.5.2.
- N. G. Nair, K. Mei, and V. M. Patel AT-DDPM: Restoring faces degraded by atmospheric turbulence using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 3434–3443. Cited by: Table 1, §3.3.4.
- J. Nawała, Y. Jiang, F. Zhang, X. Zhu, J. Sole, and D. Bull BVI-AOM: a new training dataset for deep video compression optimization. In IEEE Visual Communications and Image Processing (VCIP)), Cited by: §3.6.2.
- D. Neimark, O. Bar, M. Zohar, and D. Asselmann Video transformer network. In 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Vol. , pp. 3156–3165. External Links: Document Cited by: Table 1.
- S. Oh, H. Yang, and E. Park Parameter-efficient instance-adaptive neural video compression. In Asian Conference on Computer Vision, External Links: Document Cited by: §3.6.2.
- OpenAI, J. Achiam, S. Adler, S. Agarwal, et al. GPT-4 technical report. arXiv: 2303.08774. Cited by: §1, Table 1, §3.7.1.
- M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P. Huang, S. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856 Cited by: Table 1, Table 1, Table 1, §3.4, §3.5.1.
- Y. Ou, Z. Ma, T. Liu, and Y. Wang Perceptual quality assessment of video considering both frame rate and quantization artifacts. IEEE Transactions on Circuits and Systems for Video Technology 21 (3), pp. 286–298. Cited by: §3.7.1.
- J. Pan, B. Xu, J. Dong, J. Ge, and J. Tang Deep discriminative spatial and temporal network for efficient video deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22191–22200. Cited by: §3.3.4.
- W. Pan, Z. Yang, D. Liu, C. Fang, Y. Zhang, and P. Dai Quality-aware clip for blind image quality assessment. In Chinese Conference on Pattern Recognition and Computer Vision (PRCV), pp. 396–408. Cited by: §3.7.1.
- T. Peng, C. Feng, D. Danier, F. Zhang, B. Vallade, A. Mackin, and D. Bull RMT-BVQA: recurrent memory transformer-based blind video quality assessment for enhanced video content. In European Conference on Computer Vision (ECCV) Workshop on Advances in Image Manipulation, Cited by: Table 1, §3.7.1.
- T. Peng, G. Gao, H. Sun, F. Zhang, and D. Bull Accelerating learnt video codecs with gradient decay and layer-wise distillation. In 2024 Picture Coding Symposium (PCS), pp. 1–5. Cited by: §3.6.2.
- N. Ponomarenko, O. Ieremeiev, V. Lukin, K. Egiazarian, L. Jin, J. Astola, B. Vozel, K. Chehdi, M. Carli, F. Battisti, and C.-C. J. Kuo Color image database tid2013: peculiarities and preliminary results. In European Workshop on Visual Information Processing (EUVIP), Vol. , pp. 106–111. External Links: Document Cited by: §3.7.1.
- A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer D-NeRF: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: Table 1, §3.5.2.
- L. Qi, Z. Jia, J. Li, B. Li, H. Li, and Y. Lu Long-term temporal context gathering for neural video compression. In European Conference on Computer Vision (ECCV), Cited by: §3.6.2.
- G. Qian, J. Mai, A. Hamdi, J. Ren, A. Siarohin, B. Li, H. Lee, I. Skorokhodov, P. Wonka, S. Tulyakov, and B. Ghanem Magic123: One image to high-quality 3D object generation using both 2d and 3D diffusion priors. In The Twelfth International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.1.5.
- W. Quan, J. Chen, Y. Liu, and et al. Deep learning-based image and video inpainting: a survey. International Journal of Computer Vision. Cited by: §3.3.5.
- A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, Cited by: §3.1.1, §3.7.1.
- S. Rajput, N. Mehta, A. Singh, R. Hulikal Keshavan, T. Vu, et al. Recommender systems with generative retrieval. In Advances in Neural Information Processing Systems, Vol. 36, pp. 10299–10315. Cited by: Table 1, §3.2.3.
- D. Rao, T. Xu, and X. Wu TGFuse: An infrared and visible image fusion approach based on transformer and generative adversarial network. IEEE Transactions on Image Processing (), pp. 1–1. External Links: Document Cited by: Table 1, §3.3.6.
- N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer SAM 2: segment anything in images and videos. arXiv preprint arXiv:2408.00714. Cited by: Table 1, §3.4.1.
- S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas A generalist agent. Transactions on Machine Learning Research. External Links: ISSN 2835-8856 Cited by: §2.1.
- A. Rehman, K. Zeng, and Z. Wang Display device-adapted video quality-of-experience assessment. In Human vision and electronic imaging XX, Vol. 9394, pp. 27–37. Cited by: §3.7.1.
- J. Ren, L. Pan, J. Tang, C. Zhang, A. Cao, G. Zeng, and Z. Liu DreamGaussian4D: Generative 4d gaussian splatting. arXiv preprint arXiv:2312.17142. Cited by: Table 1, §3.1.5.
- J. Ren, Q. Zheng, Y. Zhao, X. Xu, and C. Li DLFormer: Discrete latent transformer for video inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3511–3520. Cited by: Table 1, §3.3.5.
- S. Ren, K. He, R. Girshick, and J. Sun Faster r-cnn: towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 39 (6), pp. 1137–1149. External Links: Document Cited by: §3.4.2.
- Y. Ren, X. Xia, Y. Lu, J. Zhang, J. Wu, P. Xie, X. Wang, and X. Xiao Hyper-SD: Trajectory segmented consistency model for efficient image synthesis. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.1.3.
- R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 10674–10685. External Links: Document Cited by: Figure 1, §2.3, §2.3, Table 1.
- H. Ruan, Y. Shao, Q. Yang, L. Zhao, and D. Niyato Point cloud compression with implicit neural representations: a unified framework. arXiv preprint arXiv:2405.11493. Cited by: Table 1, §3.6.2.
- J. Ryu, K. Kim, D. Heo, H. Song, C. Oh, and B. Suh Cinema multiverse lounge: enhancing film appreciation via multi-agent conversations. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25. External Links: Document Cited by: §3.2.2.
- N. Safonov, A. Bryncev, A. Moskalenko, D. Kulikov, et al. NTIRE 2025 challenge on UGC video enhancement: methods and results. arXiv preprint arXiv:2505.03007. Cited by: §3.3.1.
- C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (4), pp. 4713–4726. External Links: Document Cited by: Table 1, Table 1, §3.3.3.
- R. Sajja, Y. Sermet, M. Cikmaz, D. Cwiertny, and I. Demir Artificial intelligence-enabled intelligent assistant for personalized and adaptive learning in higher education. Information 15 (10), pp. 596. External Links: Document Cited by: §3.2.4.
- Sara Fridovich-Keil and Giacomo Meanti, F. R. Warburg, B. Recht, and A. Kanazawa K-planes: explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.5.2.
- V. Saragadam, D. LeJeune, J. Tan, G. Balakrishnan, A. Veeraraghavan, and R. G. Baraniuk WIRE: wavelet implicit neural representations. In Conf. Computer Vision and Pattern Recognition, Cited by: §2.4, §3.3.4.
- J. L. Schönberger and J. Frahm Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §3.5.2.
- J. Selva, A. S. Johansen, S. Escalera, K. Nasrollahi, T. B. Moeslund, and A. Clapes Video transformers: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), pp. 12922–12943. External Links: ISSN 1939-3539, Document Cited by: §2.1.
- K. Seshadrinathan, R. Soundararajan, A. C. Bovik, and L. K. Cormack Study of subjective and objective quality assessment of video. IEEE Trans. on Image Processing 19 (6), pp. 1427–1441. Cited by: §3.7.1.
- S. C. Shapiro and D. Eckroth (Eds.) Encyclopedia of artificial intelligence. Vol. 1, John Wiley & Sons, New York. Cited by: §3.1.
- H. R. Sheikh, M. F. Sabir, and A. C. Bovik A statistical evaluation of recent full reference image quality assessment algorithms. IEEE Transactions on image processing 15 (11), pp. 3440–3451. Cited by: §3.7.1.
- J. Shi, P. Gao, and J. Qin Transformer-based no-reference image quality assessment via supervised contrastive learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 4829–4837. Cited by: Table 1, §3.7.1.
- X. Shi, Z. Huang, F. Wang, W. Bian, D. Li, Y. Zhang, M. Zhang, K. C. Cheung, S. See, H. Qin, J. Dai, and H. Li Motion-I2V: Consistent and controllable image-to-video generation with explicit motion modeling. In ACM SIGGRAPH Conference Papers, External Links: Document Cited by: Table 1, Table 1, §3.3.7.
- Y. Shi, B. Xia, X. Jin, X. Wang, T. Zhao, X. Xia, X. Xiao, and W. Yang VmambaIR: Visual state space model for image restoration. IEEE Transactions on Circuits and Systems for Video Technology (). External Links: Document Cited by: Table 1, §3.3.4.
- Y. Shi, H. Ma, W. Zhong, Q. Tan, G. Mai, X. Li, T. Liu, and J. Huang ChatGraph: Interpretable text classification by converting chatgpt knowledge to graphs. In 2023 IEEE International Conference on Data Mining Workshops (ICDMW), Vol. , pp. 515–520. External Links: Document Cited by: Table 1, §3.2.1.
- U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman Make-A-Video: Text-to-video generation without text-video data. In The Eleventh International Conference on Learning Representations, Cited by: Table 1, §3.1.4.
- Z. Sinno and A. C. Bovik Large-scale study of perceptual video quality. IEEE Trans. on Image Processing 28 (2), pp. 612–627. Cited by: §3.7.1.
- V. Sitzmann, J. Martel, A. Bergman, D. Lindell, and G. Wetzstein Implicit neural representations with periodic activation functions. Advances in neural information processing systems 33, pp. 7462–7473. Cited by: Table 1, §3.6.1.
- V. Sitzmann, J. N.P. Martel, A. W. Bergman, D. B. Lindell, and G. Wetzstein Implicit neural representations with periodic activation functions. In Proc. NeurIPS, Cited by: §2.4.
- H. Siuzdak, F. Grötschla, and L. A. Lanzendörfer SNAC: multi-scale neural audio codec. In Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation, Note: https://openreview.net/forum?id=PFBF5ctj4X Cited by: §3.6.3.
- J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, Vol. 37, pp. 2256–2265. Cited by: §2.3.
- Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, Cited by: §2.3.
- Y. Song, Z. He, H. Qian, and X. Du Vision transformers for single image dehazing. IEEE Transactions on Image Processing 32 (), pp. 1927–1941. External Links: Document Cited by: Table 1, §3.3.4.
- Z. Song, L. Liu, F. Jia, Y. Luo, C. Jia, G. Zhang, L. Yang, and L. Wang Robustness-aware 3d object detection in autonomous driving: a review and outlook. IEEE Transactions on Intelligent Transportation Systems 25 (11), pp. 15407–15436. External Links: Document Cited by: §3.4.2.
- M. Stefanini, M. Cornia, L. Baraldi, S. Cascianelli, G. Fiameni, and R. Cucchiara From show to tell: a survey on deep learning-based image captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (1), pp. 539–559. External Links: Document Cited by: §3.1.1.
- Y. Strümpler, J. Postels, R. Yang, L. V. Gool, and F. Tombari Implicit neural representations for image compression. In European Conference on Computer Vision, pp. 74–91. Cited by: Table 1, §3.6.1.
- L. Sun, H. Guo, B. Ren, et al. The tenth ntire 2025 image denoising challenge report. arXiv preprint arXiv:2504.12276. Cited by: §3.3.4.
- X. Sun, X. Li, J. Li, F. Wu, S. Guo, T. Zhang, and G. Wang Text classification via large language models. In The 2023 Conference on Empirical Methods in Natural Language Processing, Cited by: Table 1, §3.2.1.
- H. Tan, G. Xiang, X. Xie, and H. Jia Joint frame-level and block-level rate-perception optimized preprocessing for video coding. In Proceedings of the 6th ACM International Conference on Multimedia in Asia, Cited by: §3.6.2.
- J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng DreamGaussian: generative gaussian splatting for efficient 3D content creation. In The Twelfth International Conference on Learning Representations, Cited by: Table 1, Table 1, §3.1.5.
- X. Tang, M. Yang, P. Sun, H. Li, Y. Dai, F. Zhu, and H. Lee PaReNeRF: toward fast large-scale dynamic nerf with patch-based reference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5428–5438. Cited by: Table 1, §3.5.2.
- Q. Team Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. Cited by: §4.3.
- Y. Tian, Q. Ye, and D. Doermann YOLOv12: attention-centric real-time object detectors. arXiv preprint arXiv:2502.12524. Cited by: Table 1, §3.4.2.
- R. Tous Lester: Rotoscope animation through video object segmentation and tracking. Algorithms 17 (8). External Links: Document Cited by: §3.3.7.
- H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: open and efficient foundation language models. arXiv preprint arXiv:2302.13971. Cited by: §3.7.1.
- A. van den Oord, O. Vinyals, and K. Kavukcuoglu Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 6309–6318. Cited by: §3.1.2, §3.6.3.
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30, pp. . Cited by: Figure 1, §2.1, Table 1, §3.4.3.
- R. Villegas, M. Babaeizadeh, P. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, Cited by: Table 1, §3.1.4.
- H. L. F. von Helmholtz Handbook of physiological optics. 1st edition, Voss, Hamburg and Leipzig, Germany. Cited by: §3.7.1.
- P. V. Vu, C. T. Vu, and D. M. Chandler A spatiotemporal most-apparent-distortion model for video quality assessment. In IEEE ICIP, pp. 2505–2508. Cited by: §3.7.1.
- A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding YOLOv10: Real-time end-to-end object detection. In Advances in Neural Information Processing Systems, Cited by: §3.4.2.
- D. Wang, X. Cui, X. Chen, Z. Zou, T. Shi, S. Salcudean, Z. J. Wang, and R. Ward Multi-view 3D reconstruction with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5722–5731. Cited by: Table 1.
- H. Wang, N. Anantrasirichai, F. Zhang, and D. Bull UW-GS: Distractor-aware 3d gaussian splatting for enhanced underwater scene reconstruction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Cited by: Table 1, §3.5.3.
- J. Wang, Z. Lin, M. Wei, Y. Zhao, C. Yang, C. C. Loy, and L. Jiang SeedVR: Seeding infinity in diffusion transformer towards generic video restoration. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, Table 1.
- J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang ModelScope text-to-video technical report. arXiv: 2308.06571. Cited by: Table 1, §3.1.4.
- T. Wang, L. Li, K. Lin, Y. Zhai, C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang DisCo: disentangled control for realistic human dance generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9326–9336. Cited by: Table 1, §3.1.4.
- T. Wang, K. Zhang, T. Shen, W. Luo, B. Stenger, and T. Lu Ultra-high-definition low-light image enhancement: a benchmark and transformer-based method. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 2654–2662. Cited by: Table 1, §3.3.1.
- X. Wang, W. Wang, Y. Cao, C. Shen, and T. Huang Images speak in images: a generalist painter for in-context visual learning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 6830–6839. External Links: Document Cited by: §1, Table 1, Table 1.
- X. Wang, X. Zhang, Y. Cao, W. Wang, C. Shen, and T. Huang SegGPT: towards segmenting everything in context. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 1130–1140. External Links: Document Cited by: Table 1, §3.4.1.
- Y. Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y. Wang, C. Yang, Y. He, J. Yu, P. Yang, Y. Guo, T. Wu, C. Si, Y. Jiang, C. Chen, C. C. Loy, B. Dai, D. Lin, Y. Qiao, and Z. Liu LaVie: High-quality video generation with cascaded latent diffusion models. International Journal of Computer Vision 133 (5), pp. 3059–3078. External Links: Document Cited by: Table 1, §3.1.4.
- Y. Wang, S. Inguva, and B. Adsumilli YouTube UGC dataset for video compression research. In 2019 IEEE 21st International Workshop on Multimedia Signal Processing (MMSP), pp. 1–5. Cited by: §3.7.1.
- Y. Wang, T. Isobe, X. Jia, X. Tao, H. Lu, and Y. Tai Compression-aware video super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2012–2021. Cited by: §3.6.2.
- Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. 13 (4), pp. 600–612. Cited by: §3.7.1.
- Z. Wang, E. P. Simoncelli, and A. C. Bovik Multi-scale structural similarity for image quality assessment. In Asilomar Conf. Signals Syst. Comput., pp. 1398–1402. Cited by: §3.6.1.
- Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li Uformer: a general u-shaped transformer for image restoration. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 17662–17672. Cited by: Table 1, §3.3.4, §3.3.4.
- Z. Wang, Q. Xie, T. Li, H. Du, L. Xie, P. Zhu, and M. Bi One-shot voice conversion for style transfer based on speaker adaptation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6792–6796. External Links: Document Cited by: §3.1.2.
- Z. Wang, E. P. Simoncelli, and A. C. Bovik Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, Vol. 2, pp. 1398–1402. Cited by: §3.7.1.
- Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao SIMVLM: Simple visual language model pretraining with weak supervision. In International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.1.1.
- H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, J. Yang, J. Sun, C. Han, and X. Zhang Vary: Scaling up the vision vocabulary for large vision-language models. In European Conference on Computer Vision, Cited by: Table 1, §3.1.1.
- S. Woo, J. Park, J. Lee, and I. S. Kweon CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Cited by: §2.1.
- G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and W. Xinggang 4D gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.5.3.
- H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin Fast-VQA: efficient end-to-end video quality assessment with fragment sampling. In European conference on computer vision, pp. 538–554. Cited by: Table 1, §3.7.1.
- H. Wu, L. Liao, J. Hou, C. Chen, E. Zhang, A. Wang, W. Sun, Q. Yan, and W. Lin Exploring opinion-unaware video quality assessment with semantic affinity criterion. In 2023 IEEE International Conference on Multimedia and Expo (ICME), pp. 366–371. Cited by: §3.7.1.
- H. Wu, L. Liao, A. Wang, C. Chen, J. Hou, W. Sun, Q. Yan, and W. Lin Towards robust text-prompted semantic criterion for in-the-wild video quality assessment. arXiv preprint arXiv:2304.14672. Cited by: §3.7.1, §3.7.1.
- H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20144–20154. Cited by: Table 1, §3.7.1.
- H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, et al. Q-bench: a benchmark for general-purpose foundation models on low-level vision. In The Twelfth International Conference on Learning Representations, Cited by: §3.7.1.
- H. Wu, Z. Zhang, W. Zhang, C. Chen, L. Liao, C. Li, Y. Gao, A. Wang, E. Zhang, W. Sun, et al. Q-Align: teaching LMMs for visual scoring via discrete text-defined levels. In International Conference on Machine Learning, Cited by: §3.7.1.
- J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7623–7633. Cited by: Table 1, §3.1.4.
- T. Wu, Y. Zhang, X. Wang, X. Zhou, G. Zheng, Z. Qi, Y. Shan, and X. Li CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 8469–8477. External Links: Document Cited by: Table 1, §3.1.4.
- T. Wu, S. He, J. Liu, S. Sun, K. Liu, Q. Han, and Y. Tang A brief overview of ChatGPT: the history, status quo and potential future development. IEEE/CAA Journal of Automatica Sinica 10 (5), pp. 1122–1136. External Links: Document Cited by: §1.
- W. Wu, Y. Zhao, H. Chen, Y. Gu, R. Zhao, Y. He, H. Zhou, M. Z. Shou, and C. Shen DatasetDM: Synthesizing data with perception annotations using diffusion models. In Advances in Neural Information Processing Systems, Vol. 36, pp. 54683–54695. Cited by: §3.4.2.
- W. Wu, Y. Zhao, M. Z. Shou, H. Zhou, and C. Shen DiffuMask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1206–1217. Cited by: Table 1, §3.4.1.
- Z. Wu, L. Zhu, Z. Yin, X. Xu, J. Zhu, X. Wei, and X. Yang MAFCD: Multi-level and adaptive conditional diffusion model for anomaly detection. Information Fusion 118, pp. 102965. Note: https://www.sciencedirect.com/science/article/pii/S1566253525000387 External Links: ISSN 1566-2535, Document Cited by: Table 1, §3.4.2.
- J. Wynn and D. Turmukhambetov DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.5.2.
- J. Xiang, K. Tian, and J. Zhang Mimt: masked image modeling transformer for video compression. In International Conference on Learning Representations, Cited by: Table 1, §3.6.2.
- F. Xie, Z. Wang, and C. Ma DiffusionTrack: Point set diffusion model for visual object tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 19113–19124. Cited by: Table 1, §3.4.3.
- Y. Xie, T. Takikawa, S. Saito, O. Litany, S. Yan, N. Khan, F. Tombari, J. Tompkin, V. Sitzmann, and S. Sridhar Neural fields in visual computing and beyond. In Computer Graphics Forum, Vol. 41(2), pp. 641–676. Cited by: §2.4, §3.5.2.
- D. Xu, Y. Jiang, P. Wang, Z. Fan, Y. Wang, and Z. Wang NeuralLift-360: Lifting an in-the-wild 2d photo to a 3D object with 360° views. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 4479–4489. External Links: Document Cited by: Table 1, §3.1.5.
- F. Xu, T. Zhou, T. Nguyen, H. Bao, C. Lin, and J. Du Integrating augmented reality and llm for enhanced cognitive support in critical audio communications. International Journal of Human-Computer Studies 194, pp. 103402. External Links: ISSN 1071-5819, Document Cited by: §3.1.5.
- J. Xu, X. Hu, L. Zhu, Q. Dou, J. Dai, Y. Qiao, and P. Heng Video dehazing via a multi-range temporal alignment network with physical prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18053–18062. Cited by: Table 1, §3.3.4.
- J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello Open-vocabulary panoptic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2955–2966. Cited by: Table 1, §3.4.1.
- M. Xu, J. Chen, H. Wang, S. Liu, G. Li, and Z. Bai C3DVQA: full-reference video quality assessment with 3d convolutional neural network. In IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 4447–4451. Cited by: §3.7.1.
- S. Xu, G. Chen, Y. Guo, J. Yang, C. Li, Z. Zang, Y. Zhang, X. Tong, and B. Guo VASA-1: lifelike audio-driven talking faces generated in real time. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: Table 1, Figure 4, §3.1.4.
- X. Xu, R. Wang, C. Fu, and J. Jia SNR-aware low-light image enhancement. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 17693–17703. External Links: Document Cited by: Table 1, §3.3.1.
- Y. Xu, T. Park, R. Zhang, Y. Zhou, E. Shechtman, F. Liu, J. Huang, and D. Liu VideoGigaGAN: Towards detail-rich video super-resolution. arXiv:2404.12388. Cited by: Table 1, §3.3.3.
- B. Yan, Y. Jiang, J. Wu, D. Wang, P. Luo, Z. Yuan, and H. Lu Universal instance perception as object discovery and retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 15325–15336. Cited by: Table 1, Figure 6, §3.4.
- C. Yang, L. Liang, and Z. Su Real-world denoising via diffusion model. arXiv preprint arXiv:2305.04457. Cited by: Table 1, §3.3.4.
- D. Yang, H. Guo, Y. Wang, R. Huang, X. Li, X. Tan, X. Wu, and H. M. Meng UniAudio 1.5: large language model-driven audio codec is a few-shot audio task learner. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Note: https://openreview.net/forum?id=NGrINZyZKk Cited by: §3.6.3.
- D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, X. Wu, et al. UniAudio: an audio foundation model toward universal audio generation. arXiv preprint arXiv:2310.00704. Cited by: §3.6.3.
- D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu Diffsound: discrete diffusion model for text-to-sound generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (), pp. 1720–1733. External Links: Document Cited by: Table 1, §3.1.2.
- H. Yang, L. Pan, Y. Yang, R. Hartley, and M. Liu LDP: language-driven dual-pixel image defocus deblurring network. arXiv: 2307.09815. External Links: Cited by: Table 1, §3.3.4.
- J. Yang Exploring the productivity of generative ai-powered ad campaigns: a consumer response perspective. Journal of Interactive Advertising 25 (3), pp. 222–239. External Links: Document Cited by: §1.
- J. Yang, M. Gao, Z. Li, S. Gao, F. Wang, and F. Zheng Track anything: segment anything meets videos. arXiv:2304.11968. Cited by: Table 1, §3.4.3.
- L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao Depth anything: unleashing the power of large-scale unlabeled data. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.4.2, §3.5.1.
- L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao Depth anything v2. In Advances in Neural Information Processing Systems, Cited by: Table 1, §3.5.1.
- L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M. Yang Diffusion models: a comprehensive survey of methods and applications. ACM Comput. Surv. 56 (4). External Links: Document Cited by: Figure 1.
- R. Yang and S. Mandt Lossy image compression with conditional diffusion models. Advances in Neural Information Processing Systems 36. Cited by: Table 1, §3.6.1.
- S. Yang, M. Ding, Y. Wu, Z. Li, and J. Zhang Implicit neural representation for cooperative low-light image enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12918–12927. Cited by: Table 1, §3.3.1.
- Y. Yang, L. Pan, L. Liu, and M. Liu K3DN: Disparity-aware kernel estimation for dual-pixel defocus deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13263–13272. Cited by: §3.3.4.
- Y. Yang, F. Sun, L. Weihs, E. VanderBilt, A. Herrasti, W. Han, J. Wu, N. Haber, R. Krishna, L. Liu, C. Callison-Burch, M. Yatskar, A. Kembhavi, and C. Clark Holodeck: Language guided generation of 3D embodied AI environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.1.5.
- Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang A survey on large language model (llm) security and privacy: the good, the bad, and the ugly. High-Confidence Computing 4 (2), pp. 100211. External Links: Document Cited by: §2.2.
- S. Ye, D. Kim, S. Kim, H. Hwang, S. Kim, Y. Jo, J. Thorne, J. Kim, and M. Seo FLASK: fine-grained language model evaluation based on alignment skill sets. In The Twelfth International Conference on Learning Representations, Note: https://openreview.net/forum?id=CYmF38ysDa Cited by: Figure 2, §2.2.
- A. Yi and N. Anantrasirichai A comprehensive study of object tracking in low-light environments. Sensors 24 (14). Note: https://arxiv.org/pdf/2312.16250.pdf Cited by: Table 1, §3.4.3.
- X. Yi, H. Xu, H. Zhang, L. Tang, and J. Ma Diff-retinex: rethinking low-light image enhancement with a generative diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12302–12311. Cited by: Table 1, §3.3.1.
- Y. Yin, D. Xu, C. Tan, P. Liu, Y. Zhao, and Y. Wei CLE Diffusion: Controllable light enhancement diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia, pp. 8145–8156. External Links: Document Cited by: Table 1, §3.3.1.
- G. Youk, J. Oh, and M. Kim FMA-Net: Flow-guided dynamic filtering and iterative feature refinement with multi-attention for joint video super-resolution and deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §3.3.4.
- G. Yu, A. Li, H. Wang, Y. Wang, Y. Ke, and C. Zheng DBT-net: dual-branch federative magnitude and phase estimation with attention-in-attention transformer for monaural speech enhancement. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30 (), pp. 2629–2644. External Links: Document Cited by: Table 1, §3.3.4.
- H. Yu, J. Julin, Z. A. Milacski, K. Niinuma, and L. A. Jeni CoGS: Controllable gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 21624–21633. Cited by: Table 1, §3.5.3.
- W. Yu, L. Po, R. C.C. Cheung, Y. Zhao, Y. Xue, and K. Li Bidirectionally deformable motion modulation for video-based human pose transfer. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 7468–7478. External Links: Document Cited by: Table 1, §3.1.4.
- H. Yue, C. Cao, L. Liao, and J. Yang RViDeformer: Efficient raw video denoising transformer with a larger benchmark dataset. IEEE Transactions on Circuits and Systems for Video Technology (), pp. 1–1. External Links: Document Cited by: Table 1, §3.3.4.
- S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. Yang Restormer: efficient transformer for high-resolution image restoration. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 5718–5729. Cited by: Table 1, §3.3.4.
- N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi Soundstream: an end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30, pp. 495–507. Cited by: §3.6.3.
- F. Zeng, B. Dong, Y. Zhang, T. Wang, X. Zhang, and Y. Wei MOTR: end-to-end multiple-object tracking with transformer. In European Conference on Computer Vision (ECCV), Cited by: Table 1, §3.4.3.
- G. Zhai and X. Min Perceptual image quality assessment: a survey. Science China Information Sciences 63, pp. 1–52. Cited by: §3.7.
- Y. Zhan, Z. Li, M. Niu, Z. Zhong, S. Nobuhara, K. Nishino, and Y. Zheng KFD-NeRF: rethinking dynamic nerf with kalman filter. In Proceedings of the European Conference on Computer Vision (ECCV), Cited by: Table 1, §3.5.2.
- F. Zhang and D. R. Bull A perception-based hybrid model for video quality assessment. IEEE Transactions on Circuits and Systems for Video Technology 26 (6), pp. 1017–1028. Cited by: §3.7.1.
- H. Zhang, C. Jung, D. Zou, and M. Li WCDANN: a lightweight CNN post-processing filter for VVC-based video compression. IEEE Access. Cited by: §3.6.2.
- H. Zhang, Z. Wang, D. Zeng, Z. Wu, and Y. Jiang DiffusionAD: Norm-guided one-step denoising diffusion for anomaly detection. IEEE Transactions on Pattern Analysis and Machine Intelligence (), pp. 1–13. External Links: Document Cited by: Table 1, §3.4.2.
- J. Zhang, J. Huang, S. Jin, and S. Lu Vision-language models for vision tasks: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (), pp. 1–20. External Links: Document Cited by: §3.1.1.
- N. Zhang, F. Nex, G. Vosselman, and N. Kerle Lite-Mono: A lightweight cnn and transformer architecture for self-supervised monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18537–18546. Cited by: Table 1, §3.5.1.
- R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Cited by: §3.7.1.
- R. Zhang, D. Cai, L. Qian, Y. Du, H. Lu, and Y. Zhang DiffusionTracker: Targets denoising based on diffusion model for visual tracking. In Pattern Recognition and Computer Vision, pp. 225–237. Cited by: Table 1, §3.4.3.
- T. Zhang, X. Tian, Y. Zhou, S. Ji, X. Wang, X. Tao, Y. Zhang, P. Wan, Z. Wang, and Y. Wu DVIS++: Improved decoupled framework for universal video segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence. External Links: Document Cited by: Table 1, §3.4.1.
- W. Zhang, K. Ma, G. Zhai, and X. Yang Uncertainty-aware blind image quality assessment in the laboratory and wild. IEEE Transactions on Image Processing 30, pp. 3474–3486. Cited by: §3.7.1.
- X. Zhang and Y. Demiris Visible and infrared image fusion using deep learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (8), pp. 10535–10554. External Links: Document Cited by: §3.3.6.
- X. Zhang, Z. Mao, N. Chimitt, and S. H. Chan Imaging through the atmosphere using turbulence mitigation transformer. IEEE Transactions on Computational Imaging 10 (), pp. 115–128. External Links: Document Cited by: Table 1, §3.3.4.
- Y. Zhang, T. Wang, and X. Zhang MOTRv2: Bootstrapping end-to-end multi-object tracking by pretrained object detectors. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , pp. 22056–22065. External Links: ISSN , Document Cited by: Table 1, §3.4.3.
- Y. Zhang, N. Huang, F. Tang, H. Huang, C. Ma, W. Dong, and C. Xu Inversion-based style transfer with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10146–10156. Cited by: Table 1, §3.3.2.
- Z. Zhang, Z. Hu, W. Deng, C. Fan, T. Lv, and Y. Ding DINet: Deformation inpainting network for realistic face visually dubbing on high resolution video. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 3543–3551. Cited by: §3.3.5.
- Z. Zhang, Y. Zhou, C. Li, B. Zhao, X. Liu, and G. Zhai Quality assessment in the era of large models: a survey. arXiv preprint arXiv:2409.00031. Cited by: §3.7.
- K. Zhao, K. Yuan, M. Sun, M. Li, and X. Wen Quality-aware pre-trained models for blind image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 22302–22313. Cited by: §3.7.1.
- L. Zhao, N. B. Gundavarapu, L. Yuan, H. Zhou, S. Yan, J. J., et al. VideoPrism: A foundational visual encoder for video understanding. In Proceedings of the 41st International Conference on Machine Learning, Cited by: Table 1, §3.4.
- W. X. Zhao, K. Zhou, J. Li, T. Tang, et al. A survey of large language models. arXiv:2303.18223. Cited by: §2.2, §2.2.
- Y. Zhao, W. He, C. Jia, Q. Wang, J. Li, Y. Li, C. Lin, K. Zhang, L. Zhang, and S. Ma A neural-network enhanced video coding framework beyond ECM. In 2024 Data Compression Conference (DCC), pp. 605–605. Cited by: §3.6.2.
- Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and J. Chen DETRs beat yolos on real-time object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §3.4.2.
- Y. Zhao, M. Dasari, and T. Guo CleAR: robust context-guided generative lighting estimation for mobile augmented reality. arXiv preprint arXiv:2411.02179. Cited by: Table 1, §3.1.5.
- Z. Zhao, H. Bai, Y. Zhu, J. Zhang, S. Xu, Y. Zhang, K. Zhang, D. Meng, R. Timofte, and L. Van Gool DDFM: Denoising diffusion model for multi-modality image fusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 8082–8093. Cited by: Table 1, §3.3.6.
- C. Zheng, T. Cham, and J. Cai Pluralistic image completion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1438–1447. Cited by: §3.3.5.
- Q. Zheng, Y. Fan, L. Huang, T. Zhu, J. Liu, Z. Hao, S. Xing, C. Chen, X. Min, A. C. Bovik, et al. Video quality assessment: a comprehensive survey. arXiv preprint arXiv:2412.04508. Cited by: §3.7.
- L. Zhong, Z. Wang, and J. Shang LDB: A large language model debugger via verifying runtime execution step-by-step. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics: Findings., Cited by: §4.1.
- F. Zhou, W. Sheng, Z. Lu, and G. Qiu A database and model for the visual quality assessment of super-resolution videos. IEEE Transactions on Broadcasting 70 (2), pp. 516–532. External Links: Document Cited by: §3.7.1.
- S. Zhou, C. Li, K. C.K. Chan, and C. C. Loy ProPainter: Improving propagation and transformer for video inpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10477–10486. Cited by: Table 1, §3.3.5.
- S. Zhou, C. Li, and C. Change Loy LEDNet: joint low-light enhancement and deblurring in the dark. In Computer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), pp. 573–589. Cited by: §3.3.1.
- X. Zhou, Y. Zheng, and J. Yang Bridging the metrics gap in image style transfer: a comprehensive survey of models and criteria. Neurocomputing 624, pp. 129430. External Links: ISSN 0925-2312, Document Cited by: §3.3.2.
- H. Zhu, X. Sui, B. Chen, X. Liu, Y. Fang, and S. Wang 2AFC prompting of large multimodal models for image quality assessment. IEEE Transactions on Circuits and Systems for Video Technology. Cited by: §3.7.1.
- K. Zhu, C. Li, V. Asari, and D. Saupe No-reference video quality assessment based on artifact measurement and statistical analysis. IEEE Transactions on Circuits and Systems for Video Technology 25 (4), pp. 533–546. Cited by: §3.7.1.
- L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang Vision mamba: efficient visual representation learning with bidirectional state space model. Cited by: §2.1.
- X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai Deformable DETR: Deformable transformers for end-to-end object detection. In International Conference on Learning Representations (ICLR), Cited by: Table 1, §3.4.2.
- Y. Zhu, Y. Yang, and T. Cohen Transformer-based transform coding. In International Conference on Learning Representations, Cited by: Table 1, §3.6.1.
- Y. Zhu, L. Zhang, Z. Rong, T. Hu, S. Liang, and Z. Ge INFP: Audio-driven interactive head generation in dyadic conversations. arXiv:2412.04037. Cited by: Table 1, Table 1, §3.1.4.
- R. Zou, C. Song, and Z. Zhang The devil is in the details: window-based attention for image compression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17492–17501. Cited by: Table 1, §3.6.1.
- X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y. J. Lee Segment everything everywhere all at once. In Advances in Neural Information Processing Systems, Vol. 36, pp. 19769–19782. Cited by: Table 1, §3.4.1.
- Z. Zou and N. Anantrasirichai DeTurb: Atmospheric turbulence mitigation with deformable 3d convolutions and 3d swin transformers. In Proceedings of the Asian Conference on Computer Vision (ACCV), Cited by: Table 1, §3.3.4.
- Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye Object detection in 20 years: a survey. Proceedings of the IEEE 111 (3), pp. 257–276. External Links: Document Cited by: §3.4.2.