Advanced Sensing, Perception, and Analytics for Manufacturing
Lingbao Kong*, Qiyuan Wang and Xinlan Tang
Future Information Innovative College, Fudan University, Shanghai, China
*E-mail: LKong@fudan.edu.cn
Status
Against the backdrop of the ongoing wave of Industry 5.0, intelligent manufacturing has emerged as a cutting-edge focal point within the realm of industrial manufacturing[1]. In this context, sensing technology, serving as the pivotal bridge linking the physical world to digital signal systems, is undergoing a profound transformation from traditional to intelligent paradigms. In the current era dominated by multi-domain manufacturing, traditional unimodal sensing technologies face significant limitations due to their single information dimension, weak anti-interference capability, and high calibration and maintenance costs. Correspondingly, as shown in Fig. 1, multimodal sensing technologies, which offer rich information, robust redundancy for anti-interference, and low calibration and maintenance expenses, are gradually displacing unimodal sensing technologies across a variety of complex or dynamic scenarios. This transition effectively circumvents the challenges encountered by unimodal sensing in new industrial environments[2], while being better aligned with the urgent demands of modern advanced manufacturing. In comparison to unimodal sensing, multimodal sensing technology has achieved significant breakthroughs primarily in three key aspects. The first lies in the comprehensive enhancement of perceptual capabilities. By leveraging diverse sensing detectors, multimodal sensing facilitates cross-modal information complementarity. For instance, as Fig. 1, in the inspection of surface scratches or coating defects on automotive components, integrating multiple sensing modalities such as industrial cameras, 3D laser scanners, and infrared thermal imagers can elevate the defect detection rate from 90% to 99.5%, while concurrently reducing the false alarm rate by 60%[3]. The second breakthrough is the marked improvement in robustness and anti-interference capabilities. In robotic object grasping scenarios, even when visual occlusion or blind spots occur, robots can dynamically adjust their actions through tactile and force feedback [4]. This adaptability ensures task continuity and accuracy despite environmental perturbations. The third aspect centers on innovations in intelligence and system integration. On one hand, cross-modal semantic alignment is realized through multimodal deep learning, with the incorporation of self-supervised mechanisms to reduce reliance on labeled data. For example, in video data processing, visual, auditory, and motion information are automatically correlated, thereby augmenting the model's generalization capacity[5]. On the other hand, edge computing is employed for real-time processing in system integration, mitigating dependence on cloud-based infrastructure. In intelligent logistics, leveraging AGV (Automated Guided Vehicle) navigation and obstacle avoidance technologies, obstacle response times can be minimized to as low as 100 milliseconds [6].
Figure 1. Manufacturing paradigms centred exclusively on traditional unimodal technology have become
increasingly incompatible with the development of modern advanced manufacturing. Correspondingly, multimodal sensing technologies, by leveraging multidimensional information coordination and complementary data fusion, frequently generate synergistic outcomes that exceed the cumulative performance of isolated modalities (i.e., achieving "1+1>2" effects). These technologies are progressively displacing unimodal approaches in complex industrial scenarios and have emerged as a defining characteristic of contemporary advanced manufacturing systems.
Current and future challenges
Despite the immense potential that multimodal sensing technology has demonstrated in the realm of intelligent manufacturing, its development continues to grapple with a myriad of challenges. At the hardware level, the most prominent hurdle lies in the significant disparities in the physical characteristics of diverse sensors, which greatly complicate the seamless integration and alignment of hardware components. For instance, in smart logistics applications, the high-power consumption of LiDAR (Light Detection and Ranging) sensors stands in stark contrast to the low power requirements of cameras, necessitating the design of intricate circuitry and sophisticated cooling solutions. Additionally, the exorbitant production costs pose another formidable obstacle. In predictive maintenance scenarios, the deployment cost of a single sensor can exceed 2 million units of currency, and substantial resources must be further allocated for AI model training and maintenance to analyze the subsequent data. Moving on to the data and algorithm domain, cross-modal data fusion presents a labyrinth of difficulties. These challenges encompass both semantic alignment issues across different sensors—such as in automated welding processes, where data from arc sensors, high-speed cameras, infrared thermometers, and acoustic emission sensors vary widely in terms of data types and physical significance[7],and the integration of multi- protocol heterogeneous networks. For example, the efficient fusion of 5G networks with existing industrial buses in the context of the Industrial Internet of Things remains an elusive goal[8]. Moreover, the computational complexity and real-time requirements are exceedingly demanding. Take the automatic loading task as an illustration, where the fusion of LiDAR point cloud data (points per frame) with 4K- resolution camera images must be accomplished within a stringent 100-millisecond timeframe [9], placing an enormous strain on computational resources.
Lastly, the integration of human-machine interaction also presents considerable difficulties. The new paradigms of Industry 5.0 advocate for a people-centric approach in industrial manufacturing. However, in current industrial manufacturing workflows, workers are often required to engage in complex programming tasks to adjust robotic operations. Consequently, there is a pressing need to enhance the naturalness and adaptability of human-machine interaction, as well as to improve real-time perception and decision-making capabilities in dynamic environments, all while ensuring a heightened level of safety.
Advances in science and technology to meet challenges
To address the aforementioned challenges, multimodal sensing technology must achieve breakthroughs across multiple fronts. The integration of Digital Twin technology offers a viable remedy for the exorbitant hardware costs. By leveraging multimodal sensor data to drive virtual factory simulations, this technology enables the optimization of production strategies, ultimately achieving cost reduction and efficiency enhancement[10]. The continued advancement of edge intelligence and self-learning systems provides effective solutions to mitigate the complexities of cross-modal data fusion and computational demands. Edge intelligence, through the integration of AI chips, facilitates real-time fusion of multimodal data at the terminal device level. This not only meets stringent real-time requirements but also ensures the effective integration of data [11]. On the other hand, self-learning systems employ reinforcement learning to dynamically optimize the weight allocation among multiple sensors, thereby reducing computational complexity, enhancing data reliability, and minimizing redundant data—ultimately lowering the computational burden[12]. Furthermore, human-robot collaborative monitoring technology epitomizes the people-centric ethos of the new industrial paradigm. By utilizing the Kalman filter for spatial alignment, this technology constructs dynamic safety zones that track workers' hand positions in real time. As a result, it curtails the incidence of human-robot collaborative accidents and boosts production efficiency. The development of these technologies not only compensates for the current shortcomings of multimodal sensing in practical applications but also propels mechanical manufacturing technology towards greater efficiency, intelligence, and harmony[13]. Looking ahead, intelligent manufacturing technology is poised for deeper integration with Digital Twin, edge intelligence, self-learning systems, and human-robot collaborative monitoring. This fusion will drive manufacturing systems towards the aspirational goals of "zero defects," "self-awareness," and a "people-first" approach, marking a significant leap forward in the evolution of manufacturing paradigms.
Concluding remarks
Multimodal sensing technology has propelled the evolution of traditional unimodal sensing approaches towards greater efficiency and intelligence. By integrating a diverse array of information acquisition modalities, enhancing anti-interference capabilities, and pioneering intelligent, integrated systems, this technology has effectively shattered the robustness barriers in manufacturing environments. It has not only bolstered the standardization of manufacturing processes and the precision of defect detection but also furnished manufacturing systems with high-fidelity, highly reliable data foundations. Consequently, multimodal sensing technology has emerged as the cornerstone of intelligent perception within the Intelligent Manufacturing ecosystem. Concurrently, the infusion of Digital Twin technology, edge integration architectures, and self-learning systems has further catalyzed the intelligent and miniaturized trajectory of multimodal sensing technology, laying indispensable infrastructure groundwork for contemporary industry. Beyond these technological advancements, the profound integration of human-centric principles stands as a pivotal milestone in the development of multimodal sensing technology. By anchoring human-centricity at the heart of intelligent manufacturing, this paradigm shift fortifies the foundation for IM to align seamlessly with the prevailing
ethos of the times, ensuring that technological progress remains intrinsically linked to human needs and aspirations.
Acknowledgements
The authors would like to express their sincere thanks for the support from National Natural Science Foundation of China (52375414), and Shanghai Science & Technology Committee Innovation Grant (23ZR1404200).
References
[1] AlMahasneh R, Hollósi G, Ficzere D, Bancsics M, Lukovszki C, Varga P. Uncovering Common AI Challenges Across Industrial Domains in the Transition to Industry 5.0[C]//2024 20th International Conference on Network and Service Management (CNSM). IEEE, 2024: 1-7. [2] Liang P P, Zadeh A, Morency L P. Foundations & trends in multimodal machine learning: Principles, challenges, and open questions[J]. ACM Computing Surveys, 2024, 56(10): 1-42. [3] Guclu E, Akin E. Enhanced defect detection on steel surfaces using integrated residual refinement module with synthetic data augmentation[J]. Measurement, 2025, 250: 117136. [4] Li Y, Zheng L, Wang Y, Dong E, Zhang S. Impedance Learning-based Adaptive Force Tracking for Robot on Unknown Terrains[J]. IEEE Transactions on Robotics, 2025. [5] Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S,sastry G, Askell A, Mishkin P, Clark J,Krueger G, Sutskever I. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PmLR, 2021: 8748-8763. [6] Chang Y H, Wu F C, Lin H W. Design and implementation of esp32-based edge computing for object detection[J]. Sensors, 2025, 25(6): 1656. [7] Deng F, Huang Y, Lu S, Chen Y, Feng H, Zhang J, Yang Y, Hu J, Lam TL, Xia F. A multi-sensor data fusion system for laser welding process monitoring[J]. IEEE Access, 2020, 8: 147349-147357. [8] Harmatos J, Maliosz M. Architecture integration of 5G networks and time-sensitive networking with edge computing for smart manufacturing[J]. Electronics, 2021, 10(24): 3085. [9] Sochaniwsky A. A LIGHTWEIGHT CAMERA-LIDAR FUSION FRAMEWORK FOR TRAFFIC MONITORING APPLICATIONS[D].,
[10] Durana P, Krastev V, Buckner K. Digital twin modeling, multi-sensor fusion technology, and data mining algorithms in cloud and edge computing-based Smart city environments[J]. Geopolitics, History, and International Relations, 2022, 14(1): 91-106. [11] Wang M, Zhang X, Chen S, Li X, Zhang Y. A bidirectional separated distillation-based cross-modal interactive fusion network for skeleton-based action recognition[J]. IEEE Sensors Journal, 2024. [12] Guo J, Liu Q, Chen E. A deep reinforcement learning method for multimodal data fusion in action recognition[J]. IEEE Signal Processing Letters, 2021, 29: 120-124. [13] Dani A P, Salehi I, Rotithor G, Trombetta D, Ravichandar H. Human-in-the-loop robot control for human-robot collaboration: Human intention estimation and safe trajectory tracking control for collaborative tasks[J]. IEEE Control Systems Magazine, 2020, 40(6): 29-56.