ICRA 2026 于2026年6月1日至5日在奥地利维也纳举行。作为全球机器人领域最具影响力的顶级会议之一,大会汇聚来自全球的顶尖学者与产业领袖,分享行业洞察,展示前沿成果,推动机器人技术持续突破与落地应用。

上海交通大学医疗机器人研究院多篇投稿论文参会,论文《One-Shot Autofocus Via User-Adaptive Gaze Control for Robot-Assisted Microsurgery》入选ICRA最佳医疗机器人论文(Best Paper Award in Medical Robotics)提名奖,并受邀在大会进行口头报告。研究内容涵盖手术机器人视觉、显微外科自动化、实时SLAM及微型连续体机器人控制等前沿方向,彰显了其在机器人领域持续提升的科研实力与国际影响力。
Closed-Loop Cross-Scale Motion of Decoupled Light- and Tendon- Driven Miniature Continuum Robots
Cheng Zhou, Xiaotong Qin, Haoyang Yu, Jingyuan Xia, Zheng Xu, Zecai Lin, Anzhu Gao
Small-scale robots are rapidly advancing in diverse fields such as industry and medicine. To be effective, they must be capable of accessing narrow, tortuous, or otherwise hard-to-reach environments and performing precise manipulation. This paper presents a vision-based closed-loop motion control scheme for a developed fiber-driven continuum robot for cross-scale motion. Function-multiplexed optical fibers are employed to achieve macro motion through fiber actuation and micro motion through light transmission within the fibers. An external eye-to-hand camera system observes a fiducial tag to estimate its 3D pose relative to the camera frame. The coordinate transformation between the tag and the end-effector is calibrated, along with the mapping between input laser power and light-induced joint contractions. A two-stage image-based visual servoing strategy is then implemented to guide the tag toward the target image position, thereby realizing closed-loop hybrid macro–micro motion through the developed kinematics and visual feedback. Point-tracking experiments demonstrate that the small-scale continuum robot, with an outer diameter of approximately 1.2 mm, can achieve precise cross-scale motion across workspaces ranging from tens of microns to the millimeter scale under the proposed control scheme. This work highlights the potential of hybrid macro–micro motion with visual servoing for deep access and high-precision operation in endoluminal interventions.

GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State
Guole Shen, Tianchen Deng,Yanbo Wang, Yongtao Chen, Yilin Shen, Jiuming Liu, Jingchuan Wang
DUSt3R-based end-to-end scene reconstruction has recently shown promising results in dense visual SLAM. However, most existing methods only use image pairs to estimate pointmaps, overlooking spatial memory and global consistency. To this end, we introduce GRS-SLAM3R, an end-to-end SLAM framework for dense scene reconstruction and pose estimation from RGB images without any prior knowledge of the scene or camera parameters. Unlike existing DUSt3R-based frameworks, which operate on all image pairs and predict per-pair point maps in local coordinate frames, our method supports sequentialized input and incrementally estimates metric-scale pointclouds in the global coordinate. In order to improve consistent spatial correlation, we use a latent state for spatial memory and design a transformer-based gated update module to reset and update the spatial memory that continuously aggregates and tracks relevant 3D information across frames. Furthermore, we partition the scene into submaps, apply local alignment within each submap, and register all submaps into a common world frame using relative constraints, producing a globally consistent map. Experiments on various datasets show that our framework achieves superior reconstruction accuracy while maintaining real-time performance.

Cross-view Exocentric and Egocentric Fusion for Robust Microsurgical Anastomosis Understanding
Yuxuan Liu, Yuyang Zhuge, Xinyao Zhou, Yating Luo, Yunfei Luan, Yao Guo, Guang-Zhong Yang
Microsurgical anastomosis has become increasingly prevalent in surgical autonomy, requiring accurate and stable control of suturing needles and threads while enhancing the efficiency and safety of microsurgical operations. However, current systems predominantly employ top-down-view microscopes for intraoperative imaging, which are constrained by limited field-of-view and significant occlusion caused by instrument-tissue interactions. To address these challenges, we develop a dual-view vision system for microsurgical anastomosis, integrating both conventional top-down-view microscopes and eye-in-hand cameras mounted on surgical instrument tips. Our approach involves cross-view feature fusion through different schemes to improve microsurgical scene understanding, including surgical action recognition, gripper-object interaction prediction, and instrument pose estimation. Extensive anastomosis datasets are collected on our robotic platform and several experiments are conducted for detailed evaluation of the system performance. Quantitative and qualitative results demonstrate that our dual-view microsurgical system significantly outperforms single-view microscopes in terms of robust visual perception, and cross-view feature fusion improves both the accuracy and precision of anastomosis scene understanding.

One-Shot Autofocus Via User-Adaptive Gaze Control for Robot-Assisted Microsurgery (Finalists of Best Paper Award in Medical Robotics)
Yunfei Luan, Yuxuan Liu, Yuyang Zhuge, Yating Luo, Yao Guo, Guang-Zhong Yang
Robot-assisted microsurgery (RAMS) is rapidly advancing with increasing levels of automation. Given the inherently shallow depth-of-field characteristics of surgical microscopes, integrating autofocus capabilities into RAMS has emerged as an urgent trend. Among existing solutions, gaze-induced autofocus has gained prominence due to its natural alignment with the surgeon's visual attention. However, gaze autofocus often relies on complex and non-intuitive triggering mechanisms, making it difficult to adapt for diverse users. Additionally, although the hill-climbing strategy is commonly employed to find the optimal focus plane, this process is inefficient for RMAS due to its slow convergence and inability to accommodate dynamic surgical scenarios. To address these limitations, we propose a novel gaze-controlled autofocus system featuring user-adaptive triggering and one-shot focusing. When a region is defocused and under the surgeon's gaze, our system rapidly achieves optimal focus with a single-step lens movement. Surgeons can easily adjust trigger sensitivity using a slider. Experiments validate the accuracy of our defocus estimation and triggering prediction algorithms. A user study demonstrates that the proposed system offers superior user-friendliness and operational efficiency compared to conventional systems.

