Task-Aware 6-DoF Grasp Pose Estimation through Cross-Modal Gated Fusion of Foundation-Model Semantic Embeddings

Main Article Content

Mulian Lai

Keywords

6-DoF grasp estimation, task-aware grasping, cross-modal fusion, foundation model, robotic manipulation

Abstract

Robotic grasping in cluttered scenes is dominated by methods that optimise either stability or a predefined task category, leaving a gap between “can this grasp hold the object” and “can this grasp complete the requested downstream task.” We address this gap with Task-Aware 6-DoF, a three-module detector that routes a natural- language task description through a frozen text encoder, fuses the resulting semantic vector with PointNet ++ point features through a learnable gated residual, and predicts task completion through a binary classifier trained on disjoint seeds. On a four-split procedurally generated benchmark with three random seeds and 25 trials per cell, our method reaches a task completion rate of 0.60 against seven reproduced baselines spanning two naive controllers and five recent published methods, beating six of seven decisively on the primary metric and tying the remaining LLM-only baseline within the simulator noise ban d. An honest ablation isolates the task-completion head as the dominant contribution while showing the gated attention module to be redundant. Calibration against three anchor points from prior work yields a mean absolute percentage error of 7.2 percent, comfortably below the 15 percent bound for analytic surrogates.

Abstract 9 | PDF Downloads 3

References

  • [1] Li, Yiming, Tao Kong, Ruihang Chu, Yifeng Li, Peng Wang, and Lei Li. 2021. “Simultaneous Semantic and Collision Learning for 6 -DoF Grasp Pose Estimation.” Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).
  • [2] Fang, Hao-Shu, Chenxi Wang, Minghao Gou, and Cewu Lu. 2020. “GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  • [3] Lenz, Ian, Honglak Lee, and Ashutosh Saxena. 2015. “Deep Learning for Detecting Robotic Grasps.” The International Journal of Robotics Research 34 (4–5): 705–24.
  • [4] Wang, Junfeng, Yifan Su, et al. 2024. “ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation.” arXiv Preprint.
  • [5] Yang, Yang, Yuanhao Zhang, Yan Liu, Min Tan, et al. 2024. “Attribute-Based Robotic Grasping with Data-Efficient Adaptation.” IEEE Transactions on Robotics.
  • [6] Tremblay, Jonathan, Thang To, Balakumar Sundaralingam, Yu Xiang, Dieter Fox, and Stan Birchfield. 2018. “Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects.” Conference on Robot Learning (CoRL).
  • [7] Tang, Sicong, Hamidreza Kasaei, et al. 2024. “CenterArt: Joint Shape Reconstruction and 6 -DoF Grasp Estimation of Articulated Objects.” arXiv Preprint.
  • [8] Wei, Xibai, Yang Zhang, Xinyi Geng, et al. 2024. “Learning to Generate 6 -DoF Grasp Poses with Reachability Awareness.” arXiv Preprint.
  • [9] Wang, An-Lan, Nuo Chen, Kun-Yu Lin, Yuan-Ming Li, and Wei-Shi Zheng. 2024. “Task-Oriented 6- DoF Grasp Pose Detection in Clutters.” Proceedings of the IEEE International Conference on Robotics and Automation (ICRA).
  • [10] Zhao, Jialiang, Daniel Troniak, and Oliver Kroemer. 2018. “Towards Robotic Assembly by Predicting Robust, Precise and Task-Oriented Grasps.” Conference on Robot Learning (CoRL).
  • [11] Tian, Dongying, Xiangbo Lin, and Yi Sun. 2024. “Adaptive Motion Planning for Multi-Fingered Functional Grasp via Force Feedback.” arXiv Preprint.
  • [12] Liu, Chuer, Yixuan Pan, David Held, et al. 2024. “TAX-Pose: Task-Specific Cross-Pose Estimation for Robot Manipulation.” Conference on Robot Learning (CoRL).
  • [13] Shaw-Cortez, Wenceslao, Denny Oetomo, Chris Manzie, and Peter Choong. 2018. “Robust Object Manipulation for Tactile-Based Blind Grasping.” arXiv Preprint arXiv:1709.02924.
  • [14] Zeng, Chao, Shuang Li, Yiming Jiang, et al. 2021. “Learning Compliant Grasping and Manipulation by Teleoperation with Adaptive Force Control.” arXiv Preprint arXiv:2107.08996.
  • [15] Wang, Sijie, Jianhua Zhou, et al. 2024. “Self-Supervised 6-DoF Robot Grasping by Demonstration via Augmented Reality Teleoperation System.” arXiv Preprint.
  • [16] Du, Guoguang, Kai Wang, Shiguo Lian, and Kaiyong Zhao. 2021. “Vision-Based Robotic Grasping from Object Localization, Object Pose Estimation to Grasp Estimation for Parallel Grippers: A Review.” Artificial Intelligence Review.
  • [17] Toresano, Linus et al. 2022. “A ROS2 -Based Communication Architecture for Control in Collaborative and Intelligent Automation Systems.” Proceedings of the IEEE International Conference on Industrial Informatics (INDIN).