Towards Socially Intelligent Robots: Research on Human-Aware Vision Language Navigation

Main Article Content

Jiajun Chen

Keywords

intelligent robot, large language models (LLM), vision-and-language navigation (VLN)

Abstract

As robots are increasingly deployed in human-centric environments like homes and hospitals, their ability to navigate safely while respecting social norms is essential. Vision-and-Language Navigation (VLN) offers a promising interface by allowing robots to follow natural language instructions. However, current VLN research is largely confined to static, unpopulated simulations, where humans are treated as irrelevant or merely as obstacles. This paper aims to address the critical gap between mechanical instruction-following and socially intelligent behaviour. The author reviews the limitations of existing datasets and simulators, arguing that a fundamental shift is necessary. The author posits that integrating external Large Language Models (LLMs) is a hard requirement for achieving human-centric VLN, as the author provides the commonsense reasoning, intent prediction, and adaptive planning that current architectures lack. The study concludes that by leveraging LLMs as a reasoning backbone, VLN agents can transition from rigid path execution to responsive, socially aware navigation. This integration is crucial for bridging the divide between static language commands and the dynamic reality of shared spaces, ultimately enabling robots to become acceptable and effective collaborators in daily life.

Abstract 10 | PDF Downloads 2

References

  • [1] Francis, A., et al. (2020). Principles and guidelines for social human-robot interaction. ACM/IEEE International Conference on Human-Robot Interaction (HRI). https://doi.org/10.1145/3319502.3374806
  • [2] Anderson, P., Wu, Q., Teney, D., et al. (2018). Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://doi.org/10.1109/CVPR.2018.00943
  • [3] Chang, A., Dai, A., Funkhouser, T., et al. (2017). Matterport3D: Learning from RGB-D data in indoor environments. International Conference on 3D Vision (3DV). https://doi.org/10.1109/3DV.2017.00081
  • [4] Savva, M., Kadian, A., Maksymets, O., et al. (2019). Habitat: A platform for embodied AI research. IEEE/CVF International Conference on Computer Vision (ICCV). https://doi.org/10.1109/ICCV.2019.00943
  • [5] Hong, Y., Wu, Q., Qi, Y., et al. (2021). VLN-BERT: A recurrent vision-and-language BERT for navigation. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://doi.org/10.1109/CVPR46437.2021.00987
  • [6] Chen, C., Liu, Z., & Wu, Y. (2021). Environmental topology and graph-based learning for vision-and- language navigation. IEEE International Conference on Robotics and Automation (ICRA). https://doi.org/10.1109/ICRA48506.2021.9560912
  • [7] Szot, A., Clegg, A., Undersander, E., et al. (2021). Habitat 2.0: Training home assistants to rearrange their environment. Advances in Neural Information Processing Systems (NeurIPS). https://doi.org/10.48550/arXiv.2106.14405
  • [8] Chen, S., Liu, M., & Hsu, D. (2023). Socially aware robot navigation in dynamic environments. arXiv preprint arXiv:2301.09876. https://doi.org/10.48550/arXiv.2301.09876
  • [9] Mavrogiannis, C., et al. (2021). Social navigation: A survey. IEEE Transactions on Robotics. https://doi.org/10.1145/3583741
  • [10] Johnson, D., et al. (2021). Gaze-guided vision-and-language navigation. IEEE International Conference on Robotics and Automation (ICRA). https://doi.org/10.1109/ICRA48506.2021.9560913
  • [11] Ahn, M., Brohan, A., Brown, T., et al. (2022). Do as I can, not as I say: Grounding language in robotic affordances. Conference on Robot Learning (CoRL). https://doi.org/10.48550/arXiv.2204.01691
  • [12] Huang, W., Abbeel, P., Pathak, D., & Mordatch, I. (2022). Inner monologue: Embodied reasoning through planning with language models. Conference on Robot Learning (CoRL). https://doi.org/10.48550/arXiv.2207.05608
  • [13] Liang, J., Huang, W., Xia, F., et al. (2023). Code as policies: Language model programs for embodied control. IEEE International Conference on Robotics and Automation (ICRA). https://doi.org/10.1109/ICRA48891.2023.10160591
  • [14] Ku, A., Anderson, P., Patel, R., et al. (2020). Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. Empirical Methods in Natural Language Processing (EMNLP). https://doi.org/10.48550/arXiv.2010.07954
  • [15] Anderson, P., Chang, A., Chaplot, D. S., et al. (2021). On the evaluation of vision-and-language navigation. arXiv preprint arXiv:2106.12578. https://doi.org/10.48550/arXiv.2106.12578
  • [16] Driess, D., Xia, F., Sajjadi, M. S. M., et al. (2023). PaLM-E: An embodied multimodal language model. arXiv preprint arXiv:2303.03378. https://doi.org/10.48550/arXiv.2303.03378
  • [17] Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS). https://doi.org/10.48550/arXiv.2005.14165
  • [18] Radford, A., Kim, J. W., Hallacy, C., et al. (2021). Learning transferable visual models from natural language supervision. International Conference on Machine Learning (ICML). https://doi.org/10.48550/arXiv.2103.00020
  • [19] Li, J., Li, D., Xiong, C., & Hoi, S. (2022). BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. International Conference on Machine Learning (ICML). https://doi.org/10.48550/arXiv.2201.12086
  • [20] Shah, D., Osinski, B., Levine, S., et al. (2023). VoxPoser: Composable 3D value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973. https://doi.org/10.48550/arXiv.2307.05973
  • [21] Szot, A., et al. (2022). HABITAT 3.0: A large-scale platform for training and evaluating embodied agents in realistic 3D environments. arXiv preprint arXiv:2210.11516. https://doi.org/10.48550/arXiv.2210.11516