Abstract:
This survey systematically reviews key advancements in intelligent perception, navigation, and manipulation empowered by large language models and multimodal large models. At the perception level, deep understanding of environmental semantics and geometric attributes is enhanced through multimodal fusion and language-spatial joint reasoning. For navigation, techniques such as chain-of-thought task decomposition and commonsense reasoning enable the parsing of ambiguous instructions and autonomous exploration in unknown environments. In manipulation, the dexterity and adaptability for complex interactive tasks are improved by vision-language-action models coupled with physical commonsense. Research indicates that the introduction of large models drives a paradigm shift in robotics from “perception-driven” to “cognition-driven”, significantly enhancing the system capabilities in contextual reasoning and autonomous decision-making. However, core challenges remain, including cross-modal alignment accuracy, real-time performance, safety, reliability, and Sim2Real generalization. A systematic technical reference and development roadmap are provided for building general-purpose, cognition-enhanced intelligent robotic systems.