基于强化学习的移动机器人分层有序环境探索方法

A Reinforcement Learning-based Hierarchical and Ordered Exploration Method for Mobile Robots

  • 摘要: 针对移动机器人自主探索方法存在的重复探索、环境探索不完整等问题,提出一种基于强化学习的移动机器人分层有序环境探索方法。该方法构建局部前向探索与全局逆向回溯的分层框架,将探索过程解耦为两个阶段:局部探索阶段采用一种软演员-评论家算法,以局部环境探索地图为输入,通过神经网络提取环境特征,结合探索奖励与安全性奖励,实现高效的局部深度优先探索。全局回溯阶段基于多叉树轨迹记录结构,通过势场法判定回溯点并结合图像掩模方法来优化路径,实现全局未探索区域的有序回溯与路径优化。实验结果表明,与前沿点法和下一最优视点法相比,所提方法具有更高的环境探索覆盖率与探索效率,同时也通过消融实验验证了局部探索方法对探索效率提升与全局回溯机制对探索覆盖率提升的必要性。最后实机实验验证了所提方法相较于现有环境探索方法具有更快的探索速度与更高的探索效率。

     

    Abstract: In view of the problems such as repeated exploration and incomplete environmental exploration existing in the autonomous exploration method of mobile robots, a hierarchical and ordered environmental exploration method for mobile robots based on reinforcement learning is proposed. This method constructs a hierarchical framework of local forward exploration and global reverse backtracking, and decouples the exploration process into two stages. For the local exploration stage, a soft actor-critic algorithm is proposed. Taking the local environmental exploration map as input, it extracts environmental features through neural networks, and combines exploration rewards and safety rewards to achieve efficient local depth-first exploration. In the global backtracking stage, the backtracking points are determined by the potential field method and the path is optimized by combining the image mask, based on the multi-tree trajectory recording structure. And thus, the ordered backtracking and path optimization in global unexplored areas are realized. Experimental results show that compared with the front-point method and the next-best-view method, the proposed method has higher environmental exploration coverage rate and exploration efficiency. Meanwhile, the ablation experiment also verifies the necessities of the local exploration module for improving the exploration efficiency and the global backtracking mechanism for improving the exploration coverage rate. Finally, the real-machine experiment verifies that the proposed method has a faster exploration speed and a higher exploration efficiency compared with existing environmental exploration methods.

     

/

返回文章
返回