面向人机交互的导纳与强化学习融合控制策略

Admittance and Reinforcement Learning Integrated Control Strategy for Human-Robot Interaction

  • 摘要: 针对人机交互任务中交互柔顺性不足、碰撞风险高以及跟踪性能差等问题,提出了一种基于导纳控制的强化学习策略。首先,利用软饱和函数和期望导纳模型生成光滑可微的参考轨迹。其次,引入基于AC(actor-critic)结构的强化学习方法,以应对系统动态不确定性。与现有研究不同,本文创新性地构建了包含时变非对称障碍李雅普诺夫函数的控制器,既能确保末端执行器实现精确轨迹跟踪又能严格满足位置约束,保障系统在复杂交互场景中的安全性和可靠性。最后,通过李雅普诺夫稳定性理论,证明了闭环系统的半全局一致最终有界性,基于Baxter机器人实验平台进行了一系列的实验,并与自适应阻抗控制、模糊自适应阻抗控制和传统阻抗控制进行了对比分析。对比结果表明所提出的控制方法在跟踪精度、柔顺性和防碰撞效果方面均优于对比方法,充分验证了所提出方法的可行性。

     

    Abstract: In response to the issues of insufficient interaction compliance, high collision risk, and poor tracking performance in human-robot interaction tasks, a reinforcement learning strategy based on admittance control is proposed. Firstly, a smooth and differentiable reference trajectory is generated using a soft saturation function and an expected admittance model. Secondly, a reinforcement learning method based on the AC (actor-critic) structure is introduced to deal with the system dynamic uncertainties. Unlike existing studies, a controller incorporating a time-varying asymmetric barrier Lyapunov function is constructed. This controller ensures accurate trajectory tracking of the end-effector while strictly satisfying position constraints, thus guaranteeing the system safety and reliability in complex interaction scenarios. Finally, the semi-global uniform ultimate boundedness of the closed-loop system is demonstrated through Lyapunov stability theory. A series of experiments are conducted on the Baxter robot experimental platform, and the proposed control method is compared with adaptive impedance control, fuzzy adaptive impedance control and traditional impedance control. The comparison results show that the proposed control method outperforms the other methods in terms of tracking accuracy, compliance and collision avoidance, thereby fully validating the feasibility of the proposed method.

     

/

返回文章
返回