From Kinematics to Dynamics: Learning to Refine Hybrid Plans for Physically Feasible Execution
AI 摘要
使用强化学习优化混合规划器的运动轨迹,使其满足物理约束并可在真实机器人上执行。
主要贡献
- 提出了一种基于强化学习的轨迹优化方法
- 将二阶约束显式地纳入马尔可夫决策过程
- 弥合了规划器生成的初始轨迹与真实机器人动力学之间的差距
方法论
定义包含二阶约束的MDP,使用强化学习优化混合规划器生成的初始轨迹,使其满足物理约束。
原文摘要
In many robotic tasks, agents must traverse a sequence of spatial regions to complete a mission. Such problems are inherently mixed discrete-continuous: a high-level action sequence and a physically feasible continuous trajectory. The resulting trajectory and action sequence must also satisfy problem constraints such as deadlines, time windows, and velocity or acceleration limits. While hybrid temporal planners attempt to address this challenge, they typically model motion using linear (first-order) dynamics, which cannot guarantee that the resulting plan respects the robot's true physical constraints. Consequently, even when the high-level action sequence is fixed, producing a dynamically feasible trajectory becomes a bi-level optimization problem. We address this problem via reinforcement learning in continuous space. We define a Markov Decision Process that explicitly incorporates analytical second-order constraints and use it to refine first-order plans generated by a hybrid planner. Our results show that this approach can reliably recover physical feasibility and effectively bridge the gap between a planner's initial first-order trajectory and the dynamics required for real execution.