AI Agents 相关度: 7/10

Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production

Jintao Xue, Xiao Li, Nianmin Zhang
arXiv: 2604.12667v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

提出PF-CD3Q算法,结合粒子滤波和安全强化学习,解决人机协作中疲劳预测的任务规划问题。

主要贡献

  • 提出基于粒子滤波的疲劳参数在线估计方法
  • 将疲劳预测融入强化学习决策过程,约束动作空间
  • 提出PF-CD3Q算法,实现实时疲劳预测的任务规划

方法论

使用粒子滤波在线估计疲劳参数,并将其集成到约束对偶深度Q学习中,实现安全强化学习。

原文摘要

Human-robot collaborative manufacturing, a core aspect of Industry 5.0, emphasizes ergonomics to enhance worker well-being. This paper addresses the dynamic human-robot task planning and allocation (HRTPA) problem, which involves determining when to perform tasks and who should execute them to maximize efficiency while ensuring workers' physical fatigue remains within safe limits. The inclusion of fatigue constraints, combined with production dynamics, significantly increases the complexity of the HRTPA problem. Traditional fatigue-recovery models in HRTPA often rely on static, predefined hyperparameters. However, in practice, human fatigue sensitivity varies daily due to factors such as changed work conditions and insufficient sleep. To better capture this uncertainty, we treat fatigue-related parameters as inaccurate and estimate them online based on observed fatigue progression during production. To address these challenges, we propose PF-CD3Q, a safe reinforcement learning (safe RL) approach that integrates the particle filter with constrained dueling double deep Q-learning for real-time fatigue-predictive HRTPA. Specifically, we first develop PF-based estimators to track human fatigue and update fatigue model parameters in real-time. These estimators are then integrated into CD3Q by making task-level fatigue predictions during decision-making and excluding tasks that exceed fatigue limits, thereby constraining the action space and formulating the problem as a constrained Markov decision process (CMDP).

标签

安全强化学习 人机协作 疲劳预测 任务规划 粒子滤波

arXiv 分类

cs.AI