Multimodal Learning 相关度: 9/10

From Multimodal Signals to Adaptive XR Experiences for De-escalation Training

Birgit Nierula, Karam Tomotaki-Dawoud, Daniel Johannes Meyer, Iryna Ignatieva, Mina Mottahedin, Thomas Koch, Sebastian Bosse
arXiv: 2604.11570v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

构建多模态实时交流分析系统,用于自适应VR降级训练,并应用于执法领域。

主要贡献

  • 设计并实现多模态实时交流分析系统
  • 提出将低层信号与交互结构关联的解释层
  • 评估了XR降级训练中自动线索提取的可行性和局限性

方法论

整合语音、姿势、情感、脑电和生理信号,通过Lab Streaming Layer同步,结合领域知识进行低层信号到交互结构的映射。

原文摘要

We present the early-stage design and implementation of a multimodal, real-time communication analysis system intended as a foundational interaction layer for adaptive VR training. The system integrates five parallel processing streams: (1) verbal and prosodic speech analysis, (2) skeletal gesture recognition from multi-view RGB cameras, (3) multimodal affective analysis combining lower-face video with upper-face facial EMG, (4) EEG-based mental state decoding, and (5) physiological arousal estimation from skin conductance, heart activity, and proxemic behavior. All signals are synchronized via Lab Streaming Layer to enable temporally aligned, continuous assessments of users' conscious and unconscious communication cues. Building on concepts from social semiotics and symbolic interactionism, we introduce an interpretation layer that links low-level signal representations to interactional constructs such as escalation and de-escalation. This layer is informed by domain knowledge from police instructors and lay participants, grounding system responses in realistic conflict scenarios. We demonstrate the feasibility and limitations of automated cue extraction in an XR-based de-escalation training project for law enforcement, reporting preliminary results for gesture recognition, emotion recognition under HMD occlusion, verbal assessment, mental state decoding, and physiological arousal. Our findings highlight the value of multi-view sensing and multimodal fusion for overcoming occlusion and viewpoint challenges, while underscoring that fusion and feedback must be treated as design problems rather than purely technical ones. The work contributes design resources and empirical insights for shaping human-AI-powered XR training in complex interpersonal settings.

标签

Multimodal VR Training De-escalation Law Enforcement Affective Computing

arXiv 分类

cs.HC cs.MM