Multimodal Learning 相关度: 8/10

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization

Yanis Labrak, David Grünert, Séverin Baroudi, Jiyun Chun, Pawel Cyrta, Sergio Burdisso, Ahmed Hassoon, David Liu, Adam Rothschild, Reed Van Deusen, Petr Motlicek, Andrew Perrault, Ricard Marxer, Thomas Schaaf
arXiv: 2604.06138v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出一个合成医生-病人对话数据集用于长文本音频摘要,解决数据稀缺和评估难题。

主要贡献

  • 构建合成数据生成pipeline
  • 生成大规模医生-病人对话数据集
  • 评估现有开放权重模型在长音频摘要上的性能

方法论

通过角色驱动对话生成、多说话人音频合成和LLM生成参考SOAP笔记,构建合成数据。

原文摘要

Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models.

标签

长文本音频摘要 合成数据 医疗对话 SOAP笔记 开放权重模型

arXiv 分类

cs.SD cs.AI