Multimodal Learning 相关度: 8/10

Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models

Jing Gu, Niccolò Cavagnero, Gijs Dubbelman
arXiv: 2604.08266v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出Orion-Lite,通过知识蒸馏将LLM推理能力赋予高效的视觉驾驶模型,并超越VLA教师模型。

主要贡献

  • 提出Orion-Lite,高效的视觉驾驶模型
  • 通过蒸馏LLM知识提升驾驶模型性能
  • 在复杂场景下进行闭环评估并取得SOTA结果

方法论

采用latent feature distillation和ground-truth trajectory supervision相结合的方法,将VLA教师模型的知识迁移到视觉模型Orion-Lite。

原文摘要

Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems to handle rare and complex scenarios. While integrating LLMs into Vision-Language-Action (VLA) models has yielded state-of-the-art performance, their massive parameter counts pose severe challenges for latency-sensitive and energy-efficient deployment. Distilling LLM knowledge into a compact driving model offers a compelling solution to retain these reasoning capabilities while maintaining a manageable computational footprint. Although previous works have demonstrated the efficacy of distillation, these efforts have primarily focused on relatively simple scenarios and open-loop evaluations. Therefore, in this work, we investigate LLM distillation in more complex, interactive scenarios under closed-loop evaluation. We demonstrate that through a combination of latent feature distillation and ground-truth trajectory supervision, an efficient vision-only student model \textbf{Orion-Lite} can even surpass the performance of its massive VLA teacher, ORION. Setting a new state-of-the-art on the rigorous Bench2Drive benchmark, with a Driving Score of 80.6. Ultimately, this reveals that vision-only architectures still possess significant, untapped potential for high-performance reactive planning.

标签

知识蒸馏 自动驾驶 视觉模型 LLM

arXiv 分类

cs.CV