Multimodal Learning 相关度: 9/10

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

Longgang Zhang, Xiaowei Fu, Fuxiang Huang, Lei Zhang
arXiv: 2604.08140v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

论文提出了用于可解释加密流量分析的多模态推理框架,并构建了新的基准数据集BGTD。

主要贡献

  • 提出了Byte-Grounded Traffic Description (BGTD)基准数据集
  • 提出了端到端流量-语言表示框架mmTraffic
  • 实现了高保真、可读且基于证据的流量解释报告

方法论

构建BGTD数据集,使用感知中心流量编码器和认知中心LLM生成器,通过联合优化实现多模态推理。

原文摘要

Network traffic, as a key media format, is crucial for ensuring security and communications in modern internet infrastructure. While existing methods offer excellent performance, they face two key bottlenecks: (1) They fail to capture multidimensional semantics beyond unimodal sequence patterns. (2) Their black box property, i.e., providing only category labels, lacks an auditable reasoning process. We identify a key factor that existing network traffic datasets are primarily designed for classification and inherently lack rich semantic annotations, failing to generate human-readable evidence report. To address data scarcity, this paper proposes a Byte-Grounded Traffic Description (BGTD) benchmark for the first time, combining raw bytes with structured expert annotations. BGTD provides necessary behavioral features and verifiable chains of evidence for multimodal reasoning towards explainable encrypted traffic interpretation. Built upon BGTD, this paper proposes an end-to-end traffic-language representation framework (mmTraffic), a multimodal reasoning architecture bridging physical traffic encoding and semantic interpretation. In order to alleviate modality interference and generative hallucinations, mmTraffic adopts a jointly-optimized perception-cognition architecture. By incorporating a perception-centered traffic encoder and a cognition-centered LLM generator, mmTraffic achieves refined traffic interpretation with guaranteed category prediction. Extensive experiments demonstrate that mmTraffic autonomously generates high-fidelity, human-readable, and evidence-grounded traffic interpretation reports, while maintaining highly competitive classification accuracy comparing to specialized unimodal model (e.g., NetMamba). The source code is available at https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark

标签

多模态学习 LLM 网络安全 流量分析 可解释性

arXiv 分类

cs.CR cs.AI cs.MM cs.NI