LLM Reasoning 相关度: 8/10

Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing

Ning Yang, Chuangxin Cheng, Haijun Zhang
arXiv: 2604.07148v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

COMLLM框架结合GRPO和LACS,利用LLM实现移动边缘计算中任务卸载的长期优化,提升系统性能和泛化能力。

主要贡献

  • 提出COMLLM框架,结合GRPO和LACS机制。
  • 利用多步蒙特卡洛rollout建模服务器队列动态,优化长期决策。
  • 实现零样本拓扑可扩展性,无需针对新拓扑重新训练。

方法论

采用生成式框架COMLLM,结合GRPO和LACS,通过多步rollout模拟服务器队列动态,并将其融入奖励设计中,优化LLM的决策。

原文摘要

Emerging computation-intensive applications impose stringent latency requirements on resource-constrained mobile devices. Mobile Edge Computing (MEC) addresses this challenge through task offloading. However, designing effective policies remains difficult due to dynamic task arrivals, time-varying channels, and the spatio-temporal coupling of server queues. Conventional heuristics lack adaptability, while Deep Reinforcement Learning (DRL) suffers from limited generalization and architectural rigidity, requiring retraining when network topology changes. Although Large Language Models (LLMs) offer semantic reasoning capabilities, standard Supervised Fine-Tuning (SFT) yields myopic policies that greedily minimize immediate latency without accounting for long-term system evolution. To address these limitations, we propose COMLLM, a generative framework that enables foresighted decision-making in MEC systems. COMLLM integrates Group Relative Policy Optimization (GRPO) with a Look-Ahead Collaborative Simulation (LACS) mechanism, which performs multi-step Monte Carlo rollouts while jointly modeling server queue dynamics. By incorporating these rollouts into the reward design, the framework captures the long-term impact of current decisions on future system states. Experimental results demonstrate that COMLLM achieves near-optimal latency and improved load-balancing fairness. Notably, it exhibits zero-shot topological scalability, allowing a model trained on small-scale networks to generalize to larger, unseen topologies without retraining, outperforming SFT, DRL, and heuristic baselines.

标签

LLM Mobile Edge Computing Task Offloading Reinforcement Learning

arXiv 分类

cs.LG