LLM Reasoning 相关度: 7/10

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models

Jinane Bazzi, Mariam Rakka, Fadi Kurdahi, Mohammed E. Fouda, Ahmed Eltawil
arXiv: 2604.11512v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

EdgeCIM通过软硬件协同设计,利用CIM加速边缘设备上小型语言模型的推理,提高效率。

主要贡献

  • 提出了EdgeCIM软硬件协同设计框架
  • 设计了基于CIM的宏和tile-based mapping策略
  • 在多种SLM上验证了EdgeCIM的性能和能效优势

方法论

通过CIM宏和tile-based mapping策略,优化pipeline,最大化并行性,缓解DRAM带宽瓶颈,并使用仿真器进行设计空间探索。

原文摘要

The growing demand for deploying Small Language Models (SLMs) on edge devices, including laptops, smartphones, and embedded platforms, has exposed fundamental inefficiencies in existing accelerators. While GPUs handle prefill workloads efficiently, the autoregressive decoding phase is dominated by GEMV operations that are inherently memory-bound, resulting in poor utilization and prohibitive energy costs at the edge. In this work, we present EdgeCIM, a hardware-software co-design framework that rethinks accelerator design for end-to-end decoder-only inference. At its core is a CIM macro, implemented in 65nm, coupled with a tile-based mapping strategy that balances pipeline stages, maximizing parallelism while alleviating DRAM bandwidth bottlenecks. Our simulator enables design space exploration of SLMs up to 4B parameters, identifying Pareto-optimal configurations in terms of latency and energy. Compared to an NVIDIA Orin Nano, EdgeCIM achieves up to 7.3x higher throughput and 49.59x better energy efficiency on LLaMA3.2-1B, and delivers 9.95x higher throughput than Qualcomm SA8255P on LLaMA3.2-3B. Extensive benchmarks on TinyLLaMA-1.1B, LLaMA3.2 (1B, 3B), Phi-3.5-mini-3.8B, Qwen2.5 (0.5B, 1.5B, 3B), SmolLM2-1.7B, SmolLM3-3B, and Qwen3 (0.6B, 1.7B, 4B) reveal that our accelerator, under INT4 precision, achieves on average 336.42 tokens/s and 173.02 tokens/J. These results establish EdgeCIM as a compelling solution towards real-time, energy-efficient edge-scale SLM inference.

标签

CIM 边缘计算 小型语言模型 硬件加速 低功耗

arXiv 分类

cs.AR cs.AI