Agent Tuning & Optimization 相关度: 6/10

A Comparative Study of CNN Optimization Methods for Edge AI: Exploring the Role of Early Exits

Nekane Fernandez, Ivan Valdes, Steven Van Vaerenbergh, Idoia de la Iglesia, Julen Arratibel
arXiv: 2604.14789v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

比较静态压缩和动态早退方法在边缘设备上的CNN优化,并探索其组合的优势。

主要贡献

  • 比较静态压缩和动态早退方法在边缘设备上的性能
  • 提出两种方法结合使用能进一步降低延迟和内存占用
  • 在真实边缘设备上进行ONNX推理评估

方法论

在边缘设备上,使用ONNX框架评估静态压缩(剪枝、量化)和动态早退方法,并分析其性能。

原文摘要

Deploying deep neural networks on edge devices requires balancing accuracy, latency, and resource constraints under realistic execution conditions. To fit models within these constraints, two broad strategies have emerged: static compression techniques such as pruning and quantization, which permanently reduce model size, and dynamic approaches such as early-exit mechanisms, which adapt computational cost at runtime. While both families are widely studied in isolation, they are rarely compared under identical conditions on physical hardware. This paper presents a unified deployment-oriented comparison of static compression and dynamic early-exit mechanisms, evaluated on real edge devices using ONNX based inference pipelines. Our results show that static and dynamic techniques offer fundamentally different trade-offs for edge deployment. While pruning and quantization deliver consistent memory footprint reduction, early-exit mechanisms enable input-adaptive computation savings that static methods cannot match. Their combination proves highly effective, simultaneously reducing inference latency and memory usage with minimal accuracy loss, expanding what is achievable at the edge.

标签

Edge AI CNN Optimization Static Compression Early Exit ONNX

arXiv 分类

cs.AI