Agent Tuning & Optimization 相关度: 6/10

CUTEv2: Unified and Configurable Matrix Extension for Diverse CPU Architectures with Minimal Design Overhead

Jinpeng Ye, Chongxi Wang, Wenqing Li, Bin Yuan, Shiyi Wang, Fenglu Zhang, Junyu Yue, Jianan Xie, Yunhao Ye, Haoyu Deng, Yingkun Zhou, Xin Cheng, Fuxin Zhang, Jian Wang
arXiv: 2604.11615v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出一种统一可配置的CPU矩阵扩展架构,降低硬件开销并提升AI模型性能。

主要贡献

  • 解耦矩阵单元与CPU流水线,降低集成开销
  • 可配置矩阵单元支持混合精度和适应不同带宽需求
  • 异步矩阵乘法抽象简化开发,支持统一软件栈

方法论

通过解耦设计和异步抽象,在开源CPU RTL平台上集成和评估,并在实际AI模型上验证性能。

原文摘要

Matrix extensions have emerged as an essential feature in modern CPUs to address the surging demands of AI workloads. However, existing designs often incur substantial hardware and software design overhead. Tight coupling with the CPU pipeline complicates integration across diverse CPUs, while fine-grained synchronous instructions hinder the development of high-performance kernels. This paper proposes a unified and configurable CPU matrix extension architecture. By decoupling matrix units from the CPU pipeline, the design enables low-overhead integration while maintaining close coordination with existing compute and memory resources. The configurable matrix unit supports mixed-precision operations and adapts to diverse compute demands and memory bandwidth constraints. An asynchronous matrix multiplication abstraction with flexible granularity conceals hardware details, simplifies matrix-vector overlap, and supports a unified software stack. The architecture is integrated into four open-source CPU RTL platforms and evaluated on representative AI models. Matrix unit utilization under GEMM workloads exceeds 90% across all platforms. When configured with compute throughput and memory bandwidth comparable to Intel AMX, our design achieves speedups of 1.57x, 1.57x, and 2.31x on ResNet, BERT, and Llama3, with over 30% of the gains attributed to overlapped matrix-vector execution. A 4 TOPS@2GHz matrix unit occupies only 0.53 mm\textsuperscript{2} in 14nm CMOS. These results demonstrate strong cross-platform adaptability and effective hardware-software co-optimization, offering a practical matrix extension for the open-source community.

标签

CPU 矩阵扩展 AI加速 硬件架构 异构计算

arXiv 分类

cs.AR cs.AI cs.DC cs.LG