LLM Reasoning 相关度: 6/10

LP-GEMM: Integrating Layout Propagation into GEMM Operations

César Guedes Carneiro, Lucas Alvarenga, Guido Araujo, Sandro Rigo
arXiv: 2604.04599v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

LP-GEMM通过布局传播优化GEMM序列,减少重复打包,显著提升计算性能。

主要贡献

  • 提出LP-GEMM,一种GEMM内核分解方法
  • 实现跨GEMM操作的打包布局传播
  • 在x86和RISC-V架构上验证了性能提升

方法论

通过分解GEMM内核,在连续的GEMM操作间传递打包布局,避免重复打包和解包操作。

原文摘要

In Scientific Computing and modern Machine Learning (ML) workloads, sequences of dependent General Matrix Multiplications (GEMMs) often dominate execution time. While state-of-the-art BLAS libraries aggressively optimize individual GEMM calls, they remain constrained by the BLAS API, which requires each call to independently pack input matrices and restore outputs to a canonical memory layout. In sequential GEMMs, these constraints cause redundant packing and unpacking, wasting valuable computational resources. This paper introduces LP-GEMM, a decomposition of the GEMM kernel that enables packing-layout propagation across sequential GEMM operations. This approach eliminates unnecessary data repacking while preserving full BLAS semantic correctness at the boundaries. We evaluate LP-GEMM on x86 (AVX-512) and RISC-V (RVV 1.0) architectures across MLP-like and Attention-like workloads. Our results show average speedups of 2.25x over OpenBLAS on Intel x86 for sequential GEMMs and competitive gains relative to vendor-optimized libraries such as Intel MKL. We demonstrate the practicality of the approach beyond microbenchmarks by implementing a standalone C++ version of the Llama-3.2 inference path using exclusively BLAS-level GEMM calls. These results confirm that leveraging data layout propagation between operations can significantly boost performance.

标签

GEMM BLAS Layout Propagation Linear Algebra Optimization

arXiv 分类

cs.DC cs.CV cs.LG