Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning
AI 摘要
该论文测量了生成式AI工作负载的功耗曲线,用于数据中心基础设施规划。
主要贡献
- 提出了将高分辨率工作负载功耗测量与整个设施能源需求联系起来的方法。
- 公开了基于NVIDIA H100 GPU的AI训练、微调和推理工作负载的功耗数据集。
- 构建了一个自下而上的事件驱动的数据中心能源模型,用于评估整体设施能源需求。
方法论
通过测量AI工作负载的功耗,并使用数据中心能源模型将其扩展到整个设施层面。
原文摘要
The rapid growth of generative artificial intelligence (AI) has introduced unprecedented computational demands, driving significant increases in the energy footprint of data centers. However, existing power consumption data is largely proprietary and reported at varying resolutions, creating challenges for estimating whole-facility energy use and planning infrastructure. In this work, we present a methodology that bridges this gap by linking high-resolution workload power measurements to whole-facility energy demand. Using NLR's high-performance computing data center equipped with NVIDIA H100 GPUs, we measure power consumption of AI workloads at 0.1-second resolution for AI training, fine-tuning and inference jobs. Workloads are characterized using MLCommons benchmarks for model training and fine-tuning, and vLLM benchmarks for inference, enabling reproducible and standardized workload profiling. The dataset of power consumption profiles is made publicly available. These power profiles are then scaled to the whole-facility-level using a bottom-up, event-driven, data center energy model. The resulting whole-facility energy profiles capture realistic temporal fluctuations driven by AI workloads and user-behavior, and can be used to inform infrastructure planning for grid connection, on-site energy generation, and distributed microgrids.