NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
AI 摘要
NestPipe通过嵌套流水线解决大规模推荐模型训练中的数据移动瓶颈,实现了高效且一致的训练。
主要贡献
- 提出了Dual-Buffer Pipelining (DBP) 缓解lookup瓶颈
- 提出了Frozen-Window Pipelining (FWP) 重叠通信和计算
- 在1536个workers上实现了高达3.06x的加速和94.07%的扩展效率
方法论
通过嵌套流水线,分别在批次间和批次内进行优化,利用双缓冲和冻结窗口技术,减少数据移动开销。
原文摘要
Modern recommendation models have increased to trillions of parameters. As cluster scales expand to O(1k), distributed training bottlenecks shift from computation and memory to data movement, especially lookup and communication latency associated with embeddings. Existing solutions either optimize only one bottleneck or improve throughput by sacrificing training consistency. This paper presents NestPipe, a large-scale decentralized embedding training framework that tackles both bottlenecks while preserving synchronous training semantics. NestPipe exploits two hierarchical sparse parallelism opportunities through nested pipelining. At the inter-batch level, Dual-Buffer Pipelining (DBP) constructs a staleness-free five-stage pipeline through dual-buffer synchronization, mitigating lookup bottlenecks without embedding staleness. At the intra-batch level, we identify the embedding freezing phenomenon, which inspires Frozen-Window Pipelining (FWP) to overlap All2All communication with dense computation via coordinated stream scheduling and key-centric sample clustering. Experiments on production GPU and NPU clusters with 1,536 workers demonstrate that NestPipe achieves up to 3.06x speedup and 94.07% scaling efficiency.