AI Hot BriefingThis issue · All issues
HOTHF Daily PapersSep 28, 04:00Global🤖 AI

Nereus: Adaptive Parallelism for LLM Post-Training ⚡

Nereus introduces a cost-aware runtime system that adapts reinforcement learning post-training jobs for large language models, dynamically adjusting execution plans to handle changing resource conditions and significantly improving performance.

LLM TrainingParallel ComputingReinforcement LearningGPU Optimization
Nereus: Adaptive Parallelism for LLM Post-Training ⚡
Image linked from the original article · © original publisher

Introduction to Nereus

Nereus represents a breakthrough in large language model (LLM) post-training, addressing the critical challenge of adapting reinforcement learning (RL) post-training jobs across GPU clusters. The system coordinates multiple models across generation, inference, and training phases, but faces dynamic conditions that can render initial execution plans inefficient or infeasible.

The core innovation lies in Nereus's ability to adapt parallel execution plans while jobs are running, responding to changing resource availability, sequence length, memory pressure, and stage bottlenecks. This adaptive approach ensures optimal performance throughout the training process, unlike traditional fixed-parallelism systems that become inefficient as conditions change.

Challenges in LLM Post-Training

RL post-training for LLMs presents unique challenges due to the complex coordination required across multiple models and stages. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks, causing initially suitable execution plans to become slow or infeasible over time.

Adapting jobs where models share GPU details introduces significant challenges. These include deciding whether a new plan is worth the transition cost, reusing the job's distributed state effectively, and coordinating GPU transfers across models and stages. Traditional approaches struggle with these complexities, leading to suboptimal performance and resource utilization.

Nereus's Adaptive Approach

Nereus targets these challenges as a cost-aware runtime system that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits transitions only when the current plan becomes infeasible or the potential savings justify the transition cost.

The system represents each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units across all models and stages. This structured approach enables efficient adaptation while maintaining system coherence and minimizing overhead.

Technical Implementation

Nereus's implementation focuses on efficient state representation and transition management. By modeling each model-stage replica as an Elastic Model Unit, the system can track and coordinate distributed state across complex training environments. This representation enables precise control over the adaptation process.

The global transition graph serves as a critical component, ordering transformations and GPU transfers systematically. This approach ensures that adaptation happens in a coordinated manner across all models and stages, preventing conflicts and minimizing disruption to ongoing training processes. The graph-based ordering mechanism is essential for maintaining system stability during transitions.

Performance Evaluation

Nereus's performance was evaluated using a trace built from real data, demonstrating significant improvements over traditional approaches. Online TP/PP adaptation reduced average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling, showing substantial efficiency gains.

In a 1,000-step run reaching 1,024 GPUs, only six transitions were needed, consuming just 0.079% of total runtime. This demonstrates that the adaptation overhead is minimal compared to the performance benefits. The system's ability to make efficient transitions with minimal overhead is a key factor in its success.

“A run where sequence length grows monotonically would be the clean test — static plans are provably wrong there, so adaptation should show up clearly or not at all.”

— Tech commenter
TAKEAWAYNereus dynamically optimizes LLM post-training with adaptive parallelism, delivering 2-7x throughput improvements while
Source: HF Daily Papers · always refer to the original article
AI-curated from public sources for informational purposes only; images are hotlinked originals and copyright belongs to their respective publishers.
By Chaos Lab · 妙答星球AI