Nereus: Adaptive Parallelism for LLM Post-Training π
Nereus addresses challenges in RL post-training for LLMs by adapting execution plans in real-time, improving efficiency and reducing latency.
Introduction to Nereus
Nereus is a novel approach for adaptive parallelism in the post-training phase of large language models (LLMs). It aims to optimize the execution of reinforcement learning (RL) post-training jobs on GPU clusters.
The system identifies and adapts to changes in resource availability, sequence length, memory pressure, and stage bottlenecks during the training process, ensuring that the execution plan remains efficient.
Challenges in RL Post-Training
Traditional execution plans can become inefficient or even infeasible as training progresses, due to changing conditions.
Nereus tackles these challenges by selecting a memory-feasible global plan and admitting transitions only when the current plan is infeasible or the savings repay the transition cost.
Elastic Model Units and Global Transition Graphs
Nereus represents each model-stage replica as an Elastic Model Unit, allowing for flexible and efficient management of resources.
It employs a global transition graph to order transformations and GPU transfers across all models and stages, optimizing the overall execution process.
Performance Improvements
In real-world data traces, Nereus reduces average step latency by 27.7% compared to initial fixed TP/PP layouts with DP scaling.
It also improves end-to-end 8 BPP throughput by 2.1 to 7.2 times over Open RL Hugging Face and 1.1 to 1.4 times over Verla cross-distributed clusters.
Transition Costs and Scheduler Decisions
The transition to a new execution plan incurs costs, and the scheduler must decide whether the remaining run is long enough to justify the reconfiguration.
Understanding the transition costs and the impact on optimizer state is crucial for effective scheduling and resource management.
βNereus addresses challenges in RL post-training for LLMs by adapting execution plans in real-time, ensuring that the execution plan remains efficient.β
β Songlin Jiang, Tuo Shi
By Chaos Lab Β· ε¦ηζηAI