Nereus: Adaptive Parallelism for LLM Post-Training 🤖
Adaptive parallelism addresses challenges in large language model post-training, optimizing GPU cluster performance.
Adaptive Parallelism for LLMs
Reinforcement learning post-training coordinates multiple models across generations, inference, and training on GPU clusters. Several factors can change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks.
Challenges of Job Adaptation
Adapting a job whose models share GPUs presents significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages.
Nereus' Approach
Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model.
Optimizing GPU Cluster Performance
By addressing the challenges of job adaptation, Nereus optimizes the performance of GPU clusters, ensuring efficient use of resources and minimizing bottlenecks.
By Chaos Lab · 妙答星球AI