Gaawaay bu xóot ak Megatron
DeepSpeed (Microsoft) ak Megatron-LM (NVIDIA) ñooy losisel yiy tax modeli tàggat yu am ay miliyaar ciy parametre ci ay junni GPU mëna dem.
Résumé
Without them, today's frontier models simply could not fit in memory or finish training in a reasonable time.
Plongeur bu xóot
Taggat model bu rëy ci benn GPU lu jafe la ndax poid yi, gradient yi ak stade optimiser yi mënul ànd. Stack yooyu dañu xaaj liggéey bi ci GPU yu bari. Megatron-LM moo njëkka amal paralelism tensor, di dagg matrix bu nekk ci biir layer bu nekk ci GPU yi, boole ci parallelism pipeline, biy def ay layer yu wuute ci GPU yu wuute. Siñaale DeepSpeed mooy ZeRO (Zero Redundancy Optimizer), mooy xaaj staadu optimisatër bi, degrade yi, ak parametre yi ci GPU yi ci barab bi ñu leen di toppandoo, di dagg mémoire bu GPU bu nekk. Ñaari mbir yooyu dañu leen di faral di boole (megatron-gaawaay bu xóot) ngir tàggat model yu melni BLOOM-176B ak Megatron-Turing NLG. Dañuy yokk itam njaxasu njub, saytu aktiwite, ak yebbi ci CPU wala NVMe suko defee model yu mag yi di tàggat ci hardware bu néew.
Gis-gis xarala
ZeRO amna ñatti etape ngir yokk sakkanal mémoire: etape 1 shards optimizer states, etape 2 itam shards gradients, ak etape 3 shards parametre yi ci seen bopp, dajale leen ci laaj ci diir yi ñuy jaar ci kanam ak ci ginaaw. Buñu ko boole ak paralelism tensor (ci biir couche) ak paralelism pipeline (ci biir couche), loolu dafay forme 'paralelism 3D.' Tension bi gëna am solo mooy jokkoo bi ci kaw: bépp xaaj buñ xaaj dafay yokk trafik GPU-to-GPU, kon ingenieur yi dañuy tune xaaj bi ngir mëna wéy di gaaw NVLink ak InfiniBand lëkkalekaay yu fees.
njeextalu pexe
Njëgg ak budget
Dogal yi architecture di jël dañuy indi njariñ ak njëgu liggéey bi ay at ci ginaaw.
dogal yu gëna leer
Njàngalem xarala yi dafay jàppale ekip yi ñu tànn li gën, te baña yam ci li gëna bees daal.
Xool kalite
Tanneef yu gëna baax ci wàllu ingeñër dina wàññi jafe-jafe yi ci wàllu wóor ci liggéey bi.
Ëlëgu DeepSpeed ak Megatron
Xaarandi lëkkaloo bu gëna dëgër ak FSDP (Fully Sharded Data Parallel) bu PyTorch, bi jël xalaati ZeRO yu bari, di nëbb diggante stack gëstu ak kaadar core. Xeetu jegewaale yiy doxal compilatër bi ak waajal paralelism otomatik yi seen mébet mooy dindi ajustement manuel bi. Bi clusters tàggat yaram di màgg ba yegg ci téemeeri junni accelerator, tolerans ci njuumte, elastic scaling, ak jokkoo buy jaxasoo ak ordinatër nekk nañu frontière ingenieur yi gëna am solo, ci wetu jàppale hardware yu bees yu melni NVIDIA Blackwell ak chips tàggat yaram.
Doxal ci àdduna dëgg
Taggat xeetu BLOOM-176B bu ubbeeku ci làkk yu bari, jëfandikoo Megatron-DeepSpeed buñ boole muy boole téemeeri GPU.
ak NVIDIA ñu ngi tàggat xeetu Megatron-Turing NLG bu am 530 milyaar ci paralelism 3D.
ZeRO-Offload dafay may gëstukat yi ñu mëna defar ay model yu bari ay paramet ci benn GPU station de travail ci di tuuru stade optimisateur yi ci RAM CPU bi.
Jëfandikool checkpointing aktivasioŋ ci stack yii ngir mëna ànd ak palanteer yu gëna gudd ci xaymawaat aktivasioŋ yi ci barabu denc leen ñépp.
Risk yi ak balustrade yi
Optimize benn benchmark mën na nëbb ñakk kattan yu gëna yaatu ci sistem bi.
Njëg li ñuy fay ci infrastructure yi ak ci toppatoo dañuy faral di suufeel.
Bu sistem yi di gëna xawa jafee xam, jafe-jafe yi am ci wàllu kaaraange ak seetlu mën nañu gëna bari.
Roadmap ngir samp gi
Mandargal latency, kalite, ak njëg yi laata ngay jëfandikoo.
Benchmark ci biir sargal ak done yu dëggu.
Jumtukaay bi di saytu njuumte yi, derive bi ak njeextalu jëfandikukat bi.
Waajal rollback ak yooni tontu ci jafe-jafe yi laata ngay eskale.
Weyal di banneexu
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the DeepSpeed and Megatron Training Stacks quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Gis bi ci topp
GPTQ ak AWQ ginaaw tàggat
Laaj yi ñuy faral di laaj
What is DeepSpeed and Megatron Training Stacks?
DeepSpeed (Microsoft) ak Megatron-LM (NVIDIA) ñooy losisel yiy tax modeli tàggat yu am ay miliyaar ciy parametre ci ay junni GPU mëna dem. Suñu leen amul, xeetu frontiere yu tay yi mënu ñu nekk ci memory wala jeexal tàggat ci diir bu gàtt.
Lan mooy jubluwaay bu njëkk ci optimisatëru ZeRO bu DeepSpeed?
ZeRO (Zero Redondance Optimizer) dafay dindi redondance ci mémoire bi ci xaaj etaa yi, gradient yi, ak parametre yi ci GPU yi moo gën ñu koy toppandoo ci aparey bu nekk.
Paralelism tensor, ni ñu ko njëkka amal ci Megatron-LM, naka lay xaajalee model bi?
Parallelism tensor dafay xaaj math bi ci biir benn couche (lu melni matrix bu mag buy yokk) ci GPU yu bari, muy xaaj bi ci biir couche bi.
Ban etape ZeRO mooy joxe sakkanal mémoire bu gëna mag ci sharding model parametre yi ci seen bopp?
ZeRO Stage 3 dafay xaaj parametre yi ci kaw gradient yi ak stade optimisatër yi, dajale leen ci laaj, di joxe wàññikug mémoire bi gëna mag.
Lan mooy 'checkpointing' ngir sakkanal mémoire bi ci diiru tàggat yaram?
Checkpointing biy tàmbali dafay denc lu néew ci diggante yi, ba noppi di leen xaymaat ci diiru backpropagation, di jënd yeneen xayma ngir wàññi memory bi.
Lan moo waral ñuy woowe boole pexe yooyu 'parallelism 3D'?
Paralelism 3D dafay boole ñatti pexe ortogonal: paralelism done, paralelism tensor (ci biir couche), ak paralelism pipeline (ci biir couche).