GUIDE teknik

Taggat yaram bu jaxaso

Taggat yaram bu jaxaso dafay gaawal taggat yaram ci reso neuronal bi ba noppi wàññi jëfandikoo mémoire ndax dafay def math yu bari ci 16-bit floating point ci barabu 32-bit.

2 simili jàngDañu mujjee yeesal

Résumé

It lets the same GPU train bigger models faster with almost no loss in accuracy.

Plongeur bu xóot

Taggat yaram bu yàgg dafay denc poid yi ak def math ci point flottant 32-bit (FP32). Precision mixed dafay jëfandikoo formaa 16-bit yu gëna ndaw (FP16 wala bfloat16) ngir matrix yu diis yi, boole ci denc 'kopi master' bu 32-bit ci poids yi ngir yeesali stabil. Ndax nimero 16-bit yi genn-wàll lañu ci dayo bi, moo gëna méngoo ak mémoire GPU bi, Tensor Cores yi ñoo leen di gëna gaaw ci lu tollu ci 2-8x. Li ñuy jàpp mooy FP16 ci diggante bu sew bi: gradient yu ndaw yi mën nañu wàcci ba amul dara. Fix standard bi mooy scaling perte, luy yokk perte bi ak facteur bu mag balaa backpropagation suko defee gradient yu ndaw yi des ci representable, ba noppi xaaj ko balaa yeesal diisaay bi. Apex bu NVIDIA ak AMP biñ tabax ci biir (Njàngale mu jaxaso ci boppam) ci PyTorch ak TensorFlow ñoo koy def ci boppam.

Gis-gis xarala

FP16 amul lenn ludul 5 bit exponent, loolu mooy joxe ab rang dynamique bu ndaw buy waral gradient bi di wàcci. Bfloat16 dafay denc 8 bit exponent (mënngoo ak FP32) waaye bit mantissa yu néew, moo tax daawu soxla eskalaasioŋ bu ñàkk - sabab bu mag bi TPUs ak GPUs yu bees yi bëgg ko. Tensor Cores dafay gaawal liggéey bi ci yokk ay operand 16-bit waaye di dajale ay somme yu néew ci FP32, ba noppi denc njubte gi ci barab yi njuumti somme yi di gëna yokk.

njeextalu pexe

Njëgg ak budget

Dogal yi architecture di jël dañuy indi njariñ ak njëgu liggéey bi ay at ci ginaaw.

dogal yu gëna leer

Njàngalem xarala yi dafay jàppale ekip yi ñu tànn li gën, te baña yam ci li gëna bees daal.

Xool kalite

Tanneef yu gëna baax ci wàllu ingeñër dina wàññi jafe-jafe yi ci wàllu wóor ci liggéey bi.

Ëlëgu tàggat yaram bu jaxaso

Precision mingi wéy di wàññeeku. Taggat FP8, jàppale ci NVIDIA Hopper ak Blackwell GPUs, mingi nekk standard ci model frontiere, ak gëstu ci FP4 ak formaa microscaling (MXFP) push gëna sori. Xaarandil kaadar yi di tann seen bopp njub ci couche bu nekk, aparey biy jëfandikoo formaa yu gëna sew, ak tàggat xam-xam kantite ngir dindi lignéer bi am ci digganté tàggat njub bu woyof ak inference, wàññi njëgu tàggat model yu am trillion paramètre.

Doxal ci àdduna dëgg

torch.cuda.amp.autocast bu PyTorch dafay boole ab boucle de formation ngir xaaj mémoire bi ak ñaari yoon limuy def ci benn GPU

Taggat xeetu làkk yu mag yu melni transformatër yu nuroo ak GPT ci bfloat16 ci kaw TPU yi ngir moytu tuning buy ñàkk

Defar ab dayo bu gëna mag ci GPU RTX konsomatër ci soppi ResNet nataal tàggat ci FP32 dem FP16

FP8 dafa jaxasoo ci NVIDIA H100 GPUs ngir wàññi njëgu tàggat model yu mag yi

Risk yi ak balustrade yi

Optimize benn benchmark mën na nëbb ñakk kattan yu gëna yaatu ci sistem bi.

Njëg li ñuy fay ci infrastructure yi ak ci toppatoo dañuy faral di suufeel.

Bu sistem yi di gëna xawa jafee xam, jafe-jafe yi am ci wàllu kaaraange ak seetlu mën nañu gëna bari.

Roadmap ngir samp gi

1

Mandargal latency, kalite, ak njëg yi laata ngay jëfandikoo.

2

Benchmark ci biir sargal ak done yu dëggu.

3

Jumtukaay bi di saytu njuumte yi, derive bi ak njeextalu jëfandikukat bi.

4

Waajal rollback ak yooni tontu ci jafe-jafe yi laata ngay eskale.

Weyal di banneexu

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Mixed Precision Training quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Tambalil quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Gis bi ci topp

Pseudo-Etiketu ak Tàggat Bopp

Laaj yi ñuy faral di laaj

What is Mixed Precision Training?

Taggat yaram bu jaxaso dafay gaawal taggat yaram ci reso neuronal bi ba noppi wàññi jëfandikoo mémoire ndax dafay def math yu bari ci 16-bit floating point ci barabu 32-bit. Dafay may benn GPU bi mu tàggat model yu gëna mag yi ci anam wu gëna gaaw te daanaka du ñàkk benn njub.

Lan moo waral tàggat yaram bu jaxaso di denc 'kopi master' bu 32 bit ci poid yi?

coppite ci diisaay bi dafay faral di tuuti lool; dajale leen ci 16-bit dina ñàkk njub, kon master copy bu mat njub dafay tëye yeesali njub.

Ban jafe-jafe la perte scaling di saafara ci FP16 formation?

FP16 amna rang dynamique bu gàtt, kon gradient yu ndaw yi mën nañu wër ba amul dara; yokk perte balaa backprop denc leen representable.

Lan mooy njariñu bfloat16 ci kaw FP16?

Bfloat16 dafay denc 8 bit exponent yu melni FP32, kon rang dynamique bi dafa yaatu te gradient bi bariwul lumuy wàcci.

naka lay def ba Tensor Cores mëna am njubte ci di jëfandikoo ay dugal 16-bit?

Tensor Cores yi dañuy yokk operand yu 16 bit waaye dañuy dajaloo ci 32 bit, loolu dafay tere njuumti sommage yi di yokk.

Lan mooy njariñ li gëna am solo ci jëfandikoo 16-bit ci palaasu 32-bit ci diiru tàggat yaram?

Nimero yu yam ci genn-wàll dañuy mëna ànd ak done yu bari ci mémoire GPU ba noppi may hardware buñ jagleel liggéey matrix math lu gëna gaaw yoon yu bari.