GUIDE teknik

Kubernetes ngir ML

Kubernetes sistem open-source la buy waajal, yamale ak tàmbaliwaat ci saasi prograam yiñ def ci conteneur ci masin yu bari.

2 simili jàngDañu mujjee yeesal

Résumé

For machine learning, it lets teams pack GPU-hungry training jobs and latency-sensitive model servers onto shared hardware without babysitting individual servers.

Plongeur bu xóot

Ñu ngi ko njëkka tabax ci Google ngir doxal ay sarwis web, Kubernetes dafay jàppee sa cluster ni benn pool bu mag bu CPU, mémoire, ak GPUs, ginaaw ga mu tànn ban masin mooy doxal conteneur bu nekk. Ekkipu ML yi dañu ciy wéeru ndax liggéey yi dañu bari te seer: ab tàggat yaram mën na soxla juróom ñatti GPUs ci jiroom benni waxtu, kon dara. Kubernetes dafay jàppale pod bi ci node bu am GPU yu amul fayda, su liggéey bi jeexee dafay bàyyi hardware bi. Dafay dundal itam serveur inference yi, di tàmbaliwaat conteneur yu tass yi, di tasaare ay replica ci masin yi ngir mëna dëgër. Jumtukaay yiñ tabax ci kaw, lu melni Kubeflow, Ray, ak KServe, dañuy yokk ay piyees yu jëm ci ML lu ci melni operatëri tàggat yuñ séddale, ajustement hyperparamètre, ak poñ yu mujj ci modelu autoscaling, suko defee gëstukati done yi di liggéey ak abstraction yu gëna kawe ci barabu YAML bu ñor.

Gis-gis xarala

Kubernetes dafay jox GPU yi jaaraleko ci ay plugin yuy siiwal ay jumtukaay yu melni nvidia.com/gpu, te scheduler bi dafay méngale ak laaj pod yi. Taints ak tolerans yi deñuy tere CPU yi liggéey bu yomb ci node GPU yu seer yi, ci noonu la tànneefi node yi ak sàrti affinite yi di pin tàggat ci ay hardware yuñ tànn. Ngir tàggat GPU yu bari, operatër yi dañuy sos kuréelu pod yuy gise seen biir ba noppi doxal kaadar yu melni PyTorch DDP wala Horovod, di weccoo gradient ci reso cluster bi jëfandikoo NCCL.

njeextalu pexe

Njëgg ak budget

Dogal yi architecture di jël dañuy indi njariñ ak njëgu liggéey bi ay at ci ginaaw.

dogal yu gëna leer

Njàngalem xarala yi dafay jàppale ekip yi ñu tànn li gën, te baña yam ci li gëna bees daal.

Xool kalite

Tanneef yu gëna baax ci wàllu ingeñër dina wàññi jafe-jafe yi ci wàllu wóor ci liggéey bi.

Ëlëgu Kubernetes ngir ay liggéey ML

Xaarandi mboolem ML bu gëna dëgër: gang scheduling biy ubbi pods yi ñuy séddale ci benn yoon wala benn, GPU bu fractional ak time-sliced ​​GPU di séddoo suko defee liggéey yu bari yu woyof di séddoo benn kart, ak topology-xam-xam plasement buy sargal NVLink interconnects yu gaaw. Inference bu amul serwër ci Kubernetes, escaleer poñ yu mujj yi ba zero ci diggante laaj yi, mingi màgget. Bi model yi di gëna yokk, ñiy waajal dañuy gëna déggoo ci cluster yu bari ak niir yi, ba noppi sistem yu ñuy séddoo ci rang yu melni Kueue ak Volcano ñu ngi nekk standard ngir jëfandikoo kàttanu GPU yu néew yi.

Doxal ci àdduna dëgg

Benn labo gëstu dafay jëfandikoo Operatëru Tàggat Kubeflow ngir tàmbali liggéeyu tàggat PyTorch bu 32-GPU ci ñeenti node, su ko defee mu bàyyi GPU yi ci saasi suñu daje.

Benn kompiñi e-commerce dafay jëfandikoo xeetu xalaatam ak KServe, biy autoscale replicas ci jamonoy njaay flash ba noppi wàcci ci guddi gi.

Benn bànk dafay doxal guddi gu nekk ay liggéey yuy joxe poñ ci Kubernetes CronJobs, di leen raŋ ci node CPU yu bees yi ngir ñu baña joŋante ak trafik biy liggéey ci bëccëg.

Benn startup dafay jëfandikoo Ray ci kaw Kubernetes ngir doxal ay hyperparamètre yu paralel, di wëlbati ay fukki-fukki pod yu gàtt ci misaal yi ngir wàññi njëg yi.

Risk yi ak balustrade yi

Optimize benn benchmark mën na nëbb ñakk kattan yu gëna yaatu ci sistem bi.

Njëg li ñuy fay ci infrastructure yi ak ci toppatoo dañuy faral di suufeel.

Bu sistem yi di gëna xawa jafee xam, jafe-jafe yi am ci wàllu kaaraange ak seetlu mën nañu gëna bari.

Roadmap ngir samp gi

1

Mandargal latency, kalite, ak njëg yi laata ngay jëfandikoo.

2

Benchmark ci biir sargal ak done yu dëggu.

3

Jumtukaay bi di saytu njuumte yi, derive bi ak njeextalu jëfandikukat bi.

4

Waajal rollback ak yooni tontu ci jafe-jafe yi laata ngay eskale.

Weyal di banneexu

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Kubernetes for ML Workloads quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Tambalil quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Gis bi ci topp

Test A/B ngir xeetu ML

Laaj yi ñuy faral di laaj

What is Kubernetes for ML Workloads?

Kubernetes sistem open-source la buy waajal, yamale ak tàmbaliwaat ci saasi prograam yiñ def ci conteneur ci masin yu bari. Ngir jàngu masin, dafay may ekip yi ñu defar ay liggéey yu xiif ci GPU ak ay serwër model yu sensible ci latency ci hardware buñ bokk te duñu toppatoo serwër yi.

Lan mooy liggéey bi njëkk ci Kubernetes scheduler ngir ML?

Programmeur bi dafay méngale sàrti pod bi (CPU, mémoire, GPU) ak node yi jàppandi ba noppi def pod bi ci barab bi war. Laalul kodu model wala done.

naka la ko Kubernetes di def ba GPU yi mëna am ci pod bi?

Plugin yi ci aparey yi dañuy wane GPU yi ni jumtukaay buñ mëna jàppale (lu melni, nvidia.com/gpu), may pods yi ñu laaj leen, ak jàppalekat bi di topp disponibilite bi.

Lan mooy taints ak tolerans yi ñuy faral di jëfandikoo ci cluster ML?

Taint dafay dàq pods yi ci node bi; pods yi am toleransi bu méngoo rek ñoo mëna wàcci fa. Loolu dafay denc node GPU yu néew ngir liggéey yi soxla GPU.

Ban jumtukaay mooy yokk kàttanu ML yu melni operatëri tàggat yuñ séddale ci kaw Kubernetes?

Kubeflow dafay boole ay liggéeyu ML ci Kubernetes, lu ci melni tàggat operatër yi, pipeline yi, ak tuning, suko defee ekip yi moytu bind config cluster bu woyof.

Lan moo tax Kubernetes méngoo bu baax ak liggéey yu bari ci ML?

Taggat yaram dafa spiky: GPU yu bari ci diir bu gàtt, ba noppi amul benn. Kubernetes dafay dugal liggéey bi su jumtukaay yi amul xaalis ba noppi bàyyi leen ñu jeexal, gëna baaxal jëfandikoo gi.