Bopp yu bari yu nëbbu
Multi-Head Latent Attention (MLA) jumtukaay la buy bàyyi xel ci nit ñi, ñu dugal ko ci DeepSeek-V2, mooy gëna xëcc cache bu am valeur bu am mémoire ci benn vecteur bu ndaw buñ bokk.
Résumé
It lets large language models run with far less GPU memory while keeping quality close to standard attention.
Plongeur bu xóot
Su transformatër bi defaree mbind, dafay denc ab caabi ak vecteur valeur ngir bépp token bu weesu ci 'cache KV.' Cache boobu dafay màgg ak guddaayu contexte bi ba noppi mooy ëpp doole ci jëfandikoo mémoire bi ci inference bi. MLA dafay wecci vecteur yu bari yu am caabi/valeur ak benn vecteur latent bu rang bu woyof ci token bu nekk, ba noppi mu projet latent bi dellu ci caabi bopp bu nekk ak valeur ci fly. Ndax latent bu kompact bi kese lañuy denc ci cache, DeepSeek-V2 dafa wax ni dafa dagg memory cache KV lu ëpp 90% ci wàllu xel mu bari, loolu mooy tax muy am contexte yu gëna gudd ak dayo lote yu gëna mag. Li gëna am solo mooy matris yiy wane seen kaw mën nañu leen boole ci yeneen poid, kon MLA mën na def compression bi te amul benn perte buñu mëna natt ci kalite modeling bi.
Gis-gis xarala
MLA dafay def benn compression bu rang bu woyof: stade bu nëbbu bu token bu nekk dañu koy projecte ci vecteur bu ndaw bu nëbbu, ba noppi ñu tàqale matrices yu ñuy projecte ci kaw ngir tabaxaat caabi ak valeur yu bopp bu nekk. Benn pexe bu am xel mooy 'absorbe' poids projection ci kaw ci laaj bi ak projection yiy génne, suko defee model bi du musa am caabi/valeur yu mat sëkk ci diiru inference. Rotary position embedding yi deñukoy jëfandikoo ci yoonu butoŋu decouplé, ndax rotation mënul absorbé ci anam wu mel nii, ba noppi denc leeral yi ci position.
njeextalu pexe
Gaawaay ak yaatuwaay
Liggéeyukaay yi ci làkk yi mën nañu gëna gaaw te duñu yàq deggoo gi.
Dugg ak yegg
Dafay yaatal jëfandikoo gi ci làkk yi ak ci anam yi ñuy jokkoo.
dogal yu gëna leer
Ekip yi mën nañu gëna yàgg ci àtte ci jamono ji otomatisation di liggéey ci baamtu.
Ëlëgu bàyyi xel bu nëbbu ci bopp yu bari
MLA jàppale DeepSeek-V2 ak V3 ñu gëna xéewale ngir liggéey ci escale, te pexem bi mingi tasaaroo ginaaw bi ekip yi toppee inference yu gëna yomb. Xaarandil kompresioŋ bu nëbbu bu nuroo ak MLA ngir boole ay couche Mixture-of-Expert yu bari, cache yu bari, ak decodage speculatif ci model yu ubbeeku yi ci kanam. Gëstukat yi ñu ngi jàngat itam ba fu dimension latente bi mëna wàññeeku balaa kalite bi di wàññeeku, ak ndax benn xalaat bu amul rang bi mën na gëna xëcc xel ci diiru tàggat yaram, te baña yam ci inference.
Doxal ci àdduna dëgg
DeepSeek-V2/V3 xeetu waxtaan ak emprent mémoire GPU bu gëna ndaw ci laaj bu nekk
Doxal ab këyitu laaj bu gudd di tontu fi ab cache KV bu mag di jeexal VRAM
Yokk dayo inference ci GPU bu takku ndax bu nekk ci ñoom du denc ludul benn vecteur bu nëbbu
Aktiwise palanteer yu gëna gudd ci aparey commodite ngir assistant yuñ yokk ci seet
Risk yi ak balustrade yi
Lépp lu jaarul yoon mën na dugg ci rapoor yi, jàppale ci liggéey bi, wala ci njariñu gëstu bi.
Sensibilite bu gaaw mën na jur njariñ yu wuute ci laajte yu noonu mel.
Done yu am solo mën nañu feeñ sudee seytu jëfandikoo gi néew doole.
Roadmap ngir samp gi
Mandargal formaa génne gi, melokaan bi, ak standard kalite yi laata ngay dugal ko.
Tontu yu am solo ak balluwaay yu wóor saa yu dëggu bi di am solo.
Fexeel am barabu xool nit ñi ngir am njariñ yu am solo.
Toppal anami gacce yi ak di faral di tàggataat ay laaj wala def-liggéey.
Weyal di banneexu
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Multi-Head Latent Attention quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Gis bi ci topp
Laajte yu bari
Laaj yi ñuy faral di laaj
What is Multi-Head Latent Attention?
Multi-Head Latent Attention (MLA) jumtukaay la buy bàyyi xel ci nit ñi, ñu dugal ko ci DeepSeek-V2, mooy gëna xëcc cache bu am valeur bu am mémoire ci benn vecteur bu ndaw buñ bokk. Dafay may modeli làkk yu mag yi ñu mëna dox ak memory GPU bu néew, boole ci tëye kalite bi jege bàyyi xel buñ miin.
Lan mooy jafe-jafe bi gëna mag bi ñu nara wàññi?
MLA dafay xool cache KV, biy màgg ak guddaayu contexte bi ba noppi di ëpp doole ci mémoire bi ci defar mbind.
naka la MLA di wàññi cache KV bi?
MLA dafay denc benn vecteur bu nëbbu ci token bu nekk, ba noppi defaraat caabi yi ak valeur yi ci projection bi ci kaw.
Ban model moo njëkka dugal xel mu nëbbu ci bopp yu bari?
DeepSeek moo dugal MLA ci xeetu DeepSeek-V2 ba noppi yóbbu ko ci DeepSeek-V3.
Lan moo tax MLA soxla yoon wu wuute 'decouplé' ngir position rotary?
Coppite rotary bi mënul duggu ci yeneen matrix yu diis, moo tax MLA dafay denc benn composant bu ndaw buñu decouple ngir yóbbu xibaar ci position.
Lu tollu ci ñaata KV-cache memory la DeepSeek-V2 xamle ci MLA ak bàyyi xel ci bopp yu bari?
DeepSeek-V2 dafa wax ni dafa dagg mémoire cache KV lu ëpp 90%, loolu taxna ñu mëna am contexte yu gëna gudd ak ay lots yu gëna bari.