KV Cache Optimisation
Cache KV dafay denc caabi yi ak valeur yi transformateur bi jota xayma, moo tax du delloo liggéey ci token bu bees bu nekk - waaye mën na balloon ba gigabytes.
Résumé
KV cache optimization shrinks and manages that memory so models serve longer contexts to more users at once.
Plongeur bu xóot
Ci biir transformateur, jeton bu bees bu nekk dafay toppatoo jeton yi ko jiitu jaaraleko ci butoŋu attention (K) ak valeur (V). Recomputer K ak V ci sequence bi yépp ci jéego bu nekk dina nekk quadratic ak yàqu-yàqu, kon model yi dañu leen cache: cache KV. Li ci baaxul mooy dayo bi. Cache bi dafay màgg lineairement ak guddaayu sekans bi, dayo batch bi, diisaay yi, ak bopp yi, kon laaj contexte bu gudd mën na lekk memory GPU bu gëna bari ci diisaayu model bi ci boppam. Optimisation dafay jàppale lii ci wàll yu bari: mémoire paged (vLLM's PageAttention) dafay denc cache bi ci ay blok yu jëmmal ngir dindi xaaj-xaaj bi ak mëna séddoo; kantifikaasioŋ dafay denc K ak V ci 8-bit wala 4-bit; ak coppite yi ci architecture yu melni Fexe Laajte bu Grupp (GQA) ak Fexe Laajte yu bari (MQA) may boppu laajte yu bari ñu séddoo boppu caabi/valeur yu néew, di dagg dayo cache bi ci balluwaay bi.
Gis-gis xarala
PagedAttention dafay leble paging ci mémoire virtuel ci sistem operaasioŋ yi: cache bi dafa dëkk ci ay blok yu am dayo bu takku buñ mappe ci tablo seetlu, kon laaj yi dañuy jëfandikoo blok yi ñu soxla kese ak prefix yu nuróo (lu melni ab sistem buñ bokk) mën nañu joxoñ benn blok bi. Multi-head Latent Attention (MLA), ñu koy jëfandikoo ci xeetu DeepSeek, dafay komprime K ak V ci benn vecteur bu ndaw buñ bokk, di dagg mémoire bi ba noppi di tëye njub.
njeextalu pexe
Njëgg ak budget
Dogal yi architecture di jël dañuy indi njariñ ak njëgu liggéey bi ay at ci ginaaw.
dogal yu gëna leer
Njàngalem xarala yi dafay jàppale ekip yi ñu tànn li gën, te baña yam ci li gëna bees daal.
Xool kalite
Tanneef yu gëna baax ci wàllu ingeñër dina wàññi jafe-jafe yi ci wàllu wóor ci liggéey bi.
Ëlëgu KV Cache Optimisation
Ginaaw palanteer yi dañuy dem ba yegg ci téemeeri junni wala milioŋ ciy token, cache KV dafay nekk njëgu liggéey bi gëna am solo. Xaarandil compression cache bu metti ak dàq (dagg token yu néew), séddoo prefix cross-request ni default, dindi cache bu sedd ci CPU wala NVMe, ak architecture yu melni MLA ak GQA nekk standard. Doxalal cache bi dafay gëna niroo ak hierarchie mémoire bu mat sëkk ak ay niveau ak prefetching yu xarañ.
Doxal ci àdduna dëgg
vLLM's PagedAttention dafay liggéey ci sesioŋ yu bari yuy waxtaan ci benn yoon, ci defar ay bloku KV te du am xaajalug mémoire
Grouped-Query Fexe ci xeetu Llama wàññi dayo cache KV suko defee contexte yu gëna gudd mëna méngoo ak mémoire GPU
Kantite cache KV ci 8-bit (KV8) ngir xaaj mémoire cache bi ci diiru résumé dokimaa yu gudd
Cache prefix biy jëfandikoowaat bloku KV yi ci benn sistem buñ bokk laaj ci ay junni laaj API
Risk yi ak balustrade yi
Optimize benn benchmark mën na nëbb ñakk kattan yu gëna yaatu ci sistem bi.
Njëg li ñuy fay ci infrastructure yi ak ci toppatoo dañuy faral di suufeel.
Bu sistem yi di gëna xawa jafee xam, jafe-jafe yi am ci wàllu kaaraange ak seetlu mën nañu gëna bari.
Roadmap ngir samp gi
Mandargal latency, kalite, ak njëg yi laata ngay jëfandikoo.
Benchmark ci biir sargal ak done yu dëggu.
Jumtukaay bi di saytu njuumte yi, derive bi ak njeextalu jëfandikukat bi.
Waajal rollback ak yooni tontu ci jafe-jafe yi laata ngay eskale.
Weyal di banneexu
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the KV Cache Optimization quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Gis bi ci topp
KV Cache
Laaj yi ñuy faral di laaj
What is KV Cache Optimization?
Cache KV dafay denc caabi yi ak valeur yi transformateur bi jota xayma, moo tax du delloo liggéey ci token bu bees bu nekk - waaye mën na balloon ba gigabytes. KV cache optimisation dafay wàññi ba noppi yor memory bi suko defee model yi di mëna jëfandikoo contexte yu gëna gudd ngir jëfandikukat yu bari benn yoon.
Luy denc ci cache KV te lu tax?
Cache caabi ak valeur yu bawoo ci token yu njëkk yi dafay moytu defaraat calcul attention ci token bu bees bu nekk, di sakkanal jotu liggéey.
Lu tax cache KV mën nekk jafe-jafe mémoire?
Dayo cache bi dafay méngoo ak guddaayi contexte bi ak concurrence bi, kon laaj yu gudd yi mën nañu jëfandikoo mémoire GPU bu bari.
Ban konsepti sistem jëfandikoo la PageAttention lebal?
PagedAttention dafay denc cache bi ci ay blok yu am dayo bu fiks bu nekk ci biir benn tablo, melni paging OS.
naka lay def ba Laajte yuñ boole ak Laajte yu bari di wàññi dayo cache KV?
Séddoo boppu caabi/valeur ci boppu laaj yu bari dafay tekki ni vecteur K ak V yu néew lañu wara denc.
Lan mooy benn ci njariñu séddoo prefix ci cache KV?
Benn sistem buñ bokk dafay defar ay duggal KV yu nuróo, suko defee laaj yu bari mën nañu joxoñ benn blok bi, duñu leen di ñaari yoon.