A continuaciónSiguiente guía
RAG especulativo y redacción aumentada de recuperación
Técnico
GUÍA Técnica
Cache-augmented generation (CAG) preloads an entire, fixed knowledge base into a language model's key-value (KV) cache once and reuses that cache for every question, instead of retrieving documents per query as RAG does.
It matters for small, stable corpora because it removes retrieval errors and index maintenance while answering faster than re-reading the documents each time.
Cache-augmented generation, or CAG, replaces the retrieval step with preloading. Instead of searching a knowledge base for each question, the system feeds the entire knowledge base to the model once, saves the resulting key-value (KV) cache, and reuses that cache for every subsequent query. The approach was described in a December 2024 paper titled Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks. The KV cache is the set of intermediate attention keys and values a transformer computes for every token it has read. Normally it is built during a single request and discarded. In CAG, the documents are processed ahead of time, the cache is stored, and each question is appended after the cached prefix. The model only computes the new question tokens before generating, so answers start faster than if the documents were re-read each time. After answering, the appended tokens are truncated so the cache returns to its documents-only state for the next query. The benefits are no retrieval errors, since every document is always present, no embedding index or vector database to maintain, and lower per-query latency than naive long context. The limits follow directly: the whole corpus must fit in the model's context window, the model must still reason well over that much text, and any change to the documents means rebuilding the cache. CAG therefore fits small, stable knowledge bases such as a product manual, a policy set, a syllabus or a fixed regulatory text. It is a poor fit for large, fast-changing or permission-filtered collections. A common misconception is that CAG trains the model. Nothing is fine-tuned; the weights are unchanged. Commercial prompt caching from API providers is a hosted relative of the same idea.
Las decisiones de arquitectura impulsan el rendimiento y los costos operativos durante años.
La educación técnica ayuda a los equipos a elegir la pila adecuada, no sólo la más nueva.
Mejores opciones de ingeniería reducen los incidentes de confiabilidad en la producción.
CAG's practical range grows as context windows lengthen and KV cache compression improves, but it stays bounded by memory and by how well models reason over very long inputs. Hybrid designs are a natural next step: cache a stable core of frequently needed documents and retrieve the long tail on demand. Research on cache compression, offloading caches to cheaper storage, and choosing which tokens to keep may make larger preloaded corpora workable. For now the sensible test is empirical: compare CAG, RAG and plain long context on your own questions, costs and update frequency.
A company with a 60-page employee handbook preloads it into the KV cache of an open-weight model, so every HR question is answered with the full handbook present and no vector database to maintain.
A software vendor caches its product manual for a support assistant and rebuilds the cache only when a new manual version is released.
An instructor preloads a course syllabus and lecture notes so a study assistant can answer student questions quickly throughout a semester.
A compliance team uses a hosted API's prompt caching to keep a fixed regulatory text as a cached prefix, paying reduced rates for repeated reads across many questions.
La optimización de un punto de referencia puede ocultar debilidades más amplias del sistema.
Los costos de infraestructura y mantenimiento a menudo se subestiman.
Las brechas de seguridad y observabilidad pueden crecer a medida que los sistemas se vuelven más complejos.
Defina objetivos de latencia, calidad y costos antes de la implementación.
Comparación en condiciones realistas de carga y datos.
Monitoreo de instrumentos para detectar errores, deriva e impacto para el usuario.
Prepare rutas de reversión y respuesta a incidentes antes de escalar.
Free newsletter
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Cache-augmented generation (CAG) preloads an entire, fixed knowledge base into a language model's key-value (KV) cache once and reuses that cache for every question, instead of retrieving documents per query as RAG does. It matters for small, stable corpora because it removes retrieval errors and index maintenance while answering faster than re-reading the documents each time.
CAG computes the KV cache for the corpus ahead of time and reuses it for every query.
Keys and values from attention layers are cached so earlier tokens need not be recomputed.
Removing appended tokens resets the cache to its clean preloaded state for the next question.
Because every document is preloaded, the corpus is capped by the context window and available memory.
The documents' computation is already done and stored, so only the question is processed per query.
sigue aprendiendo
Más guías seleccionadas para este tema.
A continuaciónSiguiente guía
RAG especulativo y redacción aumentada de recuperación
Técnico