GUIDE Technique

Métriques d’évaluation RAG et RAGAS

RAG evaluation scores a retrieval-augmented generation system separately on retrieval and on generation, most commonly with four metrics: faithfulness, answer relevance, context precision and context recall.

  • 3 minutes de lecture
  • Dernière mise à jour
Sur cette page3 minutes de lecture
  1. Aperçu
  2. Plongée profonde
  3. Impact stratégique
  4. The Future of RAG Evaluation Metrics and RAGAS
  5. Mise en œuvre dans le monde réel
  6. Risques et garde-fous
  7. Feuille de route de mise en œuvre
  8. Continuez à explorer
  9. Questions fréquemment posées

Aperçu

RAGAS is an open-source framework that computes these metrics, largely using an LLM as a judge. Separate scores matter because they show whether a bad answer came from fetching the wrong passages or from misusing the right ones.

Plongée profonde

A RAG system can fail in two places: retrieval can fetch the wrong passages, or generation can ignore or distort the right ones. A single overall quality score hides which happened, so RAG evaluation splits the pipeline into measurable parts. RAGAS, an open-source framework introduced in a 2023 paper by Es and colleagues, popularized a set of core metrics that cover both halves. Faithfulness asks whether every claim in the answer is supported by the retrieved context; it catches hallucination even when the answer sounds plausible. Answer relevance asks whether the answer actually addresses the question, penalizing evasive or padded replies. Context precision asks whether the retrieved chunks that matter are ranked near the top rather than buried among noise. Context recall asks whether the retrieved context contains everything needed to produce a reference answer, so it requires a ground-truth answer for each test question. Most of these scores are computed by an LLM acting as a judge, broken into small steps. For faithfulness, the judge extracts individual statements from the answer and checks each against the context; the score is the fraction supported. For answer relevance, RAGAS generates questions that the answer would respond to and compares them with the original question using embeddings. Similar ideas appear in other tools, such as the TruLens RAG triad (context relevance, groundedness and answer relevance) and DeepEval. Two misconceptions are common. First, faithful does not mean correct: an answer can faithfully repeat an outdated or wrong document. Second, LLM-judge scores are not ground truth. They vary with the judge model and prompt, so teams spot-check them against human labels and compare systems under identical settings rather than treating a score like 0.85 as an absolute grade.

Impact stratégique

Coût et budget

Les décisions en matière d'architecture déterminent les performances et les coûts d'exploitation pendant des années.

Décisions plus claires

La formation technique aide les équipes à choisir la bonne pile, pas seulement la plus récente.

Contrôle qualité

De meilleurs choix d’ingénierie réduisent les incidents de fiabilité en production.

The Future of RAG Evaluation Metrics and RAGAS

RAG evaluation is shifting from occasional offline reports toward continuous checks, where a sample of production traffic is scored and regressions trigger alerts. Frameworks have been adding metrics for multi-turn conversations, agents and tool use, and metric definitions change between library versions, so pinning versions matters for comparable results. The largest open issue is judge reliability: LLM judges can show biases such as favoring longer answers, and agreement with human reviewers varies by domain. Calibrating judges against small human-labeled sets and pairing them with deterministic retrieval metrics is likely to remain good practice.

Mise en œuvre dans le monde réel

A bank's policy assistant scores 0.95 on context recall but 0.70 on faithfulness, which tells the team retrieval is fine and the model is adding unsupported details, so they tighten the prompt rather than rebuild the index.

A support team compares two embedding models on the same 200-question test set and keeps the one with higher context recall and context precision.

A hospital knowledge tool runs faithfulness scoring on a daily sample of real questions and flags answers with unsupported claims for human review before the pattern spreads.

A startup adds a reranker after seeing high context recall but low context precision, meaning the right chunks were retrieved but ranked below irrelevant ones.

Risques et garde-fous

  • L’optimisation d’un benchmark peut masquer des faiblesses plus larges du système.

  • Les coûts d’infrastructure et de maintenance sont souvent sous-estimés.

  • Les lacunes en matière de sécurité et d’observabilité peuvent se creuser à mesure que les systèmes deviennent plus complexes.

Feuille de route de mise en œuvre

  1. Définissez les objectifs de latence, de qualité et de coût avant la mise en œuvre.

  2. Benchmark dans des conditions de charge et de données réalistes.

  3. Surveillance des instruments pour détecter les erreurs, la dérive et l'impact sur l'utilisateur.

  4. Préparez les chemins de restauration et de réponse aux incidents avant la mise à l’échelle.

Continuez à explorer

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the RAG Evaluation Metrics and RAGAS quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Démarrer le quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Questions fréquemment posées

What is RAG Evaluation Metrics and RAGAS?

RAG evaluation scores a retrieval-augmented generation system separately on retrieval and on generation, most commonly with four metrics: faithfulness, answer relevance, context precision and context recall. RAGAS is an open-source framework that computes these metrics, largely using an LLM as a judge. Separate scores matter because they show whether a bad answer came from fetching the wrong passages or from misusing the right ones.

Which metric checks whether every claim in the answer is supported by the retrieved context?

Faithfulness measures the fraction of answer claims backed by the context, which targets hallucination.

Which of these metrics requires a ground-truth reference answer for each test question?

Context recall checks whether the retrieved context contains what is needed for the reference answer, so a reference must exist.

What does context precision measure?

Context precision rewards putting useful chunks high in the ranking instead of burying them among irrelevant ones.

How does RAGAS estimate answer relevance?

If questions generated from the answer resemble the real question, the answer is on topic; embeddings measure that similarity.

A system shows high context recall but low context precision. What is the most likely fix?

The right chunks are being retrieved but ranked poorly, which is a ranking problem a reranker often addresses.