이 페이지에서3분 읽기
개요
RAGAS is an open-source framework that computes these metrics, largely using an LLM as a judge. Separate scores matter because they show whether a bad answer came from fetching the wrong passages or from misusing the right ones.
심층 분석
A RAG system can fail in two places: retrieval can fetch the wrong passages, or generation can ignore or distort the right ones. A single overall quality score hides which happened, so RAG evaluation splits the pipeline into measurable parts. RAGAS, an open-source framework introduced in a 2023 paper by Es and colleagues, popularized a set of core metrics that cover both halves. Faithfulness asks whether every claim in the answer is supported by the retrieved context; it catches hallucination even when the answer sounds plausible. Answer relevance asks whether the answer actually addresses the question, penalizing evasive or padded replies. Context precision asks whether the retrieved chunks that matter are ranked near the top rather than buried among noise. Context recall asks whether the retrieved context contains everything needed to produce a reference answer, so it requires a ground-truth answer for each test question. Most of these scores are computed by an LLM acting as a judge, broken into small steps. For faithfulness, the judge extracts individual statements from the answer and checks each against the context; the score is the fraction supported. For answer relevance, RAGAS generates questions that the answer would respond to and compares them with the original question using embeddings. Similar ideas appear in other tools, such as the TruLens RAG triad (context relevance, groundedness and answer relevance) and DeepEval. Two misconceptions are common. First, faithful does not mean correct: an answer can faithfully repeat an outdated or wrong document. Second, LLM-judge scores are not ground truth. They vary with the judge model and prompt, so teams spot-check them against human labels and compare systems under identical settings rather than treating a score like 0.85 as an absolute grade.
전략적 영향
비용 및 예산
아키텍처 결정은 수년 동안 성능과 운영 비용을 결정합니다.
더 명확한 결정들
기술 교육은 팀이 최신 스택뿐만 아니라 올바른 스택을 선택하는 데 도움이 됩니다.
품질 관리
더 나은 엔지니어링 선택은 생산 시 신뢰성 사고를 줄입니다.
The Future of RAG Evaluation Metrics and RAGAS
RAG evaluation is shifting from occasional offline reports toward continuous checks, where a sample of production traffic is scored and regressions trigger alerts. Frameworks have been adding metrics for multi-turn conversations, agents and tool use, and metric definitions change between library versions, so pinning versions matters for comparable results. The largest open issue is judge reliability: LLM judges can show biases such as favoring longer answers, and agreement with human reviewers varies by domain. Calibrating judges against small human-labeled sets and pairing them with deterministic retrieval metrics is likely to remain good practice.
실제 구현
A bank's policy assistant scores 0.95 on context recall but 0.70 on faithfulness, which tells the team retrieval is fine and the model is adding unsupported details, so they tighten the prompt rather than rebuild the index.
A support team compares two embedding models on the same 200-question test set and keeps the one with higher context recall and context precision.
A hospital knowledge tool runs faithfulness scoring on a daily sample of real questions and flags answers with unsupported claims for human review before the pattern spreads.
A startup adds a reranker after seeing high context recall but low context precision, meaning the right chunks were retrieved but ranked below irrelevant ones.
위험 및 가드레일
하나의 벤치마크를 최적화하면 더 광범위한 시스템 약점을 숨길 수 있습니다.
인프라 및 유지 관리 비용은 종종 과소평가됩니다.
시스템이 더욱 복잡해짐에 따라 보안 및 관찰 가능성의 격차가 커질 수 있습니다.
구현 로드맵
구현하기 전에 지연 시간, 품질, 비용 목표를 정의하세요.
현실적인 로드 및 데이터 조건에서 벤치마킹합니다.
오류, 드리프트 및 사용자 영향에 대한 계측기 모니터링.
확장하기 전에 롤백 및 사고 대응 경로를 준비하세요.
계속 탐색하세요
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the RAG Evaluation Metrics and RAGAS quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
자주 묻는 질문
What is RAG Evaluation Metrics and RAGAS?
RAG evaluation scores a retrieval-augmented generation system separately on retrieval and on generation, most commonly with four metrics: faithfulness, answer relevance, context precision and context recall. RAGAS is an open-source framework that computes these metrics, largely using an LLM as a judge. Separate scores matter because they show whether a bad answer came from fetching the wrong passages or from misusing the right ones.
Which metric checks whether every claim in the answer is supported by the retrieved context?
Faithfulness measures the fraction of answer claims backed by the context, which targets hallucination.
Which of these metrics requires a ground-truth reference answer for each test question?
Context recall checks whether the retrieved context contains what is needed for the reference answer, so a reference must exist.
What does context precision measure?
Context precision rewards putting useful chunks high in the ranking instead of burying them among irrelevant ones.
How does RAGAS estimate answer relevance?
If questions generated from the answer resemble the real question, the answer is on topic; embeddings measure that similarity.
A system shows high context recall but low context precision. What is the most likely fix?
The right chunks are being retrieved but ranked poorly, which is a ranking problem a reranker often addresses.
계속 학습하세요
관련 가이드
이 주제에 대해 선택된 추가 가이드