PRZEWODNIK techniczny

NDCG and Ranking Metrics

Ranking metrics evaluate the order of items rather than only whether individual labels are correct.

  • 3 minuty czytania
  • Ostatnia aktualizacja
Na tej stronie3 minuty czytania
  1. Przegląd
  2. Głębokie nurkowanie
  3. Wpływ strategiczny
  4. The Future of NDCG and Ranking Metrics
  5. Implementacja w świecie rzeczywistym
  6. Zagrożenia i poręcze
  7. Plan wdrożenia
  8. Odkrywaj dalej
  9. Często zadawane pytania

Przegląd

Precision@k, recall@k, MAP, MRR and normalized discounted cumulative gain emphasize different aspects, so metric choice should match the ranking task and its relevance judgments.

Głębokie nurkowanie

Ranking systems order candidates for a user or query. Evaluation should reflect where relevant results appear and how relevance is defined. Precision at k is the fraction of the top k results judged relevant. Recall at k is the fraction of all relevant candidates retrieved in the top k. Precision emphasizes the quality of the displayed set; recall emphasizes how much of the relevant set was found. Average precision summarizes precision at ranks where relevant items occur, and mean average precision (MAP) averages that value across queries. Mean reciprocal rank (MRR) uses the reciprocal of the rank of the first relevant result and averages across queries. MRR is useful when finding one good result quickly is the priority, but it largely ignores the quality of later ranks. Discounted cumulative gain (DCG) supports graded relevance. A gain function assigns larger value to more relevant items, while a logarithmic discount reduces credit for items lower in the list. Normalized DCG divides a ranking's DCG by the ideal DCG for that query, giving a score typically between zero and one when definitions align. This normalization makes values more comparable across queries with different relevance distributions, but aggregation choices still matter. Suppose a query has one highly relevant item and another mildly relevant item. Placing the highly relevant result first yields more DCG than placing it second. NDCG captures both graded relevance and rank position. It does not establish whether the relevance labels are unbiased or whether the candidate-generation process omitted useful items. Metrics can change with cutoff k, gain formula and label threshold. Report these choices and evaluate across the same query set. Offline metrics also do not fully capture user satisfaction, diversity, freshness, exposure bias or long-term outcomes. Use them alongside online experiments or human review when appropriate, while avoiding claims that a higher score alone proves a better user experience.

Wpływ strategiczny

Koszt i budżet

Decyzje dotyczące architektury wpływają na wydajność i koszty operacyjne przez lata.

Jaśniejsze decyzje

Edukacja techniczna pomaga zespołom wybrać odpowiedni stos, a nie tylko najnowszy.

Kontrola jakości

Lepsze wybory inżynieryjne zmniejszają liczbę incydentów związanych z niezawodnością w produkcji.

The Future of NDCG and Ranking Metrics

Ranking reports can improve by displaying top-k metrics, relevance definitions and per-query distributions rather than one aggregate number. Teams should choose MAP, MRR or NDCG based on whether the task values all relevant results, the first useful result or graded quality throughout the list. They should audit relevance judgments and candidate exposure, since metrics cannot reward items never retrieved for evaluation. Human satisfaction and diversity checks can complement offline scores. As ranking objectives evolve, preserve consistent historical definitions so trend comparisons remain meaningful.

Implementacja w świecie rzeczywistym

A search result list has five items and two are relevant. Precision@5 is 2/5, while recall@5 depends on how many relevant items exist in the full candidate set.

A user has three relevant items, and a system retrieves two within the top five. Recall@5 is 2/3 even though precision@5 is 2/5; the measures answer different questions.

For graded relevance, DCG rewards highly relevant items more when they appear near the top, using a gain and a logarithmic rank discount. NDCG divides by the ideal DCG for the same query to normalize the scale.

A recommendation team reports MRR when the first relevant result matters most and NDCG when multiple items and graded relevance across the list matter.

Zagrożenia i poręcze

  • Optymalizacja jednego testu porównawczego może ukryć szersze słabości systemu.

  • Koszty infrastruktury i utrzymania są często niedoszacowane.

  • W miarę jak systemy stają się coraz bardziej złożone, luki w bezpieczeństwie i obserwowalności mogą się zwiększać.

Plan wdrożenia

  1. Przed wdrożeniem zdefiniuj docelowe opóźnienia, jakość i koszty.

  2. Test porównawczy w realistycznych warunkach obciążenia i danych.

  3. Monitorowanie przyrządu pod kątem błędów, dryftu i wpływu użytkownika.

  4. Przed skalowaniem przygotuj ścieżki wycofywania zmian i reakcji na incydenty.

Odkrywaj dalej

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the NDCG and Ranking Metrics quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Rozpocznij quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Często zadawane pytania

What is NDCG and Ranking Metrics?

Ranking metrics evaluate the order of items rather than only whether individual labels are correct. Precision@k, recall@k, MAP, MRR and normalized discounted cumulative gain emphasize different aspects, so metric choice should match the ranking task and its relevance judgments.

Two relevant items appear in the top five. What is precision@5?

Precision@k divides relevant results in the top k by k, so it is 2/5.

Two of three relevant candidates appear in the top five. What is recall@5?

Recall divides retrieved relevant items by the total relevant set: 2/3.

Which metric emphasizes the rank of the first relevant result?

MRR averages the reciprocal rank of the first relevant result.

How does ideal DCG contribute to normalized DCG?

NDCG compares the observed order with an ideal ordering for the same query.

Which task best matches MRR?

MRR is driven by the first relevant rank and is useful when an early useful result matters.