概述
Inverse propensity scoring reweights outcomes by action probabilities, but reliable estimates require adequate policy overlap, correct propensities and careful handling of high variance and selective feedback.
深入探讨
A policy maps context, such as a user and current page, to an action such as which item to display. Off-policy evaluation (OPE) estimates a target policy's expected outcomes using logs generated by a different behavior or logging policy. It can help screen a candidate before deployment, especially in recommendation, advertising and contextual bandits. Inverse propensity scoring (IPS) weights each logged reward by the ratio between the target policy's probability of choosing that action and the logging policy's probability. For deterministic policies, only records where the logged action matches the target contribute, scaled by the inverse logging propensity. This corrects selection under assumptions, but small logging propensities create large weights and high variance. Clipping weights can stabilize estimates while introducing bias. Self-normalized IPS divides by the sum of weights, and doubly robust methods combine propensity and reward models, each with distinct assumptions. Support or overlap is crucial. If the logging policy never chose an action in a context, the log contains no direct evidence of its outcome there. No weighting formula can recover that missing counterfactual without additional assumptions. Propensities must be logged accurately, and the reward outcome must be observed. Selective labels, delayed outcomes and interference can also undermine estimates. Use OPE as evidence, not a guarantee. Report effective sample size, weight distribution, overlap diagnostics and sensitivity to estimators. A large estimate driven by a few extreme weights should be treated cautiously. Candidate policy changes that move far outside historical support may require randomized exploration or a controlled online experiment. OPE complements offline predictive metrics but does not replace experimentation for every decision. Evaluation assumptions should be documented with the logging policy version, context features, action set and reward definition.
战略影响
成本与预算
多年来,架构决策决定着性能和运营成本。
更清晰的判决
技术教育帮助团队选择正确的堆栈,而不仅仅是最新的堆栈。
质量控制
更好的工程选择可以减少生产中的可靠性事故。
The Future of Off-Policy Evaluation from Logged Data
OPE can help teams reduce risk by screening policies within the support of existing logs and identifying where evidence is too weak. Future data collection should record action propensities, exposure, context and outcomes with stable policy versions. Teams can combine OPE with small randomized experiments to improve support and validate estimates. Dashboards should show weight tails and effective sample size alongside policy value. When a candidate changes behavior substantially, the appropriate response may be to gather new evidence rather than extrapolate beyond the logs.
现实世界的实施
A recommender logged which item it showed, the context, click outcome and probability of selecting that item. An analyst uses those propensities to estimate a candidate policy's expected reward from the old log.
A target policy chooses actions that the logging policy almost never selected. The few matching outcomes receive large inverse weights, producing a noisy and potentially unstable estimate.
A team compares raw IPS with a self-normalized estimator and a doubly robust estimator, reporting assumptions and sensitivity rather than selecting whichever number is most favorable.
A policy evaluation lacks logged action probabilities. The analyst does not claim unbiased IPS results and instead collects better randomized logs or runs a controlled experiment.
风险与防护栏
优化一项基准测试可以隐藏更广泛的系统弱点。
基础设施和维护成本常常被低估。
随着系统变得更加复杂,安全性和可观察性差距可能会扩大。
实施路线图
在实施之前定义延迟、质量和成本目标。
在实际负载和数据条件下进行基准测试。
仪器监控错误、漂移和用户影响。
在扩展之前准备回滚和事件响应路径。
不断探索
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Off-Policy Evaluation from Logged Data quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
常见问题
What is Off-Policy Evaluation from Logged Data?
Off-policy evaluation estimates how a new decision policy might perform using data logged by an older policy, without deploying the new policy first. Inverse propensity scoring reweights outcomes by action probabilities, but reliable estimates require adequate policy overlap, correct propensities and careful handling of high variance and selective feedback.
What does off-policy evaluation estimate?
OPE uses logged behavior and outcomes to estimate performance under a different policy.
Why does IPS divide by the logging propensity?
Inverse propensity weighting compensates for the logging policy's action selection under assumptions.
What happens when the logging policy assigns a very small probability to a chosen action?
Small propensities lead to large inverse weights, making estimates unstable.
What does lack of support mean for a target action?
If the logger never took an action in a context, OPE cannot directly learn its outcome there from those logs.
Why report effective sample size with weighted estimates?
Weight concentration can make a large raw dataset behave like a much smaller effective sample.
继续学习
相关指南
为此主题精选的更多指南