기술 가이드

CUPED Variance Reduction

CUPED uses pre-experiment measurements correlated with an experiment outcome to reduce variance in treatment-effect estimates.

  • 3분 읽기
  • 마지막 업데이트
이 페이지에서3분 읽기
  1. 개요
  2. 심층 분석
  3. 전략적 영향
  4. The Future of CUPED Variance Reduction
  5. 실제 구현
  6. 위험 및 가드레일
  7. 구현 로드맵
  8. 계속 탐색하세요
  9. 자주 묻는 질문

개요

It adjusts each observed outcome using a pre-period covariate without changing randomized treatment assignment, but the covariate must be chosen before treatment and the analysis must preserve valid inference.

심층 분석

CUPED, or controlled-experiment using pre-experiment data, is a variance-reduction method for randomized experiments. It uses a pre-treatment variable that predicts the outcome to remove some predictable variation. Because the covariate is measured before treatment assignment takes effect, adjusting for it can improve precision without changing the randomization. This is especially useful when user behavior is correlated over time, such as prior activity predicting activity during an experiment. Let Y be the experiment-period outcome and X a pre-period covariate. A common adjustment is Y_adj = Y - theta*(X - mean(X)), where theta is often estimated as Cov(X,Y)/Var(X). If X predicts Y, the adjustment subtracts expected baseline variation while preserving the outcome's mean scale. In a simple linear setting, variance reduction is related to the squared correlation between X and Y; weakly predictive covariates provide little gain. CUPED does not create a causal effect by itself. Causal interpretation still relies on valid random assignment, correct analysis units and measurement. Covariates should be selected before treatment or based on pre-treatment data, not chosen because they make the observed result look favorable. A post-treatment variable can be affected by treatment and conditioning on it can bias the estimate. Missing baseline data, extreme values or different availability across groups require care. Use the same adjustment procedure for treatment and control groups, define how theta is estimated and report uncertainty with the adjusted outcome. Cross-fitting or pooled estimation may be appropriate depending on design. Evaluate the expected variance reduction on historical data, but do not overstate gains before the experiment. CUPED can reduce sample size or duration for a fixed power target when assumptions hold, yet the achieved precision depends on actual correlation, measurement quality and analysis plan. It complements randomization and good experiment design rather than replacing them.

전략적 영향

비용 및 예산

아키텍처 결정은 수년 동안 성능과 운영 비용을 결정합니다.

더 명확한 결정들

기술 교육은 팀이 최신 스택뿐만 아니라 올바른 스택을 선택하는 데 도움이 됩니다.

품질 관리

더 나은 엔지니어링 선택은 생산 시 신뢰성 사고를 줄입니다.

The Future of CUPED Variance Reduction

CUPED can help experiments reach useful precision sooner when stable pre-period measurements predict the outcome. Teams should select covariates before launch, verify coverage and correlation using historical periods, and report the adjustment method alongside unadjusted results. Privacy and data freshness matter when reusing user histories. Automated analysis systems can provide adjusted and raw estimates together, making the effect of variance reduction transparent. CUPED remains one part of a well-designed experiment; randomization, outcome definition and stopping rules still determine whether the result is credible.

실제 구현

A hypothetical experiment measures each user's activity before and during treatment. If pre-period activity predicts the outcome, CUPED adjusts for that baseline and can make the treatment comparison more precise.

An analyst estimates theta as covariance between pre-period metric X and outcome Y divided by variance of X, then computes Y_adjusted=Y-theta*(X-mean(X)). Centering keeps the population-scale interpretation while subtracting predictable variation.

A product team uses a pre-treatment purchase count as the covariate but avoids using a value affected by treatment, which could introduce post-treatment bias.

A team checks pre-period coverage and whether missing baseline measurements differ by experiment group before applying CUPED, rather than assuming every user has a usable covariate.

위험 및 가드레일

  • 하나의 벤치마크를 최적화하면 더 광범위한 시스템 약점을 숨길 수 있습니다.

  • 인프라 및 유지 관리 비용은 종종 과소평가됩니다.

  • 시스템이 더욱 복잡해짐에 따라 보안 및 관찰 가능성의 격차가 커질 수 있습니다.

구현 로드맵

  1. 구현하기 전에 지연 시간, 품질, 비용 목표를 정의하세요.

  2. 현실적인 로드 및 데이터 조건에서 벤치마킹합니다.

  3. 오류, 드리프트 및 사용자 영향에 대한 계측기 모니터링.

  4. 확장하기 전에 롤백 및 사고 대응 경로를 준비하세요.

계속 탐색하세요

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the CUPED Variance Reduction quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

퀴즈 시작

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

자주 묻는 질문

What is CUPED Variance Reduction?

CUPED uses pre-experiment measurements correlated with an experiment outcome to reduce variance in treatment-effect estimates. It adjusts each observed outcome using a pre-period covariate without changing randomized treatment assignment, but the covariate must be chosen before treatment and the analysis must preserve valid inference.

Which variable is appropriate for a standard CUPED adjustment?

CUPED uses a pre-experiment covariate to reduce outcome variance without conditioning on treatment effects.

What does theta=Cov(X,Y)/Var(X) represent in a common CUPED adjustment?

Theta is the variance-minimizing linear projection coefficient of Y on X under the stated setup.

What happens when the pre-period covariate has little correlation with the outcome?

Weak predictive relationship means little outcome variation can be removed by the adjustment.

Why avoid using a treatment-affected covariate?

A post-treatment variable can lie on the treatment pathway or be otherwise affected by assignment, biasing comparisons.

What does centering X by its mean do in Y_adj=Y-theta(X-mean X)?

Centering makes the adjustment subtract deviations around the mean rather than shifting the outcome by an arbitrary level.