개요
It reduces the optimistic evaluation that can arise when the same validation scores are used both to choose settings and to claim performance.
심층 분석
Searching many model settings is itself a form of learning from data. Even when every candidate uses cross-validation, choosing the candidate with the best score favors settings that benefited from random variation in those validation folds. Reporting that winning score as a fresh performance estimate can therefore be optimistic. Nested cross-validation separates selection from evaluation. Start with an outer split. Set aside its test fold and conduct the entire search using only the outer training data. Inner cross-validation compares the candidate settings. Refit the selected configuration on the outer training data, evaluate it once on the outer test fold, and repeat this process for the other outer folds. The outer scores evaluate a procedure: the model family, search space, preprocessing and selection rule used together. Different outer folds may select different settings. That is expected, because their training data differ. The result is not a competition in which you simply deploy whichever outer-fold model achieved the highest test score. Preprocessing belongs inside the procedure. A scaler, imputer or feature selector fitted on all examples before splitting can leak information. Scikit-learn's Pipeline helps keep these operations attached to model fitting, while GridSearchCV can perform the inner search. The split strategy must also match the intended use. Repeated records from one customer may need group separation. Forecasting requires respect for time order. Nesting an unsuitable random split does not repair that design problem. After evaluating the fixed procedure, perform its selection step on the available development data and fit a final model. Preserve a separate test set if the project uses one. Repeatedly changing the procedure after inspecting outer scores can turn the outer evaluation into another tuning process.
전략적 영향
비용 및 예산
아키텍처 결정은 수년 동안 성능과 운영 비용을 결정합니다.
더 명확한 결정들
기술 교육은 팀이 최신 스택뿐만 아니라 올바른 스택을 선택하는 데 도움이 됩니다.
품질 관리
더 나은 엔지니어링 선택은 생산 시 신뢰성 사고를 줄입니다.
The Future of Nested Cross-Validation
As automated model searches become easier to launch, evaluation systems should make the boundary between tuning and testing visible. A useful report records the whole selection recipe and the data it was allowed to inspect. Teams can then compare improvements to a frozen procedure instead of repeatedly optimizing against the same test feedback. Compute limits will still matter, especially for expensive models. The appropriate response is to choose an evaluation design that fits the project and document its limitations, rather than treating extra layers of cross-validation as a universal requirement.
실제 구현
A researcher compares several regularization strengths inside each outer training split. Only the selected model is evaluated on that split's outer test fold.
A hypothetical search uses five outer folds, three inner folds and ten candidate settings. It performs 150 inner candidate fits, plus five refits if the selected model is refitted once per outer fold.
A model uses multiple records from each customer. The team keeps customers separated across both inner and outer splits when it wants to estimate performance on unseen customers.
An analyst places a scikit-learn Pipeline inside GridSearchCV and evaluates that search with an outer cross-validation routine. Preprocessing is learned within the relevant training folds.
위험 및 가드레일
하나의 벤치마크를 최적화하면 더 광범위한 시스템 약점을 숨길 수 있습니다.
인프라 및 유지 관리 비용은 종종 과소평가됩니다.
시스템이 더욱 복잡해짐에 따라 보안 및 관찰 가능성의 격차가 커질 수 있습니다.
구현 로드맵
구현하기 전에 지연 시간, 품질, 비용 목표를 정의하세요.
현실적인 로드 및 데이터 조건에서 벤치마킹합니다.
오류, 드리프트 및 사용자 영향에 대한 계측기 모니터링.
확장하기 전에 롤백 및 사고 대응 경로를 준비하세요.
계속 탐색하세요
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Nested Cross-Validation quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
자주 묻는 질문
What is Nested Cross-Validation?
Nested cross-validation uses an inner loop to select a model's settings and an outer loop to evaluate the resulting selection procedure on held-out data. It reduces the optimistic evaluation that can arise when the same validation scores are used both to choose settings and to claim performance.
Within one outer split, which data may the inner hyperparameter search use?
The outer test fold must remain outside the selection process so it can evaluate the procedure on held-out data.
Why can reporting the best inner validation score overstate model performance?
Selection uses the validation results, so the winning score is not an independent assessment of that selection process.
Five outer folds, three inner folds and ten candidate settings require how many inner candidate fits in the guide's example?
Multiply the outer folds, inner folds and candidates: five times three times ten equals 150.
A scaler is fitted on the full dataset before nested cross-validation begins. Which problem remains?
Nesting does not undo information leakage introduced before splitting. Preprocessing must be fitted within the relevant training folds.
Different outer folds select different hyperparameters. How should this be interpreted?
The outer loop assesses the selection procedure, whose chosen settings may vary with its training sample.
계속 학습하세요
관련 가이드
이 주제에 대해 선택된 추가 가이드