概述
It reduces the optimistic evaluation that can arise when the same validation scores are used both to choose settings and to claim performance.
深入探讨
Searching many model settings is itself a form of learning from data. Even when every candidate uses cross-validation, choosing the candidate with the best score favors settings that benefited from random variation in those validation folds. Reporting that winning score as a fresh performance estimate can therefore be optimistic. Nested cross-validation separates selection from evaluation. Start with an outer split. Set aside its test fold and conduct the entire search using only the outer training data. Inner cross-validation compares the candidate settings. Refit the selected configuration on the outer training data, evaluate it once on the outer test fold, and repeat this process for the other outer folds. The outer scores evaluate a procedure: the model family, search space, preprocessing and selection rule used together. Different outer folds may select different settings. That is expected, because their training data differ. The result is not a competition in which you simply deploy whichever outer-fold model achieved the highest test score. Preprocessing belongs inside the procedure. A scaler, imputer or feature selector fitted on all examples before splitting can leak information. Scikit-learn's Pipeline helps keep these operations attached to model fitting, while GridSearchCV can perform the inner search. The split strategy must also match the intended use. Repeated records from one customer may need group separation. Forecasting requires respect for time order. Nesting an unsuitable random split does not repair that design problem. After evaluating the fixed procedure, perform its selection step on the available development data and fit a final model. Preserve a separate test set if the project uses one. Repeatedly changing the procedure after inspecting outer scores can turn the outer evaluation into another tuning process.
战略影响
成本与预算
多年来,架构决策决定着性能和运营成本。
更清晰的判决
技术教育帮助团队选择正确的堆栈,而不仅仅是最新的堆栈。
质量控制
更好的工程选择可以减少生产中的可靠性事故。
The Future of Nested Cross-Validation
As automated model searches become easier to launch, evaluation systems should make the boundary between tuning and testing visible. A useful report records the whole selection recipe and the data it was allowed to inspect. Teams can then compare improvements to a frozen procedure instead of repeatedly optimizing against the same test feedback. Compute limits will still matter, especially for expensive models. The appropriate response is to choose an evaluation design that fits the project and document its limitations, rather than treating extra layers of cross-validation as a universal requirement.
现实世界的实施
A researcher compares several regularization strengths inside each outer training split. Only the selected model is evaluated on that split's outer test fold.
A hypothetical search uses five outer folds, three inner folds and ten candidate settings. It performs 150 inner candidate fits, plus five refits if the selected model is refitted once per outer fold.
A model uses multiple records from each customer. The team keeps customers separated across both inner and outer splits when it wants to estimate performance on unseen customers.
An analyst places a scikit-learn Pipeline inside GridSearchCV and evaluates that search with an outer cross-validation routine. Preprocessing is learned within the relevant training folds.
风险与防护栏
优化一项基准测试可以隐藏更广泛的系统弱点。
基础设施和维护成本常常被低估。
随着系统变得更加复杂,安全性和可观察性差距可能会扩大。
实施路线图
在实施之前定义延迟、质量和成本目标。
在实际负载和数据条件下进行基准测试。
仪器监控错误、漂移和用户影响。
在扩展之前准备回滚和事件响应路径。
不断探索
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Nested Cross-Validation quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
常见问题
What is Nested Cross-Validation?
Nested cross-validation uses an inner loop to select a model's settings and an outer loop to evaluate the resulting selection procedure on held-out data. It reduces the optimistic evaluation that can arise when the same validation scores are used both to choose settings and to claim performance.
Within one outer split, which data may the inner hyperparameter search use?
The outer test fold must remain outside the selection process so it can evaluate the procedure on held-out data.
Why can reporting the best inner validation score overstate model performance?
Selection uses the validation results, so the winning score is not an independent assessment of that selection process.
Five outer folds, three inner folds and ten candidate settings require how many inner candidate fits in the guide's example?
Multiply the outer folds, inner folds and candidates: five times three times ten equals 150.
A scaler is fitted on the full dataset before nested cross-validation begins. Which problem remains?
Nesting does not undo information leakage introduced before splitting. Preprocessing must be fitted within the relevant training folds.
Different outer folds select different hyperparameters. How should this be interpreted?
The outer loop assesses the selection procedure, whose chosen settings may vary with its training sample.
继续学习
相关指南
为此主题精选的更多指南