Paper gives exact limits for validating AI systems that select the best of many outputs
A new arXiv paper argues that repeating similar tests can reduce sampling noise without revealing blind spots in AI validation. It derives exact limits for best-of-N search and proposes combining structural coverage with independent tasks.