A code-evaluation metric measuring the chance that at least one of k generated samples passes the tests.