A result that improves without learning
A team compares a hundred AI model variants on the same small set of examples. It changes a prompt, configuration or training choice and keeps the best score. The winner looks convincing. Does that score measure model quality, or also the luck of being selected from many? This matters when choosing a business system: a comparison can contain no coding errors and still create misleading expectations.
To isolate the mechanism, imagine a hundred candidates that know nothing: each guesses a binary answer with 50% success probability. Evaluate them on twenty cases and select the highest count. We show that the winner averages about 77.32% on selection data while remaining at 50% on new data. We also calculate the chance of an apparently promising 75%. This is a probabilistic experiment with explicit assumptions, not a real-network benchmark.
Before the equations: what is selected
Accuracy is correct answers divided by evaluated cases. A candidate selected before seeing results would average 50% over repetitions. But we inspect all scores and choose afterwards. This favours positive measurement fluctuations by construction: lucky fluctuations enter the winning result while unlucky ones remain among discarded results. Nobody needs to falsify a number.
We call the twenty cases the selection set. “Validation set” is another name, but use matters more than naming: data influencing model preference participate in development. An independent test evaluates a choice already fixed. This also applies to prompts, thresholds, preprocessing recipes and example selection; it is not limited to neural-network weights.
The probability model and its assumptions
Let n = 20 be the number of cases and M = 100 the number of candidates. Xⱼ counts correct answers for candidate j. Each correctness indicator is Bernoulli: one with probability p = 0.5 and zero otherwise. Assume independence across cases and candidates, with candidates fixed before inspecting scores. These are strong assumptions: near-identical configurations of a real model often have correlated errors.
C(20,k) is the binomial coefficient, counting ways to place k successes in twenty positions. Each success/failure sequence has probability 1/2²⁰; multiplying by its count gives the probability of k successes. These quantities have no physical units; accuracy is a fraction or percentage. No learning algorithm is assumed because no candidate learns in this experiment.
A rare result becomes common when searching for the maximum
A preselected candidate reaches at least 75% by getting fifteen or more cases right. Sum counts from fifteen through twenty, not just exactly fifteen, to obtain the probability. The finite calculation gives α = 0.0206947, about 2.07%. That is uncommon for one candidate. But we ask whether at least one of a hundred succeeds, not whether a particular candidate does.
The second formula uses the complementary event: nobody reaches fifteen. One candidate has probability 1 − α; independence gives (1 − α)¹⁰⁰ for all hundred. Subtracting from one gives about 87.65%. Seeing at least one 75% is therefore typical here despite every true performance being chance-level. The rare event concerned a fixed candidate, not the search winner.
This is not the probability that a real model is useless after seeing its score. We calculated outcomes under a specific no-signal hypothesis. Reversing that conditioning requires further assumptions. The 50% reference also belongs to our fair binary game: imbalanced classes, other metrics or dependent examples require a null distribution appropriate to the problem.
How good does the best look on average?
Define Z = maxⱼ Xⱼ, the winning count. Its mean needs no simulation. A nonnegative integer can be written as a sum of indicators: one if it exceeds zero, another if it exceeds one, and so on. Taking expectations makes the mean a sum of threshold-crossing probabilities. Let F(k) be the probability that one candidate gets at most k answers right.
The result is 77.3197% for a hundred candidates. The formula also reveals the direction: raising F between zero and one to a larger power cannot increase it. Hence every term 1 − Fᴹ cannot decrease as candidates are added. The observed maximum improves by construction even though each success probability stays fixed. One candidate returns exactly 50%, a useful calculation check.
| Candidates M | Expected selected accuracy | P(maximum ≥ 75%) |
|---|---|---|
| 1 | 50.00% | 2.07% |
| 5 | 62.91% | 9.93% |
| 10 | 67.03% | 18.87% |
| 50 | 74.63% | 64.85% |
| 100 | 77.32% | 87.65% |

Why a fresh test returns to 50%
Let J be the selected candidate index and Aⱼ its accuracy on a fresh test. J depends only on selection results. By assumption, test outcomes are independent of those results and every candidate has expected accuracy 0.5. Knowing who won therefore gives no favourable information about new outcomes. Separate all possible winners and weight each by its selection probability.
Independence is the decisive step, not “test” in a filename. Using new results to change the winner, choose a threshold or decide which cases to show breaks this protocol. Testing the fixed winner once does not repeat the hundred-candidate competition on the new set. It estimates that candidate, but does not remove finite-sample uncertainty or an unrepresentative sample.
The executed simulation and what it measures
The script runs 5,000 repetitions with seed 20261004 in Python 3.14.0. Each generates one hundred counts over twenty trials, selects the maximum and breaks ties by smallest index. It then generates one hundred fresh correctness outcomes for the winner. Identical candidate distributions and independent testing mean no need to generate tests for the ninety-nine discarded candidates. No real labels, trained networks or AI-service calls are involved.
Mean selected accuracy is 77.286%; fresh-test accuracy is 50.0476%. They are close to the analytical 77.3197% and 50%. Residual differences are Monte Carlo variability from finite random repetitions. JSON stores each winning index and both counts, plus seed, version and theoretical results; code reconstructs candidate draws too. The seed was not selected to produce a more favourable result.
The script also reports standard errors of simulated means, about 0.064 and 0.070 percentage points. These equal the between-repetition standard deviation divided by √5000. They describe simulation precision under independent fair-coin assumptions, not uncertainty about a real product’s performance. A very precise simulation may precisely describe a probability model that does not represent the application.
Two options: account for search or separate evaluation
Judging the maximum on the same set requires including the search in its reference distribution. Independence gives our exact formula. Without independence across candidates, but with valid marginal event probabilities, the union bound says the chance that at least one exceeds a threshold is at most the sum of individual probabilities. This underlies Bonferroni control: simple, often conservative.
For an overall bound no greater than 5%, the first usable integer threshold here is 18 of 20, or 90%. The bound is about 2.0123%; exact independent probability is about 1.9923%. Discreteness partly explains why these are below 5%: we cannot require 17.4 correct answers. This does not validate every hundred-trial experiment; its binomial law and comparison family must match the actual protocol.
Alternatively reserve a final test for the fixed procedure. With scarce data, nested cross-validation can separate choice and evaluation: inner loops select variants; the outer loop evaluates the entire procedure on data excluded from selection. We do not execute that algorithm or invent its results here. The conceptual distinction is evaluating the complete selection process rather than a winner already optimised on all available data.
When independence does not hold
Consider the opposite extreme: all hundred candidates give identical answers. Counts are identical, so the maximum equals a single count. Expected accuracy is 50%, and the chance of fifteen successes stays 2.07%, not 87.65%. A hundred spreadsheet names are not a hundred independent attempts. Many dependence structures lie between these extremes; dividing candidate count by an intuitively chosen correlation is not enough.
Cases can be dependent too: neighbouring video frames or near-duplicate documents need not provide twenty distinct observations. Here n enters one candidate’s accuracy variance p(1 − p)/n. With n = 20, standard deviation is about 11.18 percentage points. Without independence, this formula need not describe actual variability. Merely increasing row count may create false precision.
How much do more examples help?
Repeat the exact sum changing only n while keeping M = 100. With one hundred cases per candidate, expected maximum falls to 62.4762%; with five hundred, to 55.6017%. True ability stays at 50%. More independent observations narrow each estimate’s fluctuations and reduce luck’s advantage. No special size eliminates it: optimism remains about 12.48 points at one hundred cases and 5.60 at five hundred. These are probability-model results, not universal test-size recommendations.
This comparison clarifies limited evaluation budgets. Searching more configurations and measuring each more precisely are different uses of resources. Searching creates opportunities for genuinely better candidates, if they exist, but also for selecting fluctuations; better measurement reduces uncertainty under suitable assumptions. Our example cannot estimate the first benefit because all candidates are equivalent. It quantifies why the observed maximum cannot be presented as though no search occurred.
In code, getrandbits(n) produces n pseudorandom bits and bit_count counts ones: a compact success-count generator, not a document model. cdf sums binomial coefficients; mean_max applies the maximum formula. Assertions check M = 1, increasing expected maxima and approximate simulation–theory agreement. Combinatorial counts are exact integers, while final probabilities and powers use floating point. Displayed figures are rounded and claim no experimental precision on real data.
Evidence, limitations and conclusion
Cawley and Talbot’s 2010 JMLR paper experimentally documents selection overfitting and bias in evaluation protocols. We read their methods and protocol comparisons. Our binomial example isolates a simpler mechanism and does not reproduce their classifiers.
The winner is not necessarily a bad model. Searching can find better systems when true candidate abilities differ; we made them equal to isolate luck. Independent evaluation also does not guarantee behaviour in every business or future setting. The precise conclusion is that a score used to win a search is not, by itself, an unbiased estimate of the winner’s performance on new cases.
Interpret a comparison by asking which data selected the model, how many attempts were considered and which data evaluated the choice without changing it. These questions connect statistics to a concrete decision. Our winner’s 77% was not fabricated: it was correctly measured, but answered the wrong question if presented as general ability. Understanding this separates successful configuration search from convincing evidence of value.
Bibliography and code
from experiment import run
r = run()
print(round(r['exact']['expected_selected_accuracy'], 6))
print(round(r['exact']['family_tail_15'], 6))
print(r['simulation'])
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.

