ELAI S.r.l.

Do a thousand images from one machine count as a thousand independent tests?

A perfect test can measure memory rather than generalization. A reproducible experiment explains when to split machines before images.

Do a thousand images from one machine count as a thousand independent tests?

The hidden question inside an excellent result

An inspection system correctly classifies almost every image held out from training. It seems ready for a new production line. Yet test images come from familiar machines, lighting, backgrounds and acquisition geometry. Did we measure defect recognition or recognition of context? The answer depends on the test design before it depends on model architecture.

Abstract. A controlled counterexample uses a classifier that remembers machine identity only: it scores 100% on randomly split observations and 50% on unseen machines. This is not a vision benchmark; no images are processed. The deliberately elementary model isolates the statistical question. We derive overlap probability, interpret confusion matrices and distinguish known machines, new machines and future times. The reader learns which promise a score can actually support.

Before testing: what must be new?

Training uses labeled examples to build a rule; testing applies it to excluded examples. A different file row need not represent a new situation. Twenty frames from one acquisition share conditions that two machines may not share. A group is the collection tied to the unit of interest, here a machine. It is not a class: a class is the target, such as 0 or 1, and a group can generally contain both.

If deployment serves the same hundred installed machines, knowing some characteristics may be legitimate. If it targets an unseen factory, the test must represent that unfamiliarity. These are different objectives, not a moral ranking of procedures. Accuracy without the held-out unit hides this distinction. Even a machine split may be insufficient when every machine belongs to one factory but the promise concerns new factories.

A model that knows nothing about defects

Generate one hundred IDs and randomly assign fifty zero and fifty one labels. Each machine produces twenty rows with the same label, totaling two thousand. Within-machine label constancy is deliberately extreme to expose the shortcut, not a claim about real defects. Each row stores machine, frame index and label; there are no pixels or part features.

The classifier builds an ID-to-label dictionary. Unseen IDs receive the training majority class, with zero on ties. Construction takes one pass; prediction looks up a key without learning any transferable relation between appearance and defect. Hash lookup has average constant cost; storage and construction grow with groups and rows. This is a diagnostic baseline, not a proposed production architecture.

Why random row splitting makes familiarity almost certain

Assign each row independently to training with probability p = 0.8. Consider a row known to be in test. Its m-row group leaves m − 1 opportunities to see that machine in training. To never see it, all remaining rows must enter test. Multiply their probabilities because assignments are independent, then subtract from one for the opposite event.

P(seen machine | test row) = 1 − (1 − p)^(m − 1)
P(unseen) = 0.2^19 = 5.24288 × 10^−14

One row per machine gives zero familiarity; two give 80%; five give 99.84%; twenty give almost one. These describe the splitting procedure, not prediction correctness. For our dictionary, however, knowing the group reveals the answer. More frames make the test more familiar without adding knowledge about a new machine. The formula assumes independent Bernoulli assignments; fixing the exact training row count requires a different combinatorial calculation.

Two protocols, two executed results

The Python experiment with seed 20260929 assigns 1,577 training and 423 test rows. All hundred machines occur in both sets. All 423 predictions are correct. The second protocol holds out ten zero-labeled and ten one-labeled machines entirely. The remaining eighty supply 1,600 training rows. None of the twenty test IDs is known, so all 400 predictions use fallback zero and only 200 are correct.

ProtocolTrain/test rowsShared groupsAccuracy
row Bernoulli p=0.81577 / 423100100%
held-out groups1600 / 400050%

The second confusion matrix says more than 50% alone. Rows are true classes zero and one; columns are predictions zero and one. [[200, 0], [200, 0]] means all zeros and no ones are recognized. This is not moderately good recognition, but constant prediction. The first matrix is [[207, 0], [0, 216]]. Both are saved in JSON with labels, held-out groups and row indices, making the result reconstructible.

Left: theoretical probability that a test row belongs to an already-seen machine. Right: executed accuracies. The 100%–50% difference concerns the synthetic dictionary, not a real vision model.
Left: theoretical probability that a test row belongs to an already-seen machine. Right: executed accuracies. The 100%–50% difference concerns the synthetic dictionary, not a real vision model.

What the counterexample proves—and its limits

We established that separated rows do not generally establish generalization across groups. We did not establish that every neural network memorizes machines or that 50% is typical industrial performance. A model learning defect features may succeed on new groups. The counterexample refutes this deduction: no test row was in training, therefore new-machine deployment was validated. One executed case with true premise and false conclusion suffices.

Balanced group-test labels deliberately isolate memory. Real plants may have different class frequencies and group sizes. Frame-average accuracy weights machines producing more images more heavily; averaging machine accuracies weights machines equally. Neither is universally correct: match the decision. Our twenty equal-sized rows per group make the averages coincide, a property that need not hold in real data.

From demonstration to test design

For new-machine testing, choose held-out machines first, then construct training and test data. Transformations of one original stay on one side; cropping or augmentation before splitting can scatter near copies across both. Learn normalization and feature selection from training only, repeating fitting within each model-selection partition. Operation order is part of the method, not administrative housekeeping.

Scikit-learn documents GroupKFold for separated groups and StratifiedGroupKFold for also approximately preserving class proportions. They do not decide the meaningful grouping unit. Our experiment directly implements one balanced group holdout, without executing scikit-learn or cross-validation. Cross-validation repeats fitting and evaluation over partitions to help choose model settings; final test groups must remain outside that choice.

Time remains another dimension. Tomorrow a familiar machine may have worn tooling, a replaced camera or new production conditions. Random historical frames need not represent these transitions. Testing tomorrow on current machines calls for temporal ordering; tomorrow in a new factory calls for both differences. No universal split makes operational context irrelevant.

Finally, four hundred frames from twenty machines are not four hundred independent demonstrations of transfer. Uncertainty across machines requires preserving groups in resampling, and twenty units may miss rare variants. We compute no confidence intervals or population performance estimates: this is a counterexample defined by saved data and protocol. Repeating it verifies calculation without expanding the observed population.

The answer: count new situations, not only rows

A thousand images may describe one machine well without becoming a thousand tests on new machines. The dictionary reached 100% without learning defects: remember that when reading a benchmark. Useful evaluation states what remains familiar and what changes, splits accordingly and interprets scores within that scope. More images alone do not resolve the question; aligning the test with the promise does.

Sources and reproducibility

scikit-learn — Cross-validation, sections on grouped data, stratification and held-out transformations; documentation consulted 29 September 2026.

The snippet reduces the dictionary to four machines to expose its mechanism. The archive contains the hundred-machine experiment, seed, assignments and reported results, executed using Python’s standard library. These are neither EL-AI product tests nor industrial image results. Documentation describes available splitters; the dataset, probability derivation and counterexample are this article’s educational analysis.

def predict(train, test):
    memory = {group: label for group, label in train}
    return [memory.get(group, 0) for group, _ in test]
train = [('A', 0), ('B', 1)]
for test in [[('A', 0), ('B', 1)], [('C', 0), ('D', 1)]]:
    predictions = predict(train, test)
    accuracy = sum(p == y for p, (_, y) in zip(predictions, test))/len(test)
    print(predictions, accuracy)

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.