ELAI S.r.l.

Is a missing measurement really neutral? What AI learns from data collection

A wholly synthetic biomedical example explains selection, imputation and missingness signals: when they help and why changing collection makes them fragile.

Is a missing measurement really neutral? What AI learns from data collection

The problem: an empty cell also records a choice

In a biomedical dataset, some rows contain a measurement and others an empty cell. One response is to fill blanks with a mean and continue. But why is the measurement missing? It may never have been requested, arrived late or been lost during transfer. Different groups may experience these situations. A model can therefore learn the process deciding who gets measured, even when measurement values add no information.

The central question is whether retaining “present or absent” improves prediction, and what happens when collection changes. We must separate predicting an outcome, reconstructing an unseen value and describing a population. We follow a synthetic thousand-case table with explicit counts. These are not real patients, a diagnostic test evaluation or clinical decisions. The example isolates a statistical mechanism for AI design and oversight.

Notation and a deliberately simple example

Y is one for an abstract outcome and zero otherwise. In the retrospective table Y is known for all thousand cases; naturally it is unavailable as input to a future prediction. X is a measurement in arbitrary units. R is one when X is available at prediction time and zero otherwise. R is the observation indicator; the missingness indicator is 1 − R. Stating the convention prevents swapping presence and absence in equations.

To isolate collection effects, X is always 7 in our synthetic world, including unseen entries. This explicit extreme choice means values cannot distinguish outcomes. Two hundred cases have Y = 1 and eight hundred Y = 0: overall frequency 20%. X is available for 80% of Y = 1 cases and 10% of Y = 0 cases. This describes association, not an operator knowing future outcomes; correlated signals or circumstances could produce such selection.

OutcomeR = 1: presentR = 0: absentTotal
Y = 116040200
Y = 080720800
Σ2407601000

Selection changes the denominator

Complete rows contain 160 positive outcomes out of 240, about 66.67%. Rows without measurements contain 40 out of 760, about 5.26%. Both frequencies are correct for their groups, but neither describes the whole population. Dropping incomplete rows and reporting 66.67% overall would confuse selection with the total. This is often the first missing-data problem, before any imputation algorithm.

π = P(Y=1), a = P(R=1|Y=1), b = P(R=1|Y=0) P(Y=1|R=1) = πa / [πa + (1−π)b] P(Y=1|R=0) = π(1−a) / [π(1−a) + (1−π)(1−b)]

The vertical bar means “given”: conditioning restricts the group counted. The first numerator πa is the fraction with outcome one and measurement present; its denominator includes all present cases, including outcome zero. Substituting π = 0.2, a = 0.8 and b = 0.1 gives 0.16/0.24 = 2/3. For absent cases it gives 0.04/0.76 = 1/19. Information is not in the number 7, but in row selection.

Accurate filling is not the same as accurate prediction

Imputation assigns a replacement to a missing measurement. Observed mean is 7; filling with 7 perfectly reconstructs every X in this experiment, whose construction we know. Yet deleting R makes every row indistinguishable. A rule using only X can give everyone probability 0.2. Retaining R allows 2/3 for present and 1/19 for absent cases. X reconstruction is perfect either way, but information for predicting Y differs.

This proves neither that mean imputation always fails nor that an explicit indicator is always necessary: a replacement distinct from real values may already reveal absence. The narrower result is that filling which merges informationally distinct states can lose predictive signal. It does not contradict general imputation theory: constant X, for example, violates the continuous-density assumption in the paper cited below.

Measuring prediction with an interpretable error

To compare probabilities, use binary Brier score: mean squared difference between predicted probability p and outcome Y. Predicting 0.8 contributes 0.04 when Y = 1 and 0.64 when Y = 0, penalising confident mistakes. It is dimensionless and lower is better. It is neither a diagnostic error percentage nor clinical benefit.

BS = (1/N) Σᵢ (pᵢ − Yᵢ)² E[(Y−p)² | R] = q_R(1−q_R) + (p−q_R)², q_R = P(Y=1|R)

The second identity separates remaining group uncertainty q_R(1−q_R) from the penalty for predicting a different probability, (p−q_R)². Expand q_R(1−p)² + (1−q_R)p² to verify it; minimum occurs at p = q_R. In the original table constant 0.2 gives 0.16, while the R rule gives about 0.091228. These are exact table values computed with fractions before rounding, not a trained classifier’s test estimate of generalisation.

Changing the protocol can change what absence means

Keep the same thousand cases, two hundred positive outcomes and X = 7, but make measurements available for half of each outcome group. Present cases contain 100 positives and 400 negatives, as do absent cases: both have frequency 20%. R no longer carries signal even though overall outcome frequency is unchanged. Freeze the old rule, as an unupdated system would: it still predicts 2/3 and 1/19.

BS_nuovo = 0.5 [0.2(1−2/3)² + 0.8(2/3)²] + 0.5 [0.2(1−1/19)² + 0.8(1/19)²] ≈ 0.279748

The 0.5 coefficients represent equally sized groups. Inside each bracket, positive errors receive weight 0.2 and negative errors 0.8. The old rule reaches about 0.279748, worse than constant prediction’s unchanged 0.16. Recomputing conditional probabilities in the new world gives 0.2 for both, restoring Brier 0.16 for the R rule too. Retaining the indicator is not the problem; treating its outcome association as stable is.

Two synthetic tables, no clinical data. Left: outcome frequencies for present and absent measurements; the gap disappears when collection changes. Right: Brier scores for constant probability and the unchanged R rule. Lower is better; bars are neither confidence intervals nor health-benefit measurements.
Two synthetic tables, no clinical data. Left: outcome frequencies for present and absent measurements; the gap disappears when collection changes. Right: Brier scores for constant probability and the unchanged R rule. Lower is better; bars are neither confidence intervals nor health-benefit measurements.

The decomposition quantifies the missed update: new q_R = 0.2 in both groups, so additional penalty is 0.5(2/3 − 0.2)² + 0.5(1/19 − 0.2)², about 0.119748, exactly the gap between 0.279748 and 0.16. Neither population nor true measurement values need change; observation alone suffices. This constructs a possibility, not an estimate of its frequency in real services.

Do not confuse information, causality and timing

R predicting Y does not mean requesting a measurement causes the outcome. These are conditional associations. Making measurements available to more people can change R without changing Y, as in the second example. Giving the rule biological meaning would be unjustified. Real projects need collection-process knowledge: request, execution, result arrival and ingestion are different events at different times.

“Informative missingness” is not automatically Missing Not At Random, MNAR. That classification concerns which observed or unobserved variables drive absence. Here retrospective Y is observed for everyone and the mechanism depends on Y: for missing X, dependence is on an observed variable. Y is unavailable at prediction time. Mechanism labels therefore require an information set and task; an acronym alone cannot justify a strategy.

If R uses measurements arriving after the intended decision time, it leaks future information. Retrospective performance could be excellent yet impossible in deployment. Here R is defined at prediction time; delays are not simulated. Real data require timestamps and availability rules, not an assumption based on a column existing in the final database.

Estimating the population is another problem

Suppose we want overall 20% from measured cases alone and know exact selection probabilities a = 0.8 and b = 0.1. Inverse-probability weighting makes 160 positives represent 160/0.8 = 200 cases and 80 negatives represent 80/0.1 = 800. Weighted frequency is 200/(200 + 800) = 0.2. This explains selection correction, not a Y-prediction rule: these weights use the outcome unavailable in future cases.

Real selection probabilities may be unknown; misspecifying them adds error. They must be positive in represented groups: if a group is never observed, infinite weights do not create data. Very small probabilities produce large, unstable weights. Our complete table already gives the total, so weighting was unnecessary; it illustrates population composition versus individual prediction.

What observed values alone cannot reveal

Now change both question and assumptions, explicitly leaving the world where all X were known to be 7. Retain only that 24% are observed and equal 7; unseen values are merely bounded between 0 and 10. With μ₀ the missing-group mean, total mean is 0.24 × 7 + 0.76 × μ₀. Observed data do not determine μ₀: worlds with all unseen values 1 or all 9 produce the same observed archive.

μ = 0.24 × 7 + 0.76 × μ₀ 0 ≤ μ₀ ≤ 10 ⇒ 1.68 ≤ μ ≤ 9.28

The two population means are 2.44 and 8.52. The interval 1.68–9.28 follows only from value bounds; it is not a 95% confidence interval. Narrowing it requires information or assumptions, such as additional measurements or justified relationships to observed variables. Filling with 7 chooses an unseen-data assumption rather than proving it. Multiple imputations do not automatically remove ambiguity; they reflect their model and assumptions.

Reproducibility, research and limits

The download contains both tables, formulas and results. Python standard-library Fraction avoids intermediate rounding in identities; conversion to decimals is only for storage and plots. There is no random sampling or seed. The snippet prints conditional probabilities, Brier scores and mean bounds for the separate problem. Assertions check 2/3, 1/19, lost advantage and weighted frequency recovery. These are not tests on clinical records or trained models.

Le Morvan, Josse, Scornet and Varoquaux, NeurIPS 2021, distinguish imputation quality from prediction. They study consistency assumptions and compare synthetic-regression procedures including missingness information. We reproduce neither NeuMiss nor their experiments; our discrete table asks a different, narrower question.

A real model needs temporally valid data, a documented protocol and test groups covering plausible collection or site changes. Compare imputed values, values plus indicators and native missing-data methods with other factors fixed. Fit means and transformations on training data, without adapting to the test. Evaluate important missingness patterns separately: good aggregate performance can hide a small problematic group. This is a proposed validation design, not executed here.

An operational limit remains: monitoring missingness rates can flag change but cannot automatically reveal a changed relationship with Y. Reliable outcomes are needed, often arriving later. More complete collection may benefit a service while making a model based on old selection obsolete. Data and process analysis must therefore proceed together, without treating historical correlation as a permanent rule.

Conclusion: preserve the data’s context

An empty cell need not be neutral. Here availability improved prediction by revealing selection; after a protocol change the same rule became worse than constant probability. This justifies neither ignoring absence nor always exploiting it. Preserve its meaning and timing, separate imputation from prediction, test collection stability and state what observations cannot reveal. AI quality depends on this knowledge as well as model sophistication.

References and reproducibility

Le Morvan, Josse, Scornet & Varoquaux — What’s a good imputation to predict with missing values? NeurIPS 2021.

from experiment import run
r = run()
for row in r['scenarios']:
    print(row['scenario'], row['conditional_p'])
    print(round(row['brier_prior'], 6),
          round(row['brier_frozen_mask'], 6))
print(r['separate_nonidentifiability'])

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.