The question the average cannot answer
An optical sensor inspects a component and an AI model must produce a result within 10 milliseconds of acquisition. A test report gives average latency of 8 ms. Can we say the device is on time? Not yet: the mean distributes total time across trials, but does not say how many answers miss the deadline. A correct classification may be useless if the component has passed the useful action point.
Abstract. We construct two thousand-entry latency datasets with equal means, calculate percentiles and misses, then add other device stages. We derive what zero misses in a finite sample means. The path leads from the practical question to binomial probability without assuming statistical expertise. All data are explicitly constructed synthetic values, not measurements of a board, accelerator or EL-AI product.
First define which clock we are reading
Latency is time between chosen events. “Function call to return” differs from “sensor sample to result available to the actuator.” A deadline is when the result is needed. Inference time may exclude sensor reading, preparation, memory copies and waiting. Set D=10 ms and define a miss as L>D; a result exactly at 10 ms is on time. State that convention explicitly.
First isolate inference only. We do not simulate queues: trials are separated enough not to overlap. No neural model runs. Dataset A contains 1000 entries of 8 ms; B contains 990 entries of 7.8 ms and ten of 27.8 ms. We store integer microseconds, so 7.8 ms is 7800, avoiding decimal-rounding effects on counts.
Same mean, different answers
The mean sums times and divides by trial count. In B, many fast cases exactly compensate for a few slow ones. The formula is correct; it answers a different question from punctuality.
A has no misses in the trace; B has ten, or 1%. B is faster in most cases yet misses deadlines in a minority hidden by the mean. We do not know whether A would remain regular on hardware: we constructed it that way. This proves a mathematical possibility, not architectural superiority. Equal classification accuracy would not remove the timing difference.
A percentile is a threshold, not the worst case
For the 99th percentile, sort the thousand times and take position 990. We use nearest rank: mathematical index ceil(q n), with q between zero and one. Code subtracts one because lists start at zero. At least 99% of values are at or below this threshold. The remaining 1% need not resemble it. Some libraries interpolate between positions; comparisons need the convention.
| Trace | Mean ms | p99 ms | p99.9 ms | Observed max ms | Misses / 1000 |
|---|---|---|---|---|---|
| A | 8.0 | 8.0 | 8.0 | 8.0 | 0 |
| B | 8.0 | 7.8 | 27.8 | 27.8 | 10 |
B even has a better p99 than A despite ten misses. At p99.9 we enter the slow tail and reach 27.8 ms. “P99 under 10 ms” can be sensible if some misses are acceptable; it does not mean every result is under 10 ms. The observed maximum of 27.8 ms is not a proven bound for every future input and device state.

In the left panel B stays at 1% between 8 and 27.8 ms: raising the threshold from 10 to 20 ms does not rescue those ten cases. It reaches zero at 27.8 ms because we count L>D. A reaches zero at 8 ms. The whole curve avoids reducing the decision to one percentile. The right panel follows after distinguishing sample and population.
The device is more than the model
Add a constant 1.5 ms for acquisition, preparation and result delivery, with no overlap. This is an educational assumption, not a measurement. Total latency is Ltot=Linf+1.5 ms. Both means become 9.5 ms; A remains at 9.5, while B has 990 results at 9.3 and ten at 29.3. Miss rate remains 1%, but normal-case margin changes. Timing only the neural kernel would miss the other stages.
Addition is simple because the extra cost is constant. Generally, adding two stage p99 values does not give system p99: the sum distribution depends on which latencies occur together. Measure stages on the same request with identifiers, or use valid per-stage bounds. Overlapping pipelines and asynchronous accelerators also require separating submission from actual result completion.
Zero misses does not mean zero probability
Now imagine n real independent trials under unchanged conditions with no misses. This statistical hypothesis does not turn our synthetic traces into measurements. Let p be the unknown, constant miss probability per trial. No miss in one trial has probability 1−p; for n independent trials probabilities multiply. A rare event can therefore escape a short campaign.
K counts misses, n counts trials and α is the chosen probability threshold. For a one-sided 95% upper bound use α=0.05. We seek pU for which zero misses has 5% probability. With n=1000, pU≈0.002991, about 0.299%. Zero in a thousand therefore does not justify “below one in a thousand with 95% confidence”: the bound is still about three in a thousand.
Confidence describes procedure coverage over repeated campaigns under the model; it does not assign subjective probability to a fixed parameter or guarantee the next run. The right panel shows the bound falling with n. This needs more informative observations, not copies of the same case. If load and temperature create clustered misses, independence is doubtful and this formula is no shortcut.
How many trials would a target require?
Invert the calculation before a campaign. To obtain an upper bound no greater than ε under a zero-miss acceptance rule, require (1−ε)ⁿ≤α. Taking logarithms and noting log(1−ε) is negative gives the minimum below. Dividing by a negative reverses the inequality, crucial to avoid an undersized test.
At least 2995 independent zero-miss trials would support that specific one-sided bound. One miss invalidates the zero-event formula and requires the appropriate binomial calculation. This does not certify an embedded system; it specifies a statistical experiment. If p were truly 1%, a hundred independent requests would have probability 1−0.99¹⁰⁰≈63.4% of at least one miss. Rare per request can become common across a sequence.
From the example to a device protocol
First fix the timing boundary. For asynchronous calls, stopping immediately after submission measures submission, not completion. A counter needs sufficient resolution, correct conversion and wraparound handling. Official Zephyr timing documentation reads a counter before and after a block and converts cycles to time; the timer can depend on architecture, SoC or board. We have not run that API on hardware.
For each series I would record board revision, processor, clocks, OS or runtime version, compiler flags, model precision, input shapes, memory use and sample count. Separate cold start and steady state, documenting warm-up, concurrency, interrupts and thermal conditions. Do not automatically discard slow cases: determine whether they are measurement errors or real service events. Millisecond units alone do not make different protocols comparable.
Model weights differ from peak RAM; latency, power and energy per inference also differ. We infer no consumption from synthetic times. Hard timing requirements need a worst-case argument under allowed conditions, including blocking and interference, not merely a percentile. For delay-tolerant requirements, measured distributions and an explicit expired-result policy may fit better. The choice depends on delay consequences, which our example does not quantify.
Reproduce the result and read the code
The package includes both full lists, their generator, JSON results and plotting code. No randomness means no seed. Comparing t with 10000 counts misses in microseconds; dividing by 1000 converts outputs to milliseconds. Percentiles use sorting and the declared rank. Probability bounds use expm1 and log1p to reduce precision loss near one. These are numerical choices, not new statistical models.
Executed checks verify mean 8 ms, ten B misses, p99 7.8 ms and minimum 2995 trials. Code checks that 2995 meets the inequality and 2994 does not. No Zephyr or neural-network performance test is hidden behind these results. Real data require traceable measurements replacing the lists and reassessment of independence, inputs and condition stability.
Answer to the opening question
An 8 ms mean does not prove timeliness. Define start and end, count deadline misses, examine the tail and separate observation from guarantee. B looks better at p99 yet has ten misses per thousand; even zero misses in a thousand independent trials leaves a statistical bound near 0.3%. Measure the actual requirement instead of asking the mean a question it cannot answer.
Sources and attribution limits
NIST exact-binomial documentation is the statistical reference; we derive the special zero-event one-sided case directly. Official Zephyr pages explain counters, conversions and measurement boundaries, consulted as available on 28 September 2026, not as firmware used in our experiment. This is an educational monograph with reproducible calculations, not a hardware benchmark, original research or evidence of an available EL-AI embedded product.
NIST — Exact Binomial Confidence Limits.
Zephyr — Executing Time Functions.
from math import ceil, log, log1p, expm1
x = [7800]*990 + [27800]*10 # synthetic microseconds
q99 = sorted(x)[ceil(.99*len(x))-1]
print("mean ms:", sum(x)/len(x)/1000)
print("p99 ms:", q99/1000)
print("deadline misses:", sum(t > 10000 for t in x))
print("95% upper p after 1000 independent zero-miss trials:",
-expm1(log(.05)/1000))
print("zero-miss trials for upper p <= .001:",
ceil(log(.05)/log1p(-.001)))
# Constructed data; no hardware timing was performed.
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 28 September 2026.

