A plausible result from data that never existed
A small device monitors a machine’s vibration. Every ten milliseconds it collects a block of measurements and hands it to an artificial intelligence model. The model seems fast enough if we look at its average runtime. Yet some answers are inexplicable. The problem can arise before any statistical error: while the model reads the measurements, the acquisition circuit is already replacing them with the next ones. The model therefore receives fragments from two different times, without any instruction necessarily raising an error.
Abstract. We study when a memory region can be reused without changing an input that is still being read. We build a synthetic timeline, derive a data-lifetime condition, compare two and three buffers, then introduce explicit ownership. The central result is specific: extra memory can absorb a delay, but an ownership protocol prevents overwriting an occupied block; when resources run out, the loss must be explicit. Every number comes from a Python simulation, not measurements on a board or an EL-AI product.
Two activities proceeding together
A buffer is a memory region that temporarily holds data. A producer fills it; a consumer reads it. Here the producer is acquisition assisted by DMA, or direct memory access: a device transfers data without asking the CPU to copy every sample. Meanwhile the CPU, the processor executing our computation, can run inference, meaning application of the model to the data. Concurrency is useful, but the two actors may access the same memory.
Imagine two trays: one is filled while the other is read. This is the intuition behind double buffering, often called ping-pong buffering. The analogy has an important limit: memory has no hand automatically signalling that a tray is still occupied. A valid address remains valid even after its contents change. Keeping a pointer remembers where to read; it does not freeze what is there. We need the entire block to remain stable throughout all reads that depend on it.
The experiment: eight blocks and one isolated delay
Our hypothetical sensor produces one channel at 16,000 samples per second. We group 160 samples, each represented by two bytes. A block therefore contains 320 bytes and takes 160/16,000 seconds, or 10 ms, to acquire. These are parameters chosen to make the case readable, not specifications of tested hardware. Blocks are numbered from zero: block 0 is written from 0 to 10 ms, block 1 from 10 to 20 ms, and so on.
The consumer processes one block at a time in arrival order. The eight assigned service times are 6, 6, 14, 6, 6, 6, 6 and 6 ms: the average is 7 ms. We assume the input must remain available throughout processing, including any wait before the CPU starts. There are no additional interrupts, copies, caches or synchronization costs. We are not estimating network speed; we isolate the relationship between input lifetime and memory reuse.
With two cyclic buffers, one holds even-numbered blocks and the other odd-numbered blocks. Block 2 finishes arriving at 30 ms; the CPU reads it from 30 to 44 ms. But block 4 starts filling that same buffer at 40 ms. Reading and overwriting overlap for four milliseconds. An average below 10 ms does not prevent this: correctness concerns every block, not the average behaviour of the sequence.
How much time is actually available before reuse?
To calculate this interval, let T be the acquisition period, B the buffer count and k the block index. Block k completes at rₖ=(k+1)T. Its buffer is overwritten when block k+B begins, at dₖ=(k+B)T. Here d denotes the deadline for data availability, not a clinical or control deadline. Their difference is the usable time after acquisition:
Cₖ is the time spent processing the block and Wₖ the wait before starting; both are in milliseconds, like T. The second line says that waiting plus processing must finish before reuse. With B=2, only 10 ms remain: the buffer being filled is not free waiting space. For block 2 we obtain 0+14≤10, which is false. With B=3 it becomes 14≤20: this delay fits within the available window.
To include the effect on later blocks, we calculate start sₖ as the later of the block’s arrival and the previous finish fₖ₋₁. We then add Cₖ. This is a serial processor without preemption, meaning it does not interrupt one block to run another:
Margin mₖ is positive when the data are no longer needed before reuse, and negative when the lifetimes overlap. Block 3 arrives at 40 ms but waits until 44; it finishes at 50, exactly when the same buffer must receive block 5. Its margin is zero. The simulation handles equality with an ideal event ordering; a real system should not treat it as a guarantee because no margin remains for jitter, synchronization or the actual final access.
| Block | Arrival ms | Start ms | Finish ms | Reuse with B=2 ms | Margin ms |
|---|---|---|---|---|---|
| 0 | 10 | 10 | 16 | 20 | 4 |
| 1 | 20 | 20 | 26 | 30 | 4 |
| 2 | 30 | 30 | 44 | 40 | -4 |
| 3 | 40 | 44 | 50 | 50 | 0 |
| 4 | 50 | 50 | 56 | 60 | 4 |

Read the figure from left to right: grey is acquisition, blue is a read within the window and red a read extending beyond reuse. CPU timings are identical in both panels. Only the time at which the producer returns to the same memory changes. The table also exposes block 3’s wait: comparing each Cₖ with T alone would hide the effect of queued work.
What the model can actually read
An overlap does not automatically prove how much the prediction changes: that depends on the model’s read order. We can nevertheless construct a minimal counterexample to data consistency. Consider four representative positions at indices 0, 40, 80 and 120. New block 4 begins overwriting them at 40, 42.5, 45 and 47.5 ms respectively. At 44 ms, an ideal snapshot contains generation labels [4, 4, 2, 2]. These labels identify the source block of each value; they are not measured sensor amplitudes.
The device could therefore construct features from a signal never acquired as such. We have not run a neural network and assign no diagnostic error rate to this example. We demonstrated a narrower but fundamental point: the contract “this inference uses block 2” is violated. Improving model accuracy on a clean dataset does not repair that contract.
Giving each buffer an owner
Now change the rule: instead of automatically choosing the next cyclic address, the producer may use only a free buffer. We distinguish four states. FREE means available; FILLING reserved for acquisition; READY complete and waiting; READING reserved for the consumer. The normal path is FREE → FILLING → READY → READING → FREE. “Ready” does not mean “free”: waiting data already belong to a job that still needs to read them.
The second simulation implements these states with an event queue. At 40 ms, with two buffers, block 2 is still being read and block 3 has just become ready. No buffer is free for block 4. Our chosen policy drops the entire incoming block and records its index. The run completes blocks 0, 1, 2, 3, 5, 6 and 7: no overwrite is permitted, but one loss is explicit. With three buffers, no block is lost on this same finite trace.
This separates two frequently confused requirements: integrity of delivered data and completeness of acquisition. We can satisfy the first while losing the second. A block sequence counter helps the receiver notice the jump from 3 to 5; it does not reconstruct the missing measurement. In time-series analysis the gap must retain its timestamp: silently concatenating the remaining blocks would invent continuity the sensor did not provide.
Why three buffers do not solve every delay
A reserve absorbs a transient disturbance, not a permanent imbalance. Repeat the cyclic calculation with every processing time equal to 12 ms and T still 10 ms. After the first block, each arrival adds 2 ms of waiting: Wₖ=2k ms. With three buffers the window is 20 ms, so the margin becomes 20−(2k+12)=8−2k ms. Block 4 reaches zero margin; block 5, the sixth in the sequence, exceeds reuse by 2 ms. The program verifies exactly this first exceedance.
No finite number of buffers can preserve every block indefinitely when this single consumer remains slower than the producer. Explicit ownership protects occupied blocks but eventually requires dropping data, slowing the source if possible, or increasing processing capacity. Here “drop a block” is an abstract decision: on a real peripheral, one must verify how to stop, divert or ignore acquisition without DMA continuing to write into an occupied buffer.
Copying, retaining or reducing work
Immediately copying the input to private memory can shorten the time for which the sensor buffer must remain intact. In the reuse condition, Cₖ must then be replaced by the time until the final read from the original buffer, not automatically the entire inference. But the copy must finish before overwriting starts, its destination must remain occupied while needed, and transfer costs time and memory bandwidth. We did not measure these costs and do not assume copying is always best.
Moving from two to three buffers in our example increases block storage alone from 640 to 960 bytes: 320 extra bytes. This is not the application’s peak RAM. Model weights and activations, stack, queues, alignment and other structures remain. “Zero-copy”, meaning reading acquired memory directly, does not mean “zero waiting”: it removes a copy but may prolong buffer occupancy. The right choice depends on the limiting resource and which losses the task can tolerate.
From simulation to firmware: what remains to be checked
Zephyr’s primary ring-buffer documentation illustrates the separation between accessing memory, transferring data and confirming produced or consumed bytes. Its concurrency section states that there is no general internal concurrency control; it distinguishes one producer and one consumer from multiple actors and discusses write visibility across CPUs. This helps explain the contract; it does not prove that our simulator implements that kernel. We read Concepts, Instantiation and Usage, Concurrency and Internal Operation; the APIs depend on the selected firmware version.
In the educational program, simultaneous events are ordered as CPU completion, acquisition completion, then new acquisition start. State changes occur in one Python process. This makes the case deterministic but proves neither atomicity, cache coherence nor operation ordering on a microcontroller. A physical test would require an identified board and peripheral, runtime version, DMA configuration, buffer placement, synchronization strategy and an access timeline. Saturation behaviour would also need testing, not just normal operation.
Answer and reproducibility
The sensor can indeed overwrite data the model is still reading, even when average processing appears fast. What matters is how long each block remains necessary and who may reuse its memory. In our example, a third buffer absorbs an isolated delay; explicit ownership avoids corruption but requires one drop with two buffers. Neither observation establishes a general real-time guarantee. The practical consequence is to design the model, acquisition and loss policy together: an AI answer is meaningful only if we know which data produced it.
The short listing below reproduces the cyclic timeline: max preserves waiting, while reuse−end measures margin. The archive also includes ownership-based event simulation, the full trace, the mixed-generation counterexample and persistent overload. Assertions verify the overlapping block, the dropped block and the first exceedance with constant 12 ms service. There is no randomness, so no seed is needed. The timeline scan is linear in block count; the event queue uses a heap and costs O(n log n + nB), with O(nB+n+B) memory: n is the block count and B the buffer count. We search for free buffers and record all B states at each event. For fixed B these costs become O(n log n) and O(n).
Primary source and code
Zephyr Project Documentation — Ring Buffers.
T = 10 # milliseconds
costs = [6, 6, 14, 6, 6, 6, 6, 6]
for B in (2, 3):
end = 0
for k, C in enumerate(costs):
release = (k + 1) * T
start = max(release, end)
end = start + C
reuse = (k + B) * T
print(B, k, start, end, reuse - end)
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 3 October 2026.

