Spare time can arrive too late
A device handles a sensor and periodically runs an AI model on the same processor. Counting the work suggests sufficient overall capacity: the processor should not remain busy all the time. Yet a response arrives after it was needed. There is no contradiction. Spare time may lie after the deadline, while a higher-priority task interrupts inference during the useful interval. Understanding this requires more than asking how many milliseconds the model takes alone: we must ask when it is allowed to use them.
We will study two tasks with deliberately simple numbers and reconstruct every execution interval. We compare a priority fixed in advance with a rule that considers the nearest deadline. Then we reduce AI computation time and add a short initial block to examine remaining margin. These are synthetic simulator timings, not measurements on microcontrollers or accelerators. By the end we can distinguish average capacity, computation time, and response time, and interpret lateness without immediately blaming the neural model.
Three different clocks in one device
A task is a recurring activity, such as reading and processing a sensor. Each concrete activation is a job, one request for work. Period T says how often a request is released; cost C is the processor time needed to finish without interruption; relative deadline D is how long we may wait after release. Response time R spans release to actual completion, including waiting. Here all are in milliseconds, abbreviated ms: one millisecond is one thousandth of a second.
Assign sensor S a cost of 2 ms every 5 ms and AI inference a cost of 4.5 ms every 8 ms. Both must finish before their next request, so D equals T. They start together at time zero. Assume one processing core, independent activities, constant costs, no memory or accelerator waiting, and immediate, zero-cost preemption. These assumptions make the problem analyzable; they do not automatically describe an operating system or board. In particular, 4.5 ms is a chosen experimental input, not a proven maximum execution time for a real model.
| Activity | C (ms) | T (ms) | D (ms) |
|---|---|---|---|
| S | 2 | 5 | 5 |
| AI | 4.5 | 8 | 8 |
Average capacity tells us how much work, not when it finishes
To check total load, divide each activity’s cost by its period and sum. C/T is dimensionless: the sensor needs two milliseconds of computation for every five available, or 40% of the processor. AI needs 4.5/8, or 56.25%. Total utilization is 96.25%, below 100%. This excludes average overload by periodic work, but does not establish whether the priority rule places that work in the right intervals.
Consider 40 ms, a common multiple of 5 and 8. Eight S requests and five AI requests arrive, requiring 8 × 2 + 5 × 4.5 = 38.5 ms of work. That leaves 1.5 ms. But if a response was needed at 8 ms, spare time at 39 ms cannot rescue it. Think of appointments: half an hour free in the evening does not fix incompatible morning appointments. Unlike that analogy, computation can be interrupted and resumed; the compared rules exploit this possibility.
Fixed priority: follow the first job until it is late
The first rule gives higher priority to the shorter-period activity: S always precedes AI. This is rate monotonic, RM, ordering priority by activation frequency. At zero, the sensor occupies 0–2 ms. AI runs from 2 to 5, completing 3 of its 4.5 ms. At 5, a new S interrupts it until 7. AI resumes and finishes at 8.5, half a millisecond after its deadline. It actually consumed only 4.5 ms but spent 8.5 ms between release and completion.
The processor is neither faulty nor slower: it follows the assigned rule exactly. The sensor released at 5 ms has deadline 10, whereas the older AI job has deadline 8. Fixed priority ignores this relative urgency and serves the sensor first. This creates the counterexample to “below 100% load means everything is on time.” The simulator does not cancel expired jobs: they finish and lateness is recorded, avoiding concealment of a violation by deleting its request.
Counting interruptions before knowing the duration
We can reconstruct the result iteratively. AI response must include its own C milliseconds plus all sensor work that precedes it. If response lasts R, the interval starting at zero contains ceil(R/5) S jobs, each costing 2 ms. ceil rounds upward to an integer. The new estimate of R therefore depends on the previous estimate. Start with AI cost alone and repeat until the count stops changing: this is a fixed point, a value that reproduces itself in the relation.
At an assumed 4.5 ms we count one sensor and obtain 6.5. But 6.5 ms includes the sensor release at 5: now there are two and the total becomes 8.5. There are still two within 8.5 ms, so iteration stops. A job released exactly at completion cannot delay a job already completed: this convention explains ceil and why interval endpoints matter. The deadline test is R ≤ D. Here 8.5 > 8, so failure is explicit even though C = 4.5 is much smaller than D.
For our independent, periodic, preemptible tasks, simultaneous release is the fixed-priority critical case: it concentrates higher-priority requests ahead of the observed one. We do not extend this to arbitrary locks, suspensions, or accelerators. C. L. Liu and James W. Layland’s 1973 Journal of the ACM paper provides the theoretical reference for these assumptions and the fixed-versus-deadline-driven comparison. It is mathematical analysis, not a modern inference benchmark; our simulation uses separately chosen data.
Changing the order while keeping the same work
The second rule chooses the ready job with the nearest absolute deadline: Earliest Deadline First, EDF. Absolute deadline is release time plus D, not D alone. Execution is identical until 5 ms. At 5, however, the running AI must finish by 8 and the new sensor by 10. EDF lets AI continue until 6.5, then serves the sensor, which finishes at 8.5, still before its own deadline. No computation has accelerated. We have moved an interruption that the first rule imposed too early.
Over the simulated 40 ms, EDF completes every request by its deadline. For ideal tasks under the stated assumptions, the classical result links EDF feasibility to U ≤ 1; this does not authorize using 100% of a real board while ignoring interrupt overhead. Simulation confirms our case, while the theorem has more precise assumptions than a load slogan. A dynamic policy also needs correct implementation and defined overload handling: it does not replace analysis of actual resources.
Half a millisecond less can remove lateness without creating margin
Return to RM and reduce only AI cost from 4.5 to 4 ms. Now U = 2/5 + 4/8 = 0.9. Iteration becomes 4 → 6 → 8 → 8. AI finishes exactly at its deadline. The simulator records no violations over 40 ms, but the first job has zero margin. In a real project, replacing an assumed cost with a measured mean would make this equality unjustified reassurance: a longer execution, uncounted context switch, or other interference can change the outcome.
The familiar sufficient RM bound for two tasks, 2(√2 − 1), is approximately 0.8284 under classical assumptions. Exceeding it does not mean every task set fails; it means that general test no longer guarantees feasibility. Our on-time 90% case demonstrates the distinction. By contrast, observing a request completed after D proves failure of that simulated configuration. An unmet sufficient condition and an actual violation are different evidence and require different wording.
A short block changes the problem
In the final variant, retain C_AI = 4 ms but impose an initial 1 ms interval in which neither task can execute. This is an abstract nonpreemptible occupation already present at release, not a simulation of a mutex or specific driver. S now runs 1–3, AI 3–5, the second S 5–7, and AI 7–9. Inference misses its deadline although the periodic pair still requires 90% capacity. The unrelated 1 ms is additional work and is counted in the trace.
B denotes only the initial block imposed in this variant. We have not proved that 1 ms bounds blocking in a real system or analyzed every synchronization protocol. Over 40 ms, periodic work is 36 ms and blocking adds 1: total executed occupancy is 37 ms, not 36. The comparison shows that nonpreemptible work belongs in the model and that a complete trace prevents confusing selected-task load with everything occupying the processor.
Reading the diagram without confusing requests
The figure shows the first 12 ms of four simulations; the results file retains all 40 ms. Blue marks sensor execution, green AI execution, and orange initial blocking. The red vertical line is the first AI job’s deadline; the black dot is its completion. Green after that dot belongs to other requests: color identifies an activity, not a single job. Half-millisecond subdivisions are simulator units, not measured interrupts. The first two rows show identical work arranged differently.

The simulator uses integer 0.5 ms ticks, placing every example duration and release on the grid without time rounding. At each tick it inserts new requests, chooses by RM or EDF, and subtracts one tick of remaining work. It records release, deadline, and completion for every request. Assertions verify first-AI responses of 8.5, 6.5, 8, and 9 ms, plus no violations for EDF and unblocked 4 ms RM. Late-request counts are trace results, not failure probabilities.
What remains unknown before discussing a board
On a real device, C is the first difficult input. Average timing on some inputs is not automatically an upper bound. Different operator paths, contested memory, interrupts, variable frequency, and temperature can change observed work. Evaluation should state platform, runtime and model versions, compiler, clock, inputs, batch, threads, and measurement conditions. With an accelerator, waiting for completion does not always occupy the CPU: our single-processor model cannot be applied by indiscriminately substituting total elapsed time for C.
Request arrivals may also differ from periodic assumptions. Communication can burst, a sensor can deliver blocks, and acquisition can delay inference release. The system deadline may start at the physical measurement, whereas our R starts at job release. Time before release does not disappear: it belongs in the sensor-to-response path alongside relevant communication and actuation. We have not simulated these elements and claim no end-to-end timing guarantee.
Which intervention addresses the observed cause?
Reducing C with a smaller model may help, as the 4 ms variant shows, but must preserve predictive usefulness and include other costs. Priority changes act on order: EDF improves our case without changing the model. Shortening nonpreemptible sections acts on B; moving work to another core introduces communication and shared-resource questions. These remedies are not interchangeable. The trace helps choose which hypothesis to investigate before buying faster hardware or blaming only the neural network.
A future device test should measure execution and response separately, record releases and preemptions, and include expected interference scenarios. Results should be compared with an explicit timing budget, keeping means, percentiles, and demonstrable bounds distinct. We have no power, energy-per-inference, peak-RAM, or board-performance measurements here. Embedded systems are an editorial topic and area of interest to explore; this study establishes no available EL-AI product or installation. No physical test or certification is claimed.
The answer: time must be available before the deadline
A processor can have spare capacity and still deliver late because capacity is an aggregate quantity, whereas a deadline concerns a specific interval. In our example, 96.25% load gives first completion at 8.5 ms with RM and 6.5 ms with EDF for identical work. Reducing AI cost to 4 ms removes RM lateness but leaves zero margin; a 1 ms initial block brings lateness back. The practical conclusion is to reconstruct the full timing path and processor-access rules as well as measuring the isolated model.
Sources, code, and reproducibility
Code and results can be downloaded below the listing. The experiment is deterministic and needs no random seed. The snippet prints periodic load, first-job response, and lateness count for each case, then calculation iterations. Versions used are Python 3.14.0 and Matplotlib 3.11.2. The historical reference is a 1973 publication, not a 2026 research development; its assumptions, relevant proofs, and limitations were read, and our example is not presented as reproducing author measurements.
from experiment import run
r = run()
for name, case in r['cases'].items():
print(name, case['periodic_utilization'],
case['first_AI_response_ms'], case['late_jobs'])
print(r['recurrences'])
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 9 October 2026.

