ELAI S.r.l.

Why an AI service slows down before capacity runs out

A 100 ms computation can become a 750 ms response. Queues, variability and capacity headroom through derivation and a reproducible experiment.

Why an AI service slows down before capacity runs out

Eight requests per second for a server that can handle ten

An inference service takes 100 milliseconds per request on average, so it could complete ten per second without idle time. Only eight arrive per second: why do users wait much longer than 100 milliseconds? Average capacity ignores arrival timing and duration variability. Three closely spaced requests can queue even if the server later sits idle. Future idle time cannot process a request before it arrives.

Abstract. Study one server processing requests in arrival order. Separate service, waiting and response; derive the M/G/1 queue formula and compare three duration distributions with equal mean. Mean response becomes 300, 500 or 750 ms without changing arrival rate or 100 ms mean service. Event simulation checks the scale and exposes statistical variability. This educational model is neither an EL-AI load test nor a prediction for a GPU using dynamic batching.

First the concrete sequence, then the model

Four requests arrive at 0, 40, 80 and 500 ms, each requiring exactly 100 ms. The first starts immediately; the second starts at 100 ms after waiting 60 ms; the third starts at 200 ms after waiting 120 ms; the fourth starts immediately at 500 ms. Processing duration is constant, yet responses take 100, 160, 220 and 100 ms. The model itself need not slow down for perceived service to deteriorate.

inizioᵢ / startᵢ = max(arrivoᵢ / arrivalᵢ, fineᵢ₋₁ / endᵢ₋₁) fineᵢ / endᵢ = inizioᵢ / startᵢ + Sᵢ Tᵢ = Wᵢ + Sᵢ

S is service time, W the wait before service, and T total time in the system. The first relation drives the simulator: start only after arrival and the previous completion. Network transport, authentication and external preprocessing are excluded; if present, measure and include them consistently. The word latency alone does not identify the interval being observed.

Assumptions that make the queue tractable

M/G/1 assumes Poisson arrivals at rate λ requests per second: independent exponential interarrival times, not regularly spaced requests. G denotes general independent identically distributed service times, also independent of arrivals. One denotes one server. Assume unlimited queue, no abandonment, no priorities, no interruption of service, and FIFO—first in, first out.

Write m=E[S] for mean service and dimensionless ρ=λm for load. λ=8 and m=0.1 s give ρ=0.8. A stable stationary distribution requires ρ<1; finite mean waiting in our formula additionally requires finite E[S²]. Stability means no indefinite long-run queue growth, not that every request meets a deadline. This 80% is ideal-server load, not automatically a GPU-utilization counter.

Why squared duration appears

An arrival may find one request in service and others queued. It waits for residual service R and full service of jobs ahead. Long jobs not only last longer; an arrival is more likely to encounter them still running. Mean duration alone therefore does not suffice. Over a job lasting s seconds, residual time falls linearly from s to zero, with area s²/2. Long durations weigh quadratically in time-average residual service.

E[R] = λ E[S²] / 2 E[W] = E[R] + m E[N_q] E[N_q] = λ E[W] E[W] = λ E[S²] / (2(1−ρ))

N_q counts waiting requests, excluding the one in service. The third line uses Little’s relation among mean count, flow and waiting. Poisson arrivals see time averages—PASTA—linking arrival snapshots to time averages. Substitute the third line into the second and collect E[W] to obtain Pollaczek–Khinchin. Units agree: λ is s⁻¹ and E[S²] is s², yielding seconds.

Same mean, three different waits

First, every service lasts exactly 0.1 s: E[S²]=0.01 s². Second, exponential duration has mean 0.1 s and second moment 0.02 s². Third, nine in ten jobs take 0.05 s and one in ten takes 0.55 s, independently. Mean remains 0.9·0.05+0.1·0.55=0.1 s, but second moment is 0.9·0.05²+0.1·0.55²=0.0325 s². These are synthetic alternatives isolating variability, not company logs.

DistributionE[S] (ms)E[W] (ms)E[T] (ms)
S = 0.1 s100200300
S ~ Exp(mean=0.1 s)100400500
P(S=0.05 s)=0.9; P(S=0.55 s)=0.1100650750

For the mixture, 8·0.0325/[2·(1−0.8)]=0.65 s waiting, plus 0.1 s service. The server works the same average fraction of time in every case. What changes is the delay a long job imposes behind it. A useful indicator is C_s²=Var(S)/m², squared coefficient of variation: 0, 1 and 2.25. Since E[S²]=m²(1+C_s²), mean waiting is ρm(1+C_s²)/[2(1−ρ)].

Near capacity, small changes become costly

The denominator 1−ρ is remaining headroom. With exponential mean service 100 ms, increasing arrivals from 8 to 9 per second raises mean response from 0.5 to 1 s: 12.5% more traffic doubles response. At 9.5 per second it is 2 s. There is no universal magic 80% threshold; the curve depends on distributions, objectives and architecture. Choose acceptable response and verify model fit before selecting reserve capacity.

Theoretical mean response versus load ρ with mean service fixed at 100 ms. Curves differ only in duration distribution. These are stationary model means, not percentiles or production measurements.
Theoretical mean response versus load ρ with mean service fixed at 100 ms. Curves differ only in duration distribution. These are stationary model means, not percentiles or production measurements.

Faster computation can affect response nonlinearly. In the exponential case, reducing mean service from 100 to 80 ms at eight arrivals per second lowers ρ from 0.8 to 0.64 and mean response from 500 to about 222 ms. This is a parameter change under fixed assumptions, not a hardware promise. Adding a second server changes the model; simply replacing m by m/2 is incorrect. Load balancing and shared queues matter too.

A simulation that exposes its own uncertainty

We executed eight independent simulations per distribution. Each starts empty, discards 20,000 warm-up arrivals and retains 300,000. Interarrival times are exponential with mean 1/8 s; the recurrence above drives service. NumPy 2.5.3 creates 24 distinct streams with SeedSequence(20260929).spawn(24), ordered constant, exponential, mixture. Mean waits averaged over eight runs are 199.0, 403.7 and 658.0 ms, versus theoretical 200, 400 and 650 ms.

Standard errors estimated from eight independent run means are about 1.0, 4.7 and 6.1 ms. Consecutive jobs in one queue are dependent: a backlog affects many later arrivals. Treating them as independent would understate uncertainty. Replicates expose this limitation, but eight do not prove asymptotic behavior or eliminate every warm-up effect. The formula is analytic under assumptions; simulation is a finite numerical check with disclosed deviations.

Where this model stops describing inference

Real inference engines may process requests together. Batching makes service depend on queue state, violating independence here. Different output lengths, cancellations, caches and accelerator sharing introduce further dependencies. Arrivals may be bursty rather than Poisson. A finite queue changes interpretation too: rejected requests mean low latency among accepted requests can conceal poor service.

Application requires arrival, start and completion timestamps, request class, outcome and scheduling policy. Check interval distributions, service variability and temporal correlations, then compare predictions with load tests outside parameter-fitting data. None was performed on EL-AI. Mean formulas do not determine 95th or 99th percentiles: distributions can share mean and variance but differ in tails. Percentile objectives need measurements or a further distributional model.

Available capacity is not guaranteed response time

A service slows before saturation because it must absorb irregular timing, not merely average work volume. Headroom drains backlogs; variability determines how often and how long they form. At eight requests per second, reducing variability changes waiting despite unchanged mean computation. Design therefore needs separate service and queue measurements and headroom chosen against a verifiable objective, rather than maximum utilization as the sole efficiency metric.

Reference and reproducibility

Eytan Modiano, MIT — M/G/1 Queues, Lectures 8–9, course 6.263J, Fall 2002, slides 2–6.

The reference provides assumptions, formula and residual-service proof. AI-service numbers, mixture and simulations are educational examples executed for this article. The snippet reproduces the analytic comparison; the archive includes the full simulator, per-run results and plots. These are neither reproductions of an author benchmark nor original peer-reviewed research.

mean_service = 0.1  # seconds
arrival_rate = 8.0  # requests per second
rho = arrival_rate * mean_service
for name, second_moment in [('constant', .01), ('exponential', .02), ('mixture', .0325)]:
    waiting = arrival_rate * second_moment / (2 * (1-rho))
    print(name, 'queue seconds:', waiting, 'total seconds:', waiting+mean_service)
# Analytical M/G/1 values, not a real inference benchmark.

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.