The problem: doubling resources without halving the wait
A business wants shorter AI response times. Computation appears divisible: if one device needs twelve milliseconds, two should need six. Yet devices must exchange intermediate results, wait for one another and combine the answer. That time can consume the entire benefit. Our question is concrete: how much work is needed before splitting a model over two devices makes the same operation faster?
We first build a small splittable calculation, then a time budget and a break-even threshold. Without reconstructing every equation, follow three questions: what is split, what must travel and how much time is actually saved? All times below are hypothetical parameters in a Python-executed arithmetic model. We used no two-GPU platform, measured no network and reproduced no benchmark; these are not EL-AI product or installation results.
Splitting a model is not replicating it
Data parallelism uses model replicas on different examples; training also requires coordinating updates. Model parallelism instead distributes parts of the network involved in the same operation. Here we study matrix-level model splitting, called tensor parallelism. A tensor is a multidimensional array of numbers; a matrix has two dimensions. Increasing requests processed per second is not the same as reducing one request’s wait.
Take a minimal network: row vector x, matrix A producing four intermediate values, ReLU replacing negative values with zero, and matrix B producing two outputs. This is not a trained language model; it exposes where communication arises. Shapes are x: 1 × 2, A: 2 × 4 and B: 4 × 2. Numbers have no physical units.
Split A’s first two columns from its last two, and B’s corresponding first two rows from its last two. Device 1 computes ReLU([1,1]) and final contribution [1,1]. Device 2 computes ReLU([0,4]) and contribution [−4,8]. Their sum is [−3,9], exactly the original output. Summation must happen somewhere; if both devices need the full output next, both must receive it. An all-reduce collective combines contributions and makes the result available to all participants.
The sum cannot move freely through a nonlinear function. ReLU(2 − 3) is zero, whereas ReLU(2) + ReLU(−3) is two. Row and column partitioning therefore changes mandatory communication points. Our Python integer example checks algebraic equivalence; floating-point precision and summation order may introduce numerical differences. A distributed implementation needs output-quality checks as well as timing.
A time budget with explicit units
Now model a larger work phase. Let S be indivisible time and C divisible time on one device. With two identical devices, the ideal latter time is C/2. Introduce local efficiency η: at η = 0.8 each half works at 80% of assumed ideal efficiency, giving C/(2η). η is neither probability nor dashboard GPU utilisation: it summarises computation-shape and implementation effects, excluding separately counted communication.
Q counts serialised communications in the phase; ℓ is each startup cost, V bytes transferred along each effective critical path and B effective bytes per second. H is total communication time. V need not equal logical tensor bytes: a collective algorithm may send multiple pieces or rounds. Q and V must come from the actual configuration. The formula assumes equal communications, no compute overlap, no competitors and no request queue.
| Hypothetical parameter | Value |
|---|---|
| S | 0.8 ms |
| C | 12 ms |
| η | 0.8 |
| Q | 4 |
| ℓ | 0.03 ms |
| V | 2 000 000 bytes |
| B | 10 GB/s = 10¹⁰ bytes/s |
Use decimal gigabytes: ten billion bytes per second, not gigabits or gibibytes. Each transfer takes 2,000,000 / 10,000,000,000 seconds, or 0.2 ms, plus 0.03 ms startup. Four cost H = 0.92 ms. Parallel computation takes 12/(2 × 0.8) = 7.5 ms. Thus T₁ = 12.8 ms and T₂ = 9.22 ms, a speedup about 1.388 rather than two. These numbers evaluate the table’s assumptions, not an existing card.
Deriving the break-even point
To decide whether splitting helps, require T₂ smaller than T₁. S occurs in both and cancels, leaving computation savings and H. Rearranging gives the following condition: the computation time saved must exceed the communication time added.
At η = 0.8 the bracket is 0.375, giving threshold 0.92/0.375 = 2.45333 ms. Times are equal at that value; C must exceed it to help. With C = 1 ms and other assumptions unchanged, one device takes 1.8 ms and two take 2.345: more resources increase waiting. If η is at most 0.5, the denominator is not positive; splitting saves no computation time, and positive communication prevents speedup in this model.
Reducing effective bandwidth to 1 GB/s while C stays 12 ms makes H = 8.12 ms and T₂ = 16.42 ms: even the larger job slows down. At 50 GB/s, H falls to 0.28 ms and T₂ to 8.58 ms. Infinite bandwidth removes V/B but not four startup costs. Large messages are often bandwidth-sensitive; many small ones can be startup-dominated. Grouping messages may help but can delay data availability, so it is not free improvement.

Latency, throughput and efficiency are different
Speedup is T₁/T₂; parallel efficiency for two devices is that quantity divided by two, about 0.694 in the base case. This does not mean 30.6% of energy is wasted: we compare times and resource count, not power. Both devices are occupied together. If the model fits one card and requests are independent, two replicas may increase service capacity; comparing replicas against sharding requires workload, queues and memory constraints absent from this calculation.
S also matters. It cancels from break-even because it is common, yet reduces relative benefit. With S = 100 ms, C = 12 ms and 10 GB/s, times become 112 and 108.42 ms: only about 1.033 times faster. Optimising a small part of total waiting may barely affect users. If distribution introduces extra serial work, add it to T₂; break-even worsens, and that new cost cannot be cancelled as common.
Overlap does not make communication disappear
A system can communicate some results while computing other work. Imagine perfectly overlapping blocks of duration C/(2η) and H. Their combined time cannot be smaller than the longer block, giving optimistic lower bound S + max[C/(2η),H], or 8.3 ms instead of 9.22 in the base case. This is not a prediction: dependent computation must await arriving results, and traffic may compete for memory or execution resources.
Increasing batch size—processing more examples together—can improve local computation efficiency, but also changes communicated data and waiting to form a batch. Multiplying C while keeping all else fixed is invalid if V and η also change. Moving from two to four devices likewise changes collective algorithms, topology, message count and local shapes. Our plot varies C with other parameters fixed to isolate a cause, not promise that curve for every real workload.
From derivation to verifiable measurement
The attached code uses no distributed library: it computes whole and split networks with integer lists, checks equality and evaluates timing formulas. The final snippet prints output [−3,9] and six JSON-stored cases. The curve is a deterministic grid with no random seed. We did not time Python and call that GPU time; that would be a different, misleading comparison.
Hardware validation should fix model, precision, input shapes, devices, links, drivers and software versions. Separate startup and warmup from stable repetitions, wait for actual asynchronous completion and inspect compute/communication traces. Mean, median and percentiles answer different questions; a lower mean may hide long tails. This proposed protocol makes the analysis testable, but was not executed.
Megatron-LM, arXiv version 4 dated 13 March 2020, describes coordinated matrix partitions and Transformer communications. Its scaling experiments also vary model size. They therefore do not measure our fixed-work comparison, and we transfer none of their percentages to this example.
PyTorch documentation describes all-reduce and communication profiling tools. It guides observation of real implementations; it does not prove that our formula models every backend or that asynchronous calls remove dependencies.
The same formula can frame a network question
We asked how much computation is needed. Reverse the question: for fixed work, what minimum effective bandwidth makes splitting useful? Isolate V/B in the previous inequality. Here C and ℓ are in seconds, V in bytes and B in bytes per second. The result concerns effective bandwidth for this communication, not an advertised port rating.
D is the time available for each byte transfer after startup. With C = 0.012 s and base parameters, D = 0.001095 s; bandwidth must exceed about 1.82648 GB/s. If D is zero or negative, no finite bandwidth suffices: startup alone consumes the margin. A theoretically wider link therefore cannot fix every slowdown. Measure the full path with the same message sizes and algorithm.
We can isolate η too. For C greater than H, it must exceed C/[2(C − H)]. With C = 12 ms and H = 0.92 ms, the threshold is about 0.54152. At C = 1 ms it becomes 6.25, impossible under our efficiency-at-most-one assumption. Removing all local inefficiency would still not help. These thresholds describe the same budget from different perspectives, exposing optimisations that could not change the decision even if perfect.
Answering the initial question
Two devices accelerate the same phase only when saved computation exceeds added costs. Our declared assumptions make this testable: divisible computation must exceed 2.45333 ms; at 12 ms speedup is about 1.388. Bandwidth, startup, efficiency and dependencies change the threshold. Splitting may be needed to fit a model in memory without accelerating it, a different objective. Infrastructure choices should distinguish capacity, response time and cost, then measure the path that matters to users.
References and reproducibility
PyTorch — Distributed communication package: all_reduce and profiling collective communication.
from experiment import run
r = run()
print(r['matrix']['full'])
for case in r['cases']:
print(round(case['one_ms'], 6), round(case['two_ms'], 6),
round(case['speedup'], 6))
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.

