ELAI S.r.l.

An AI model activates few experts: why can it still become overloaded?

Sixteen tokens explain routing, local capacity and balance in MoE models: reproducible calculations, strategy comparisons and critical reading of research.

An AI model activates few experts: why can it still become overloaded?

The paradox of idle resources

Imagine an assistant reading many technical documents together. Its model contains several specialized modules and uses only one for each text fragment. This seems a natural way to save computation. Yet one part of the system fills up while others remain almost empty. Total available resources do not explain the problem: we need to know where requests go. How can a model activate little computation overall while overloading one specific destination?

Abstract. We study a Mixture of Experts, or MoE, layer selecting one expert per token. A synthetic sixteen-token example derives capacity, unprocessed assignments and unused space. We separate router probabilities, discrete assignments, balance over time and physical accelerator placement. We compare the historical Switch approach with developments documented in DeepSeek-V3 and the 2026 LLEP preprint, without transferring their benchmarks to our example. Our results are executed Python counts: no trained network and no GPU measurements.

An expert is not a person; a token is not a request

A token is a unit into which text is segmented: a word, fragment or symbol. In our layer, an expert is a feed-forward network, a block of numerical transformations applied to the token representation. Experts have different parameters; they are not agents debating with one another and need not correspond to professions such as doctor or engineer. The term refers to possible learned specialization, not certified competence.

The router scores modules and chooses which to use. Top-1 selects the highest score; top-k selects k. Sparsity concerns expert activation per token, not an otherwise inactive model: attention, routing and other components still operate. Sixteen tokens in a batch can also belong to one conversation or several documents. Routed-token counts are not counts of users served.

Constructing a case we can inspect

Take T=16 tokens and E=4 experts. The router prefers expert 1 for twelve tokens, expert 2 for two, and experts 3 and 4 for one each. Thus n=[12,2,1,1], where n_i counts assignments to expert i. The sum is sixteen because every token is assigned once. This invented batch exposes the mechanism; it is not a trace from a commercial model. Mean load is four and maximum load twelve, three times the mean.

Whether all assignments can be processed depends on local capacity C. Assume equally sized rigid buffers: each expert admits at most C assignments in the current pass. Excess assignments do not automatically move to a different expert. Four service counters suggest the queueing intuition, but the analogy has a limit: experts have different weights, so changing the selected module can change the computed function.

Why sixteen slots do not suffice for sixteen tokens

Let c be a dimensionless capacity factor multiplying mean load. Our simulator rounds up: C=ceil(cT/E). At c=1, C=4 and total capacity EC=16. Total capacity matches demand but its distribution does not: expert 1 admits only four of its twelve assignments. Others admit 2, 1 and 1. Eight assignments are processed and eight exceed capacity.

C = ceil(c × T / E) A = Σ_i min(n_i, C) D = Σ_i max(n_i − C, 0) = T − A U = E × C − A

A counts processed assignments, D excess assignments and U unused slots. These are counts, not seconds, watts or bytes. The second formula sums what each expert admits, not what the system could admit if slots were interchangeable. At c=1, A=8, D=8 and U=8: unmet demand coexists with idle space. This quantitatively resolves the opening paradox.

cC per expertD excessTotal slotsU unused
148168
1.25572011
1.5662414
2843220
2.51024026
31204832

The table varies c while holding routing fixed. Doubling capacity gives C=8 but leaves four excess assignments. This batch needs C=12, reached at c=3, to eliminate them: forty-eight slots for sixteen assignments. Real software need not use exactly three times the memory or time; buffer representation and kernels matter. This is a rigid-allocation count, not an accelerator measurement.

Executed synthetic example: with routing fixed, increasing capacity factor reduces excess assignments while increasing reserved slots. The line at 16 is requested assignments, not a memory limit. Points are six computed configurations; connecting lines are guides only.
Executed synthetic example: with routing fixed, increasing capacity factor reduces excess assignments while increasing reserved slots. The line at 16 is requested assignments, not a memory limit. Points are six computed configurations; connecting lines are guides only.

Dropping a token does not mean deleting a word

Within an MoE layer, token dropping can mean skipping an expert contribution when capacity is exhausted. In the Switch setup we read, the representation continues through the residual connection. No word is deleted from the document. The model can still answer, but that token missed its intended transformation at that layer. We count missing contributions without computing their effect on language quality.

Other implementations process all assignments with dynamic shapes, grouping or different scheduling. Then D is not actual runtime behavior: it counts assignments that would exceed hypothetical C. The problem may reappear as memory use, waiting or work on the busiest device. A reported dropped-token percentage therefore requires the underlying policy. Identical load distributions can have different consequences in different systems.

Nearly uniform probabilities, highly uneven decisions

Give each token probability .28 for its winner and .24 for each other expert; the sum is one. The preference is weak, but top-1 still chooses one winner rather than splitting the token four ways. Let f_i=n_i/T be actual assignment fraction and P_i mean router probability in the batch. Here f=[.75,.125,.0625,.0625], whereas P=[.27,.245,.2425,.2425]. Confusing these distributions hides overload.

f_i = n_i / T P_i = (1/T) × Σ_t p_i(t) B = E × Σ_i f_i P_i

B is a Switch-style balance term before its training-objective coefficient. Uniform assignments and probabilities give one. Our numbers give B=1.05375 even though half the assignments exceed C=4. This does not make the term useless; its value is not a count of missing tokens or a hard per-batch constraint. We also do not claim that one is a global lower bound for every admissible f and P pair.

The router learns through continuous scores while the winning choice changes discretely. Statistical regularization and capacity guarantees therefore answer different questions: one guides learning; the other must specify every assignment’s treatment. Increasing balance weight arbitrarily can shift the objective from task quality toward uniformity. Evaluating that tradeoff requires controlled training runs, which we did not perform.

A daily average can hide every local peak

Change experiments to four successive eight-token windows: all tokens go to expert 1, then 2, then 3, then 4. Across the whole interval each expert receives eight; the aggregate histogram looks perfectly balanced. Yet at c=1 every window has C=8/4=2. Six assignments exceed capacity each time, totaling twenty-four of thirty-two. Aggregate balance does not reveal the peak relevant to the buffer.

This does not show short windows are always worse or that every document should use every expert. It demonstrates information lost through aggregation. Real measurements need layer-, batch- and device-level loads at a scale matching the constraint. When memory is allocated per micro-batch, a daily average answers a different question from instantaneous exhaustion.

Balancing experts and balancing machines are different

An expert is a logical module; a GPU is a physical device and can host several. Our second count has eight experts with loads [8,8,0,0,4,4,4,4] on four devices, two experts each. Consecutive pairs yield device loads [16,0,8,8]. Placing pairs (1,3), (2,4), (5,6), (7,8) yields [8,8,8,8]. No token changes expert; only execution location changes.

This is an allocation possibility, not a measured speedup. Moving or replicating weights costs memory and communication; outputs must return to their origin, and training gradients must be accumulated correctly. The next batch may have different loads. Physical balancing preserves the router’s semantic choice while adding a scheduling problem. Active parameter count alone therefore determines neither latency, peak memory nor service cost.

What research adds, and what we have not reproduced

Fedus, Zoph and Shazeer’s Switch Transformers is a 2022 JMLR paper, not a 2026 novelty. We read routing, capacity, objective and experiments; Table 1 uses C4 pretraining and 32 TPUv3 cores. Its timings describe neither our counts nor EL-AI infrastructure.

DeepSeek-V3’s technical report, v2 dated 18 February 2025, separates selection bias from combination weights and retains a small sequence auxiliary term. Its two-scale ablations hold data and architecture comparable. We reproduced no training or benchmarks; “loss-free” does not mean every auxiliary loss is absent.

For a current connection we read methods and experiments in Nguyen et al.’s Least-Loaded Expert Parallelism, arXiv v1, 23 January 2026. It redistributes work and weights; controlled tests use eight H200s and distinguish layers from full models. Gains depend on communication, batch size and thresholds; this is not an exhaustive frontier survey.

From reproducible counts to a design decision

The Python snippet repeats the first table: ceil defines rounding, min limits admitted assignments, and subtraction yields excess and idle slots. The full package adds probabilities, balance score, windows and placements. Inputs are deterministic, with no seed and no model-generated tokens. Probability accounting costs O(TE) and stores a T×E matrix; capacity-only counting is O(E) per configuration. These costs describe the teaching script, not an MoE kernel.

A real evaluation must fix checkpoint, layer, top-k, precision, hardware, runtime, batch and document distribution. It should separately measure expert/device loads, peak memory, communication and latency, then assess answer quality and less frequent domains. Comparisons need identical inputs and must include balancing overhead. These are proposed checks, not performed here; the experiment identifies questions to ask before costly measurement.

The answer is concrete: activating few experts per token limits some work without ensuring it reaches available capacity. In our batch, sixteen slots leave eight assignments without an expert contribution; nearly uniform probabilities and global averages do not reveal this alone. Capacity, routing learning and physical placement are different levers with different costs. A business evaluating an assistant should ask not only how many parameters activate, but what workload its own documents create, at what quality and under which constraints.

References and reproducible material

Fedus, Zoph, Shazeer — Switch Transformers, JMLR 23, 2022, sections 2.1–2.4.

DeepSeek-AI et al. — DeepSeek-V3 Technical Report, arXiv v2, 18 February 2025, sections 2.1.2 and 4.5.

Nguyen, Pandit, Xu, Xiong, Joty — Least-Loaded Expert Parallelism, arXiv v1, 23 January 2026, sections 3–5.

from math import ceil
loads = [12, 2, 1, 1]
tokens, experts = sum(loads), len(loads)
for factor in (1, 1.25, 1.5, 2, 2.5, 3):
    capacity = ceil(factor * tokens / experts)
    accepted = sum(min(n, capacity) for n in loads)
    print(factor, capacity, tokens-accepted, experts*capacity-accepted)

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.