ELAI S.r.l.

8-bit AI: when the sum no longer fits the accumulator

Small values do not guarantee safe sums. Deriving INT32 bounds for an INT8 dot product, including zero points, bias and intermediate sums.

8-bit AI: when the sum no longer fits the accumulator

The model is small—but is the result?

Imagine moving a neural network to an embedded device, a computer inside a machine or sensor. Weights and activations use 8-bit integers, so each number takes little space. A neuron, however, multiplies many pairs and adds their products. Where does the sum go? Into an accumulator, a variable that must hold much larger results than individual inputs. Our question is how to determine, before execution, whether that variable has enough range.

Abstract. Start with a quantized dot product: a sum of products of integer codes representing real values. Derive a conservative signed 32-bit accumulator bound and exceed it using valid 8-bit inputs. Distinguish rounding, saturation and overflow, then show how bias, actual weights and operation order change the analysis. Results are exact integer calculations and explicit arithmetic models executed in Python, not microcontroller measurements or TensorFlow Lite backend tests.

The integer code is not yet the value

Affine quantization represents real x as x≈s(q−z): q is the integer code, z the code for real zero, and s a positive scale. With z=−128, q=127 lies 255 steps from zero. Subtracting two individually valid 8-bit values can therefore produce a value outside signed 8-bit range. Widen the type before subtraction. Compact storage does not imply every operation can retain the same width.

Use activations qₐ from −128 to 127, weights q𝑤 from −127 to 127, and weight zero point zero, consistent with the consulted TensorFlow Lite INT8 specification. Activations are intermediate network values; weights are learned coefficients. For one output channel with scales sₐ and s𝑤, integer A collects the sum before output-format conversion.

A = Σᵢ₌₁ᴷ (qₐ,ᵢ − zₐ) q𝑤,ᵢ + b_int y ≈ sₐ s𝑤 A

K counts products contributing to this output. b_int is the neuron’s constant bias expressed at product scale sₐs𝑤. A, codes and integer bias are dimensionless counts; y has the units established by model and scales. A different channel weight scale changes its bias scale too. We study A; later INT8 conversion may add rounding and clipping, which cannot repair an already wrong sum.

How large can a sum grow?

A signed 32-bit integer ranges from −2,147,483,648 to 2,147,483,647. For an initial guarantee, ignore cancellation and bound each term’s absolute value. |qₐ−zₐ| can reach 255 and |q𝑤| 127, so a product has absolute value at most 32,385. The triangle inequality says a sum’s absolute value cannot exceed the sum of absolute values.

|A| ≤ K·32385 + |b_int| K ≤ floor((2147483647 − |b_int|) / 32385)

The second line is sufficient, not necessary, and assumes nonnegative headroom after bias. Using the positive limit for both signs conservatively forgoes the one extra negative value. With zero bias, K≤66,311 guarantees this direct sum and its partial sums fit for every allowed input. K=66,312 does not always fail; this universal bound simply no longer guarantees it.

A case that crosses the limit

Set every activation code to 127, zero point to −128 and every weight to 127. These valid values produce 32,385 per term. At 65,536 terms the sum is 2,122,383,360 and fits INT32. At 66,311 it is 2,147,481,735; one more product gives 2,147,514,120, above the maximum. This is not gradual loss of decimal precision but inability to represent the integer result in the chosen type.

KExact sumModeled wraparoundModeled saturation
65536212238336021223833602122383360
66311214748173521474817352147481735
663122147514120−21474531762147483647
700002266950000−20280172962147483647

The table compares three explicitly calculated arithmetics. Python retains the exact integer. The wraparound model retains the low 32 bits and interprets them as signed, potentially turning a large positive sum negative. Saturation clamps at the representable maximum. Neither recovers the exact sum. This does not assert either behavior for a particular embedded kernel; its contract, instructions and implementation require verification.

Exact sum and modeled INT32 wraparound as K grows. The vertical axis uses billions of integer counts, not time or energy. The discontinuity is arithmetic, not a board-performance measurement.
Exact sum and modeled INT32 wraparound as K grows. The vertical axis uses billions of integer counts, not time or energy. The discontinuity is arithmetic, not a board-performance measurement.

In C++, do not assume an out-of-range ordinary signed-integer sum wraps: the consulted working draft defines out-of-range arithmetic-expression behavior as undefined. Accelerator-specific instructions may have different rules. Our program therefore uses Python integers and an explicit modular formula, rather than provoking native overflow and mistaking one observed result for a portable rule.

Bias and weights change the headroom

A bias of 2,000,000,000 fits INT32 but consumes most positive headroom. Our extreme case then permits only 4,554 products under the same bound. Checking bias and product types separately is insufficient; bound their sum. Conversely, activation zero point zero lowers the maximum absolute product to 128·127=16,256, giving 132,104 guaranteed terms without bias. Layer geometry and quantization parameters must be analyzed together.

The global bound is often overly conservative for known weights. For fixed wᵢ and activation code between lᵢ and uᵢ, calculate both endpoint products and select minimum and maximum. A negative weight reverses which endpoint is larger. Summing minima and maxima encloses every possible sum, even if some inputs cannot occur together. This guarantee depends on valid intervals; it does not predict error frequency.

Lᵢ = min((lᵢ−zₐ)wᵢ, (uᵢ−zₐ)wᵢ) Uᵢ = max((lᵢ−zₐ)wᵢ, (uᵢ−zₐ)wᵢ) b_int + Σᵢ Lᵢ ≤ A ≤ b_int + Σᵢ Uᵢ

With weights (127,−127,64,−64), zₐ=−128 and the full INT8 range, product intervals are [0,32385], [−32385,0], [0,16320], [−16320,0]. Without bias A lies in [−48705,48705]. The global bound gives |A|≤129540: valid but less precise. This uses actual coefficients and all allowed inputs, not observed averages or hoped-for cancellation. Apply the same calculations to the actual program’s partial sums too.

A valid final sum does not protect intermediate steps

A kernel may rewrite A as Σqₐq𝑤−zₐΣq𝑤+b_int. Algebraically equivalent expressions still require storage for their intermediate sums. With K=150,000, qₐ=zₐ=−128 and q𝑤=127, every centered term is zero and A=0. Expanded contributions are −2,438,400,000 and +2,438,400,000, neither fitting INT32. This demonstrates no particular library defect; it shows why verification must follow the implemented arithmetic path, including precomputed corrections.

Saturating every addition also changes the problem. In a small 8-bit example, 100+100−100 equals exactly 100. Saturating each step produces 100,127,27; reordering to 100−100+100 returns 100. This demonstrates order dependence in saturating arithmetic, not a recommendation for 8-bit network accumulation. Pure modular addition and subtraction instead preserve the result modulo the base: wraparound must not be assigned saturation’s behavior. The rules differ.

Which alternative solves which problem?

Widening the accumulator to 64 bits adds headroom if products, corrections and intermediate conversions also use appropriate types. No cycle or energy cost follows without platform measurements. Blocking a dot product limits local sums, but blocks must still be combined without overflow. Rescaling and rounding each block first introduces a different error requiring separate analysis. Changing precision, scales or architecture may help but changes the numerical tradeoff and requires model-quality evaluation.

A future device check has a concrete protocol: identify layer and channel; extract actual weights, bias, scales and zero points; reconstruct K and kernel arithmetic; compare a sufficiently precise reference with ordinary and extreme cases; record library version, compiler, flags and hardware. This analysis executes only the mathematical reference. It reports no device latency, peak RAM, power or energy per inference because none was measured on a board.

The answer: counting data bits is not enough

An INT8 model may require sums much wider than its inputs. Check product count, post-zero-point ranges, bias and every intermediate step. A conservative bound gives a simple guarantee; weight-aware analysis tightens it; kernel inspection establishes whether it describes the program. The aim is to turn a vague precision question into verifiable pre-release conditions, not to cast suspicion on every quantized network.

Sources and reproducible program

TensorFlow — TensorFlow Lite 8-bit quantization specification, definitions and CONV_2D requirements.

C++ working draft — Expressions, arithmetic range and evaluation rules.

Sources define representation and semantics; bounds, counterexamples and figures were derived and executed for this educational monograph. The code finds the last guaranteed K and compares three arithmetic rules. It is deterministic, requiring no seed. The archive also includes bias, weight-wise intervals, expanded-form cancellation and saturation examples. These are neither EL-AI embedded-product results nor original peer-reviewed research.

LIMIT = 2**31 - 1
product = (127 - (-128)) * 127
safe_k = LIMIT // product
for k in (safe_k, safe_k + 1, 70000):
    exact = k * product  # Python arbitrary-precision integer
    wrapped = (exact + 2**31) % 2**32 - 2**31
    saturated = min(LIMIT, max(-2**31, exact))
    print(k, exact, wrapped, saturated)
# Explicit mathematical models, NOT execution of an embedded kernel.

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.