ELAI S.r.l.

AI on microcontrollers: why one rounding step can change the answer

From INT32 accumulation to INT8 output: a complete example of scales, zero points, rounding and saturation, with executed code.

AI on microcontrollers: why one rounding step can change the answer

The same model, a different answer

An intelligent sensor on a machine must turn a few signals into a score. On the development computer the result appears correct, but after porting to a microcontroller some values differ by one integer unit. Is this harmless noise or a bug? That depends on what the integer represents. Near a threshold, one step of the numerical representation can change a decision. Before changing the neural network, it helps to understand the step that returns a wide sum to the small output format.

We study requantization: rescaling followed by rounding and limiting to the representable range. We derive a three-input neuron, execute the calculation with exact fractions and construct counterexamples for halfway, negative and saturated values. The goal is not to prove INT8 unreliable: it is to expose the numerical contract two implementations must share. We used no board, measured no latency and evaluated no complete network; the code is an educational reference, not a bit-exact emulator of a commercial runtime.

An integer is a label on a scale

INT8 is a signed eight-bit integer: in two’s complement it represents −128 through 127. In a quantized network that number need not equal the mathematical value of the quantity. A positive scale s and zero point z are also needed. Scale specifies one step’s size; z identifies the integer corresponding to real zero. Think of numbered marks on a shifted ruler: the mark number alone does not determine length. The analogy explains conversion, not the network’s operation.

r = s(q − z)

q is the stored integer and r the reconstructed value. All signals in our example are normalised and dimensionless: we are not directly converting volts or degrees. With s = 0.1 and z = 10, q = 13 means r = 0.3; q = 10 means zero; q = 0 means −1. This matters when adding zero-valued image borders: mathematical zero must be encoded as z, not necessarily the byte whose numerical value is zero.

Build the full calculation before rounding

The neuron sums three products and adds a bias, a constant shifting the result. Integer inputs [13, 9, 12], with sₓ = 0.1 and zₓ = 10, represent [0.3, −0.1, 0.2]. Integer weights [2, −3, 1], with s𝑤 = 0.2 and z𝑤 = 0, represent [0.4, −0.6, 0.2]. Real bias is 0.02. We chose numbers represented exactly on their grids to isolate final rounding without mixing it with earlier input quantization error.

Real arithmetic gives 0.3 × 0.4 + (−0.1) × (−0.6) + 0.2 × 0.2 + 0.02 = 0.24. To obtain the same value from integers, first subtract the input zero point. Products become 3 × 2, (−1) × (−3), and 2 × 1: 6, 3 and 2. Each unit of their sum represents sₓs𝑤 = 0.02. Bias must use that same scale: bq = 0.02/0.02 = 1. Adding 6 + 3 + 2 + 1 gives accumulator A = 12, which still reconstructs 12 × 0.02 = 0.24.

A = Σᵢ (qxᵢ − zₓ)qwᵢ + bq; bq = b/(sₓs𝑤) y = sₓs𝑤 A

The relation assumes zero-centred weights and a bias representable at the product scale; otherwise bias needs its own specified rounding. A is an integer: a typical pipeline accumulates in INT32 to provide more range than individual INT8 operands. This article assumes accumulation does not overflow. Sufficient bit width is necessary but does not resolve subsequent rescaling issues.

From the sum to the output byte

Output uses sᵧ = 0.08 and zᵧ = −7, so it has a different ruler. To express A in output steps, divide its real value by sᵧ. Ratio M = sₓs𝑤/sᵧ = 1/4 converts the scales. Define R as rounding to nearest, with exact ties away from zero. Define clamp as limiting between −128 and 127. Our contract, in the stated order, is:

M = sₓs𝑤/sᵧ qy = clamp(R(MA) + zᵧ, −128, 127) ŷ = sᵧ(qy − zᵧ)

For A = 12, MA = 3 is already integral. Adding −7 gives stored qy = −4. That negative byte represents a positive result: ŷ = 0.08 × (−4 + 7) = 0.24. The code’s sign is therefore not the quantity’s sign. Adding bias directly at the output scale would not generally be equivalent either: here bias belongs to the accumulator before conversion.

Halfway between values: a complete rule is necessary

Now change only A to 10. Before rounding, MA = 2.5. Integers 2 and 3 are equally close. Our rule chooses 3, while ties-to-even chooses 2 because it is even. Adding zᵧ gives −4 and −5, reconstructing 0.24 and 0.16. Before requantization the value was 0.20: both are 0.04 away, but they are different results.

If a hypothetical decision triggers at ŷ ≥ 0.20, the first version triggers and the second does not. We have not measured accuracy loss on a dataset: this counterexample demonstrates why a one-code difference can matter. No rounding rule is universally best for every network. The compatibility issue is validating with a different rule from the device, or ignoring the tolerance required by the decision.

Order also matters. With our rule, R(2.5) − 7 = −4, but R(2.5 − 7) = R(−4.5) = −5. Moving zero-point addition before rounding looks like harmless real-number algebra, but changes a tie relative to zero. The compact formula must specify where rounding occurs. If a runtime uses another convention, its comparison reference must reproduce that convention; we do not claim all runtimes use R.

Negative numbers, truncation and distortion

Discarding the fractional part is not always rounding to nearest. With A = −11, MA = −2.75: nearest is −3, truncation toward zero gives −2, and rounding toward negative infinity gives −3. With A = 11, truncation and rounding toward negative infinity both give 2, while nearest gives 3. The differences are not symmetric for every rule. The table reports final codes after adding zᵧ = −7; interpreting them as values still requires subtracting zᵧ and multiplying by 0.08.

AMAR: ties awayTies to evenTruncationToward −∞
-11−2.75-10-10-9-10
-10−2.5-10-9-9-10
00-7-7-7-7
102.5-4-5-5-5
112.75-4-4-5-5

The plot evaluates every integer A from −16 through 16, without saturation, showing reconstructed error ŷ − y. Lines between points aid reading; they do not represent tests on continuous inputs. Truncation pushes values toward real zero; rounding toward negative infinity produces only nonpositive errors. On this symmetric grid, mean error for R and truncation is zero, but individual errors still differ, and this proves no absence of bias on nonsymmetric real data.

Error on identical synthetic accumulators: M = 1/4, sᵧ = 0.08, zᵧ = −7. Grey lines mark ±0.04, half an output step; nearest rounding stays within them without saturation.
Error on identical synthetic accumulators: M = 1/4, sᵧ = 0.08, zᵧ = −7. Grey lines mark ±0.04, half an output step; nearest rounding stays within them without saturation.

Two algebraically exact steps, two different roundings

A device may implement a fractional multiplier with integer multiplication and bit shifts. Understanding the risk does not require emulating a particular processor. Take a second example, M = 3/8, with two specified algorithms: one rounds only at the end; the other computes 3A/4, rounds, divides by two and rounds again. Without rounding, 3/8 equals (3/4)/2. For A = 1 the first gives R(0.375) = 0; the second gives R(R(0.75)/2) = R(0.5) = 1.

This is not a claim of a library defect: we constructed two different numerical contracts. Some implementations intentionally use different, validated intermediate steps. Comparing results requires knowing multiplier, shift, intermediate widths and rounding rules at every stage. Checking only final real multiplier M is insufficient. The package evaluates both algorithms for every A from −16 through 16, including negatives.

What error can we bound?

With one nearest rounding, an exact multiplier and no saturation, the distance between R(MA) and MA is at most half an integer. Multiplying by the output scale gives a local reconstructed-value bound:

|ŷ − y| ≤ sᵧ/2 = 0.04

This proves a property of this step, not error relative to an original floating-point model: weights, inputs and biases upstream may already be approximate. It is not a guarantee on a many-layer final decision either. Later operations may amplify the half-step; near a threshold it may already suffice to cross it. Whole-network analysis requires the complete operation path and representative data.

If the device uses approximate multiplier M̂ without further intermediate roundings, the triangle inequality separates two contributions: |ŷ − y| ≤ sᵧ(1/2 + |A|·|M̂ − M|), still without clipping. The first is final rounding, the second multiplier error amplified by the accumulator. A small relative constant error therefore does not justify ignoring A’s range. The bound excludes overflow and additional hidden roundings.

Saturation breaks the half-step guarantee

Increase A to 600 while keeping the original scales. Real value is 12, MA = 150 and the code before limiting is 150 − 7 = 143. INT8 cannot hold it: clamp returns 127. Reconstruction is 0.08 × (127 + 7) = 10.72, an error of −1.28. This is not rounding error: the representation boundary has been reached. The 0.04 guarantee does not apply because its no-saturation assumption is violated.

Changing sᵧ to widen the range also widens the spacing between consecutive values. Range and resolution must therefore be chosen together using data appropriate to intended use. An activation function may further restrict the interval: for ReLU, which zeroes negative values, the zero boundary’s code is zᵧ, here −7, not necessarily 0. Numerical saturation and activation are conceptually different even when a kernel fuses them.

From sources to device verification

Jacob and colleagues describe affine quantization, bias scale and the integer pipeline in their 2018 work, whose sections 2 and 4 we read. Their MobileNet comparisons concern Snapdragon hardware and specified protocols; we do not transfer those results to a microcontroller. gemmlowp documentation explains accumulator conversion, while LiteRT’s specification distinguishes operator types, ranges and zero points. These references explain the mechanism; our small examples are separate from their benchmarks and certify no specific-version compatibility.

Python’s Fraction keeps 1/10, 1/5 and 2/25 exact during calculation: even writing 0.1 as a float may introduce a binary approximation, an unnecessary variable here. Function away rounds the absolute value with integer arithmetic, then restores its sign; requant applies multiplier, rounding, zero point and saturation in order. Assertions check the full example, ties, double rounding and the half-step bound over the complete grid. Data are deterministic, with no seed.

For L inputs, the dot product needs O(L) operations; converting N available accumulators needs O(N) steps. Script memory grows with saved rows, which does not measure firmware peak RAM. Python uses arbitrary-precision integers: an embedded implementation must explicitly define intermediate product widths, valid shifts, saturation and casts. Translating educational code line by line does not establish that a compiler preserves every property.

Proposed device verification starts with shared reference vectors: real zero, negatives, halfway values, saturation boundaries and channels with different scales. Record converted model, converter version, runtime, kernel, compiler and options. If bit-exact equality is not expected, document per-operator tolerance and separately evaluate decision impact. Only after correctness should mean and percentile latency, peak RAM, power and energy per inference be measured with an explicit hardware protocol. Those tests remain unperformed: our calculation yields none of those measurements.

The answer: check what the byte means

A correct INT32 sum does not by itself guarantee equivalent INT8 output across two systems. Preserve scales and zero points, specify tie handling, respect operation order and distinguish rounding from saturation. In our example accumulator 10 reconstructs either 0.24 or 0.16 depending on the rule; at 600 the format boundary introduces much greater error. The practical lesson for small-device AI is to verify what numerical operation the model actually performs before asking how fast it runs.

Sources and reproducibility

Benoit Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (2018), sections 2 and 4.

Google gemmlowp, Building a quantization paradigm from first principles.

Google LiteRT, 8-bit quantization specification.

from fractions import Fraction
from experiment import requant, run
print(requant(10), requant(10, rounder=round))
print(requant(600))
print(next(r for r in run()['double_rounding'] if r['a'] == 1))

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.