ELAI S.r.l.

The model is updating and power fails: which version restarts?

A sensor, two model copies, and a reproducible experiment: separating download, verification, and activation while distinguishing simulation from hardware.

The model is updating and power fails: which version restarts?

The problem: an update is not yet a usable model

An industrial sensor reads temperature and vibration, then uses a small artificial intelligence model to classify a machine’s state. A new version of its weights, the numbers learned during training, arrives. While the device saves them, someone disconnects its power. When power returns, the sensor must decide what to load. The problem is not simply finishing the transfer: it is preventing an incomplete copy from being mistaken for a ready model, while avoiding destruction of the only working copy.

The article asks a concrete question: can we arrange the update so that every restart finds either the intact old version or the intact new version? We will build a state machine: a description of operations and the persistent states they leave behind. We will execute every interruption point included in that model, comparing direct overwrite with two separate areas and a final activation step. The result will be a property conditional on explicit assumptions, not a reliability measurement of a commercial board.

Before the details: three meanings of “ready”

Received means that the transfer has delivered some data. Verified means that those data satisfy defined checks, such as expected size, a cryptographic digest, and compatibility with the program that executes the model. Active means that the startup procedure is allowed to select them. These three events need not coincide. A completed-download notification does not prove that the data are already persistent; successful verification does not require immediate activation. Separating these events lets us reason about what happens between them.

In our scenario, the inference program, called the runtime, stays installed: we update the weights as data. Flash memory retains information without power; RAM normally does not. We will call a memory region reserved for one copy a slot. An activation log, or journal, stores small descriptions of selectable copies. Think of it as an index of available editions of a book, but the analogy ends there: real memory imposes erase blocks, programming constraints, and potentially interrupted writes that a paper index does not represent.

The contract that makes the reasoning possible

We assume durable, ordered writes: if the program completes one step before starting the next, that step remains after restart. We also assume an atomic final marker: after restart it is either unset or set, never ambiguously accepted. The two slots are independent, the confirmed old copy stays intact, and metadata and the expected digest come from a trusted source. The generation counter increases without overflow. These are preconditions of the mathematical model; a hardware project must establish how to implement them or change the protocol when they do not hold.

We introduce interruptions only between completed abstract operations. We do not simulate the flash circuit as voltage drops, a half-completed write, an unflushed cache, or damage to a neighboring sector. This distinction is decisive: enumerating every state in our Python program does not mean testing every microcontroller failure. The small experiment answers a logical question about the protocol. Electrical, timing, and memory-endurance tests remain a second verification layer, necessary before using the design on a device.

The minimal case: four parts and one copy

We reduce the old model to four bytes [1, 1, 1, 1] and the new model to [2, 2, 2, 2]. These are not useful inference weights: they represent four file parts, making completeness visible. The naive program erases the slot and then writes one part at a time. It has five operations and six observable interruption points, including the point before the first operation. Before erasure the old model exists; after all writes the new one exists. At the four intermediate points, neither complete copy exists.

The sequence [2, 2, empty, empty] does not become a correct model merely because transfer will resume later. At restart, necessary material is already missing. We can stop inference and wait for recovery, sometimes an acceptable choice; we cannot claim operational continuity. Writing a “version 2” label first does not solve the problem either: a label does not complete the data. If interpreted as permission to load them, writing it early makes partial-copy selection easier. The useful property concerns both the data and the selection procedure.

Two copies and a decision made at the end

We retain the confirmed copy in slot A and prepare the candidate in slot B. The eight operations are: erase B, write its four parts, verify its digest, append a still-inactive journal record, and set the commit marker. Commit means making the decision visible after restart. The record contains a generation number, slot name, required runtime version, and expected digest. The generation number orders activations: it does not measure model quality and must not be confused with a scientific performance assessment.

At restart we read records from newest to oldest. We accept the first that is committed, compatible, and linked to complete data with the correct digest. We repeat verification on the data actually present: remembering that the download was checked before interruption is insufficient. If the candidate fails, we try the previous record. If none passes, the result is “no valid model.” This explicit outcome prevents the procedure from converting the absence of a usable copy into arbitrary loading that appears successful.

eligible(r) = committed(r) AND compatible(r) AND complete(r) AND hash_ok(r) selected = first eligible record in descending generation order

The formula answers “which copies may be selected?” The letter r denotes a record; AND requires every condition to hold. complete checks that no parts are missing, while hash_ok compares the computed and expected digests. In the code, both checks are collected in valid. There are no physical units in this expression: these are logical predicates. For example, a committed record requiring runtime 2 is ineligible on a runtime-1 device, even when every byte matches the received data.

What the experiment actually shows

The figure collects the execution of both programs. Each point is a restart after a given number of completed operations. In the upper panel, red points indicate an unusable copy; in the lower panel, the previous version remains selected until final commit. The horizontal scales count different operations in the two protocols, not seconds: point 5 in the first panel is not the same instant as point 5 in the second. The comparison concerns possible selection outcomes, not update speed.

Interruptions between abstract steps: old copy, new copy, or no valid copy. Points represent neither probabilities nor electrical power-cut tests.
Interruptions between abstract steps: old copy, new copy, or no valid copy. Points represent neither probabilities nor electrical power-cut tests.
ProtocolPoints examinedOldNewNone valid
A: overwrite6114
B: dual + commit9810

Four problematic points out of six do not mean a 66.7% failure probability. We have not assigned a temporal distribution to interruptions, measured operation durations, or modeled supply voltage. Erasing a block and writing a byte need not take the same time. Likewise, zero invalid states among the nine examined is not a statistical estimate of zero risk. It exhaustively checks only the cut points admitted by our abstraction. Confusing these readings would turn a useful logical result into a reliability promise unsupported by data.

Why it works: preserving an invariant

An invariant is a property that remains true after every allowed step. Here it is: before commit, at least one selectable copy exists, namely A. It is initially true by construction. Erasing and writing B do not change A because the slots are independent. Verifying B does not change A. Adding an inactive record does not make B selectable. This gives induction over the steps: if the property holds before each pre-commit operation, it also holds afterward. Restarting at any such point returns to A.

When the atomic marker becomes active, B is already complete and verified; its newer record makes it preferable. If startup checks instead detect an unusable candidate, A remains available. Order is essential: activating the record before completing the data destroys the argument. Independence is equally essential: if erasing B physically erases part of A because both regions share a sector, the invariant fails. Drawing two rectangles in a memory map is insufficient; the actual granularity of destructive operations must be respected.

Three negative tests, including one that must fail

The program also executes three cases separate from interruption. It corrupts one byte of the committed candidate: the digest no longer matches and A restarts. It changes the candidate record’s required runtime from 1 to 2: the content is intact but incompatible and A restarts. Finally, it corrupts B and erases A: selection returns None, meaning no usable copy. The last test is as important as the first two. It exposes the property’s boundary: two slots do not protect against losing both copies, and the procedure must not hide that condition.

The digest is SHA-256, a function that summarizes bytes into a fixed-length sequence. In our example it detects the intentional test-data modification, but does not establish who supplied the model. If an adversary can replace both the file and its expected digest, comparison may succeed for an unauthorized file. Package authenticity, key protection, signature verification, and update authorization require a separate design. Here we assume trusted metadata and do not perform a security evaluation of the distribution channel.

How much storage does a way back cost?

We move from four educational bytes to a hypothetical sizing example, without attributing it to a board. A model occupies W = 1,048,576 bytes, or 1 MiB; each copy has an H = 256-byte header. Erasure uses E = 4,096-byte blocks, and we reserve J = 8,192 bytes for the journal. To prevent an erase operation from crossing between copies, we round each slot up to the next multiple of E. The following formula answers a precise physical question: how many flash bytes must this scheme reserve?

S = E × ceil((W + H) / E) F = 2S + J S = 1,052,672 bytes = 1,028 KiB F = 2,113,536 bytes = 2,064 KiB = 2.015625 MiB

S is the reserved size of one slot and F the total; ceil means rounding upward to an integer. KiB is 1,024 bytes and MiB is 1,048,576 bytes. The executed calculation returns 2,113,536 bytes: already 16,384 bytes beyond a 2 MiB memory, before adding runtime, startup code, or other data. Saying “the model is one megabyte, so two megabytes suffice” would be wrong even in this minimal example. The total concerns persistent storage, not peak inference RAM, which also depends on activations, buffers, and operator implementations.

Integrity is not full compatibility or quality

Our compatibility check uses a single runtime integer: it is intentionally elementary. In an AI application, a weights file may require specific input shapes, channel order, sensor units, normalization, operators, and label mappings. Changing normalization while retaining old weights can produce a system that starts normally but makes bad decisions. The package manifest should identify a coherent set of these dependencies. Restoring weights alone is not complete rollback if the rest of the configuration has already changed incompatibly.

Our program’s commit makes B selectable; it does not implement a trial period followed by confirmation. A practical extension could distinguish candidate, trial startup, and confirmed version, using a health check and a limit on attempts. It would then also need to define what happens if power fails during confirmation. Passing a startup test does not prove that the model will be accurate on future data. Startup continuity, interface correctness, and predictive validity are different properties: none follows automatically from the other two.

Real systems and alternatives

Official MCUboot documentation describes firmware-image updates with trial, confirmation, and revert modes, plus persistent information for recovering interrupted swaps. It helps illustrate the gap between an abstract rule and a configurable bootloader. Our code does not execute MCUboot, reproduce its algorithm, or update firmware.

Alternatives depend on the dominant constraint. One slot with network recovery saves local space but accepts an interval without inference and depends on recovery availability. Two slots retain a ready copy at the cost of extra storage. A differential update transfers only changes, but fewer transmitted bytes do not guarantee that applying them can be interrupted without damage: a strategy for preserving the previous state is still needed. A temporary file followed by name replacement shifts part of the problem to the filesystem contract, including persistence after power loss.

The journal also needs an endurance-aware design. In the code it is a Python list sorted at each startup, and the counter does not overflow. Real storage is finite: deleting old records, recycling sectors, and managing wear introduce further transitions to analyze. With R records, the educational sort costs O(R log R); verifying W bytes requires O(W) hashing work and fallback to A may require reading two copies. These complexities do not provide milliseconds or millijoules: platform measurements are needed to turn them into latency and energy.

How to reproduce it and what to test on a device

The attached archive contains experiment.py, plot.py, JSON results, and instructions in all four languages. The experiment is deterministic, so it uses no random seed. Running experiment.py enumerates states and automatically checks the stated properties; plot.py draws outcomes rather than measuring hardware. The short code below invokes the same calculation and prints point/outcome pairs, the three negative cases, and total bytes. In the full file, the decisive loading condition accepts a record only after commit, compatibility, and copy verification: this is where the explanation becomes an executable rule.

On hardware we would instead propose controlled interruptions during erasure, programming, metadata updates, and confirmation, not just between software functions. Board, flash, versions, configuration, voltages, and protocol would need specification, with checks for corrupt candidates, wrong dependencies, lost journals, and failed trial startup. This is an unexecuted verification plan. We have no restart-time, energy, wear, or failure-probability measurements here. Nor do we attribute this experiment to an EL-AI embedded product: it is an educational analysis, not documentation of a company installation.

The answer: protect the copy before choosing the new one

If power fails, our protocol restarts A until commit and B after commit provided the candidate passes checks; otherwise it returns to A if still valid. The reason is not a special capability of artificial intelligence: we preserved an intact copy and delayed choosing another until its data were ready. This clarifies what to demand from an embedded update: a verifiable selection rule, a preserved previous state, and implementable memory assumptions. The simulation demonstrates the reasoning under those assumptions; the device still has to demonstrate that it meets them.

Technical source and reproducible material

MCUboot — Bootloader design.

from experiment import run
r = run()
for name in ['naive', 'dual']:
    print(name, [(x['cut'], x['outcome']) for x in r[name]])
print(r['negative_cases'])
print(r['memory']['total_flash_bytes'])

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 6 October 2026.