ELAI S.r.l.

When AI creates mass from nowhere: correcting predictions without confusing constraints with truth

Three compartments and ten kilograms explain projection, conservation and nonnegativity in scientific models, with derivations and executed code.

When AI creates mass from nowhere: correcting predictions without confusing constraints with truth

A plausible prediction with an impossible total

An AI model replaces part of a simulator: it receives a process state and predicts how a substance will be distributed. Each value looks reasonable in isolation. Added together, however, they give twelve kilograms where only ten were available, with no external inflow. The prediction is not merely inaccurate: it violates information already known. Can we correct it without restarting training? And once the total balances, can we consider it correct?

This article answers through three compartments and a mass constraint. We derive the smallest correction, show why it can produce negative quantities and add a second constraint to prevent them. We then distinguish an actual mathematical guarantee from two promises we cannot make: knowing the true distribution and repairing an incorrect initial balance. All numbers are synthetic and calculated in Python; we trained no network, simulated no fluid and validated no EL-AI plant.

What we actually know about the system

Consider a substance in three control volumes. We use masses in kg, not concentrations: x₁, x₂ and x₃ are the quantities to predict. The system is closed for this substance, with no sources, reactions producing or consuming it, losses or external fluxes. Known total mass is M = 10 kg. We are not claiming every substance in every process is individually conserved: this assumption defines the example. With inflows or outflows, total mass must be updated using their balance.

x₁ + x₂ + x₃ = M = 10 kg; xᵢ ≥ 0

A network minimising average data error can learn many relationships well without exactly satisfying this sum for every new input. A soft constraint adds a training penalty; a hard constraint admits only outputs satisfying the equality. These are different objectives. Even a large penalty may leave residual error because it competes with other objectives and imperfect numerical optimisation.

The smallest correction has a precise definition

Let raw prediction be y = [8, 3, 1] kg. Its total is twelve, an excess of two kilograms. We could remove both from the first compartment or distribute the correction. Choosing requires defining “change as little as possible”. Here we minimise the sum of squared corrections, treating a one-kilogram change equally in every compartment. This is a metric choice, not a physical law.

minₓ ½Σᵢ(xᵢ − yᵢ)²; Σᵢxᵢ = M

This operation is called projection: it finds the feasible point nearest the prediction under the chosen distance. First impose only the sum, leaving nonnegativity for later. Introduce multiplier λ and form L = ½Σᵢ(xᵢ − yᵢ)² + λ(Σᵢxᵢ − M). At the minimum each derivative is zero: xᵢ − yᵢ + λ = 0. The correction is therefore identical across compartments; imposing the total determines its size.

λ = (Σᵢyᵢ − M)/n; xᵢ = yᵢ − λ n = 3; λ = (12 − 10)/3 = 2/3 kg

The result is x = [22/3, 7/3, 1/3] kg, approximately [7.3333, 2.3333, 0.3333]. Total mass is exactly ten in the script’s rational arithmetic. Removing all excess from the first compartment gives [6,3,1], with total squared correction 4 kg²; distributing it uniformly costs 4/3 kg². Hence projection prefers the latter. The quadratic objective is strictly convex and the constraint affine, so the minimum is unique, independent of random search or local minima.

A useful guarantee, distinct from knowing the truth

If truth were t = [6,3,1], the intuitive correction to the first compartment would be perfect, whereas projection would not. Conservation does not reveal where the error occurred. It nevertheless guarantees something precise: if t satisfies the correct total, Euclidean projection cannot increase overall squared distance to t. We can prove this without knowing its three values.

‖y − t‖² = ‖y − x‖² + ‖x − t‖²

Double bars denote Euclidean distance; its square sums squared errors. Vector y − x has every component equal to λ, while x − t sums to zero because both totals equal M. Their dot product is zero, so the cross term disappears when expanding the square. In case A, error to t falls from 4 to 8/3 kg². Two previously exact compartments change: improvement concerns the aggregate error, not each component separately.

Now consider y = [8,1,1]. Its sum is already ten, so correction is zero. Against t = [6,3,1], however, 8 kg² of error remains. A model can conserve mass perfectly while distributing it poorly. Conservation is necessary for our system, not a complete accuracy measure. It also certifies no transfer times, local fluxes, energy or reactions absent from the model.

The total balances, but negative mass appears

In case B the prediction is y = [−1,4,9] kg. Subtracting 2/3 from every entry gives [−5/3,10/3,25/3]. Total mass is ten, but the first value is even more negative. Clipping it to zero afterwards makes the total 35/3, about 11.6667 kg, destroying the constraint just imposed. Sum and nonnegativity must be satisfied together, not repaired sequentially with incompatible operations.

minₓ ½Σᵢ(xᵢ − yᵢ)²; Σᵢxᵢ = M; xᵢ ≥ 0 xᵢ = max(yᵢ − θ, 0); Σᵢmax(yᵢ − θ, 0) = M

θ is a common threshold, measured in kg. Compartments large enough relinquish the same amount θ; those that would become negative stop at zero. In B, fix the first at zero and require the others to total ten: (4 − θ) + (9 − θ) = 10, hence θ = 1.5 kg and x = [0, 2.5, 7.5]. This enforces both conditions simultaneously.

The max form is not unproved intuition. Optimality conditions add a nonnegative multiplier μᵢ for each bound xᵢ ≥ 0. Stationarity gives xᵢ − yᵢ + θ − μᵢ = 0, and complementarity requires μᵢxᵢ = 0. If xᵢ is positive, μᵢ is zero and xᵢ = yᵢ − θ. If xᵢ is zero, μᵢ = θ − yᵢ must be nonnegative. These cases produce the formula. The feasible set remains convex, so these conditions identify the unique global minimum.

CasePrediction (kg)Sum only (kg)Sum + nonnegativity (kg)
A[8; 3; 1][22/3; 7/3; 1/3][22/3; 7/3; 1/3]
B[−1; 4; 9][−5/3; 10/3; 25/3][0; 2.5; 7.5]
C[8; 1; 1][8; 1; 1][8; 1; 1]
Synthetic data in kg. Both projections coincide in A; sum-only projection retains a negative value in B. The second constraint changes the required correction, not merely the bar colours.
Synthetic data in kg. Both projections coincide in A; sum-only projection retains a negative value in B. The second constraint changes the required correction, not merely the bar colours.

The metric determines where correction goes

We treated all three errors equally. With trustworthy independent uncertainties σᵢ², we could minimise ½Σᵢ(xᵢ − yᵢ)²/σᵢ² instead. Correcting a highly uncertain compartment would cost less than correcting a well-known one. With only the sum constraint, derivation gives xᵢ = yᵢ − σᵢ²(Σⱼyⱼ − M)/Σⱼσⱼ². For y = [8,3,1] and hypothetical variances [4,1,1] kg², the result is [20/3,8/3,2/3] kg. We did not estimate these variances from data: they illustrate a different optimisation question.

The distance guarantee must use the projection’s metric: changing weights does not automatically preserve the previous Euclidean result. Correlated uncertainties change the problem through a covariance matrix. Switching from masses to concentrations cᵢ also changes total mass to ΣᵢVᵢcᵢ, where Vᵢ is volume in m³ and cᵢ is kg/m³. Simply summing concentrations over unequal-volume cells imposes the wrong constraint. Discretisation is part of the formula’s physical meaning.

A wrong constraint can worsen a good prediction

Take an almost-correct prediction, [5.9,3,1.1] kg, against synthetic truth [6,3,1]. Its sum is ten and total squared error is 0.02 kg². If we wrongly impose M = 12, projection adds 2/3 kg to every entry and error rises to about 1.3533 kg². This does not contradict the proof: truth lies outside the imposed set. A biased measurement, forgotten flux or inconsistent units can make a formal guarantee mathematically precise but physically wrong.

When the balance is uncertain, a soft penalty or an uncertainty-aware formulation may therefore be more appropriate than an arbitrary hard equality. In our simplified quadratic correction problem, adding α/2(Σᵢxᵢ − M)² leaves residual (Σᵢyᵢ − M)/(1 + nα). With initial excess 2 kg and n = 3, α = 1 leaves 0.5 kg; α = 100 leaves about 0.00664 kg. This derives a correction problem, not a PINN training benchmark. It shows why penalising and enforcing are different.

To verify the residual, set r = Σᵢxᵢ − M. Differentiation gives xᵢ − yᵢ + αr = 0. Summing over n components gives r = Σᵢyᵢ − M − nαr, yielding the previous fraction. α is dimensionless because both objective terms have units kg². This calculation excludes nonnegativity; if required, that constraint must also be included in the penalised problem.

The connection to research, with its limits

Baez and colleagues’ PINN-Proj preprint, arXiv v1 dated 12 November 2025, studies integral projections in physics-informed networks. We read methods and experiments: five differential-equation problems, penalty/projection comparisons and ten-trial averages. Conservation and state accuracy are separate metrics; tables show no universal victory on the latter. Training cost increases. We did not reproduce the study or extend our linear proof to all its quadratic constraints.

PINN means physics-informed neural network: training also incorporates residuals of physical equations. Not every surrogate network is a PINN. Output correction may be post-processing or part of a differentiable training chain; these are different protocols. For our affine constraint and fixed M, the projection derivative is I − 11ᵀ/n, removing the uniform component of variations. I is the identity matrix and 1 the all-ones vector. Nonnegativity introduces active-set changes and nondifferentiable points, requiring further training choices.

Reproduction and verification: what the code computes

The package uses exact fractions without an external numerical solver. affine subtracts mean excess in O(n). simplex sorts values descending, determines how many remain positive and finds θ from prefix sums; sorting dominates O(n log n) cost. Optimality conditions are checked component by component. For M = 0 it returns zeros; negative total is incompatible with nonnegative masses and rejected. Vectors and sorting require O(n) storage.

A real application should check units and balances, distinguish global from local conservation, document measurement and discretisation errors, and compare accuracy before and after correction on independent cases. Exactly satisfying a discretised integral does not make the continuous field exact. Correction at each instant also does not guarantee temporal dynamics: it might move mass without plausible flux. These are limits of the solved problem, not details to hide behind a perfect total.

Answering the opening question

Yes, a mass-creating prediction can be corrected without retraining if the balance is known and acceptable corrections are defined. Projection makes that choice explicit and, under precise assumptions, guarantees something about aggregate distance to feasible truth. But making ten kilograms add up does not reveal where they belong. The useful result is a prediction consistent with more known information, with separate checks on distribution and dynamics. Scientific AI requires asking not only how well it predicts, but which properties can be proved, from which data and within what boundaries.

Sources and executed code

Baez, Zhang, Ma, Nguyen, Das, Daniel — Guaranteeing Conservation of Integrals with Projection in Physics-Informed Neural Networks, arXiv v1, 12 November 2025 (preprint), sections 3–5.

Stephen Boyd and Lieven Vandenberghe — Equality constrained minimization, Convex Optimization lecture notes.

from experiment import affine, simplex, run
print([str(v) for v in affine([8, 3, 1], 10)])
print([str(v) for v in simplex([-1, 4, 9], 10)])
print(run()['wrong_total'])

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.