When the slope matters as much as the value
A model reconstructs process temperature within one hundredth of a degree. Can we also trust its heating-rate estimate? Physical AI models may fit values closely while oscillating between them. A derivative measures how rapidly a quantity changes and exposes oscillations almost invisible in the values. Small temperature error does not automatically imply small rate-of-change error.
Abstract. First derive the noise-versus-time-window tradeoff when estimating derivatives from noisy measurements. Then compare functions that are close in value but far apart in derivatives. A local polynomial filter provides an alternative to differencing two samples. Results are calculations on synthetic functions and a stated statistical model, not laboratory measurements, network training or plant validation. The goal is to identify the extra evidence needed before using a model inside a physical equation.
A simple process with explicit units
Use f(t) = 20 + 2 sin(ωt), temperature in Celsius, t in seconds and ω = 2π/60 radians per second. It oscillates with amplitude 2 °C and period 60 s; it is not a complete heat-transfer model. Its derivative is f′(t) = 2ω cos(ωt), in °C/s, about 0.20944 at t = 0. Zero is a local time origin, so negative sample times simply precede that point.
Measurements add error: y(t) = f(t) + ε(t). Assume independent zero-mean errors with standard deviation σ = 0.05 °C at every sample. The variance calculations do not require Gaussian noise. σ is an educational assumption, not a measured instrument specification. Correlation or systematic drift would change the formula.
Why closer samples can worsen the estimate
Estimate the central rate using measurements h seconds before and after. Divide temperature difference by elapsed time 2h to obtain central difference D_h. Without noise, smaller h brings average slope closer to the derivative. With noise, it also divides the error difference by an increasingly small number, potentially destabilizing an apparently more time-local estimate.
Variance measures random spread in (°C/s)². Independent error variances add; subtraction does not cancel them. Standard deviation is σ/(√2 h): about 0.03536 °C/s at h = 1 s, but 0.35355 at h = 0.1 s, larger than the true 0.20944 derivative. These are numerically evaluated theoretical moments, not observations from repeated trials. No random samples are generated.
Wider windows reduce noise but introduce bias
Why not use an enormous interval if distant samples reduce spread? Because average slope across a curved segment differs from central slope. For our sinusoid at zero, mean D_h is exactly 2 sin(ωh)/h; subtracting 2ω gives bias, the systematic mean error. For a sufficiently smooth function, Taylor expansion gives leading bias h² f‴(t)/6, growing quadratically while the local expansion is valid.
Subtracting the two Taylor expansions cancels even terms; dividing by 2h leaves the derivative and a cubic term reduced to order h². O(h⁴) denotes residual order four or higher with sufficient smoothness. Mean squared error, MSE, equals squared bias plus variance because the random deviation from its own mean averages to zero. Its square root, RMSE, returns to °C/s for comparison with the true rate.
| h (s) | Mean D_h (°C/s) | RMSE (°C/s) |
|---|---|---|
| 0.1 | 0.209436 | 0.353553 |
| 1 | 0.209057 | 0.035357 |
| 3 | 0.206011 | 0.012274 |
| 5 | 0.200000 | 0.011794 |
| 10 | 0.173205 | 0.036407 |
Increasing h from 0.1 to 5 seconds improves RMSE here; increasing to 10 worsens it. This is not a universal step choice: noise and signal curvature determine the tradeoff. Central differencing also uses future sample t+h, requiring h seconds of delay in real time or replacement with a causal estimator using only past and present. Numerical accuracy does not remove that timing constraint.
A good value fit does not control derivatives
Now consider the approximating function itself, rather than measurement noise. Define g(t) = f(t) + 0.01 sin(8πt). Its value differs by at most 0.01 °C everywhere, with perturbation period 0.25 s. Yet g′(t) − f′(t) = 0.08π cos(8πt), whose amplitude is about 0.25133 °C/s, exceeding the original derivative amplitude 0.20944. Close values did not prevent markedly different slopes.
This is an analytical counterexample, not a badly trained neural network. Value-only validation error cannot by itself guarantee derivatives used in physical models. More generally ε sin(kt) has amplitude ε but derivative amplitude εk: increasing frequency can enlarge the latter without enlarging the former. k has inverse-time units, giving εk temperature-per-second units.
Both left-panel axes are logarithmic: equal distances represent equal ratios, not differences. Moving from 0.1 to 1 second multiplies the step by ten. Watch which error component dominates and where the total stops decreasing. Right-panel axes are linear: g′ even becomes negative while f′ stays positive, suggesting local cooling where the reference still warms.

An alternative: fit a local slope from several points
At five equally spaced times t+jΔ, j from −2 to 2, fit a+bjΔ+c(jΔ)² by least squares. We want central derivative b. Symmetry makes sums of j and j³ zero, so the linear column is orthogonal to constant and quadratic columns. Its normal equation gives b = Σ_j j y_j /(Δ Σ_j j²). Since Σ_j j² = 10, weights are [−2, −1, 0, 1, 2]/(10Δ).
Independent equal-variance noise gives variance σ² times squared-weight sum: 0.1σ²/Δ², versus 0.5σ²/Δ² for central difference at h=Δ. Random spread falls, but the window is wider and assumes a low-degree polynomial describes it. At Δ=1 s, mean estimate is 0.20814 °C/s and RMSE 0.01586, versus central difference RMSE 0.03536. This is not universal superiority: fast changes may be smoothed away.
This mechanism underlies local polynomial differentiation, including Savitzky–Golay filters. SciPy distinguishes derivative order, window length, polynomial degree and sample spacing delta. We implement the five weights directly rather than running SciPy. Series boundaries lack symmetry, and padding or fitting choices change results. Saying data were filtered is not enough to reproduce a derivative.
What to verify before the curve enters a physical model
If an equation uses a rate, evaluation must cover that quantity with its units and time scale. Curvature penalties, derivative data or physical constraints can restrict oscillations, but introduce assumptions and weights to validate rather than automatic guarantees. Even exact differentiation of a network may give the wrong physical slope: calculation can be correct for an inadequate learned function.
The answer is no: accurate temperature reconstruction alone does not establish derivative accuracy. The first example amplifies noise through too narrow a window; the second turns a small model oscillation into a large slope oscillation. Different causes require different checks. Real-data work needs characterized noise, acquisition timing, boundary treatment and an independent derivative reference. These were not measured here.
Full code saves the table, error curves and analytical trajectories, using Python’s standard library for calculations and Matplotlib for figures. No hardware or EL-AI results are included. The snippet reproduces the essential table: mean difference, bias, variance and their combination. It separates the known synthetic derivative from the estimator under noisy measurements.
import math
omega = 2*math.pi/60
sigma = 0.05
true = 2*omega
for h in [0.1, 1, 5, 10]:
mean = 2*math.sin(omega*h)/h
bias = mean-true
variance = sigma**2/(2*h*h)
print(h, mean, math.sqrt(bias*bias+variance))
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.

