A correct answer can contain too little information
A camera must distinguish three component types: similar brackets A and B, and a very different spacer C. A large model correctly recognises A, while considering B a plausible alternative and C almost impossible. We want a smaller model to learn the same task. Giving it only label A discards part of the lesson: to the teacher, confusing A with B is not the same as confusing it with C.
Distillation transfers information from a teacher model’s outputs to a student. Our question is precise: what information travels through alternative probabilities, and why does temperature change learning? A three-class example will lead to the objective, update signal and high-temperature limit. The result helps design a comparison; it does not establish that a student is already accurate, fast or production-ready.
The teacher does not transfer its weights
A model assigns an input a set of scores, called logits, before converting them to probabilities. The student may have a different architecture: it need not copy every teacher parameter. Instead it tries to reproduce the teacher’s behaviour on transfer inputs. The teacher–student analogy concerns this functional relationship, not human intention or understanding. If the teacher is systematically wrong, that behaviour can also become part of the lesson.
In our synthetic example the teacher logits are v = (4, 3, 0), and the student’s initial logits z = (2, 0, 0), ordered A, B, C. A is correct. These numbers isolate the mechanism; they do not come from real images, a trained model or an EL-AI project. The student distinguishes A from the alternatives but gives B and C identical scores. The teacher clearly distinguishes those alternatives.
From scores to probabilities: what temperature does
To compare models we need distributions: positive numbers summing to one. Softmax exponentiates scores and normalises them. Temperature T, positive and dimensionless, divides logits before exponentiation. It is not physical temperature or processor heat. At larger T, differences matter less and the distribution becomes less concentrated. Class ranking does not change when the same T applies to every score of a model.
p denotes the teacher, q the student; i selects a class and j runs over all classes in the denominator. At T=1 the teacher yields about (0.721399, 0.265388, 0.013213). At T=2 it yields (0.574097, 0.348207, 0.077696). C remains least plausible, but its probability is less suppressed. We have not added teacher information; we have changed how its differences are exposed to the learning objective.
| T | p(A) | p(B) | p(C) |
|---|---|---|---|
| 1 | 0.721399 | 0.265388 | 0.013213 |
| 2 | 0.574097 | 0.348207 | 0.077696 |
| 4 | 0.465836 | 0.362793 | 0.171371 |
| 20 | 0.361016 | 0.343409 | 0.295575 |
Read across rows to compare alternatives and down columns to see T’s effect. The ratio p(B)/p(C) is exp((3−0)/T): about 20.09 at T=1 and 4.48 at T=2. Temperature does not make alternatives more separable by ratio; it softens them. Any benefit comes from how the objective uses the whole distribution, not a magical amplification of differences.
Two lessons in one objective
The student can receive both the correct label and teacher probabilities. The first term penalises low probability for A. The second measures a directional distribution discrepancy using Kullback–Leibler divergence, KL. For fixed p, minimising KL(p||q) is equivalent to minimising teacher cross-entropy: they differ by a term independent of the student. We use natural logarithms and sums over classes, without dividing by the number of classes.
α controls teacher weight relative to the label: α=0 ignores the teacher; α=1 ignores the label in the loss. It is not a certified trust probability. The supervised term uses T=1; the transfer term uses the same T for both models. The teacher stays fixed during this update. Accidentally updating its parameters would implement another algorithm. Why T² appears will become clear from the gradient, the actual direction of model correction.
The learning signal: which score rises?
The gradient with respect to a logit tells how loss changes when that score increases slightly while the others stay fixed. Gradient descent subtracts a small proportional amount: a positive gradient lowers the score, a negative one raises it. The label term gives q_i(1)−y_i, with y=(1,0,0). Here it is about (−0.213014, 0.106507, 0.106507). The label pushes A up and lowers B and C equally.
For the teacher term, start from log q_i(T) = z_i/T − log Σ_j exp(z_j/T). Differentiating with respect to z_k gives (δ_ik−q_k)/T; δ_ik is one when i=k and zero otherwise. In the p_i-weighted sum, teacher terms are constant and Σ_i p_i=1. What remains is the difference between what the student assigns and what the teacher asks, divided by T. This assumes both models describe identical classes in identical order.
At T=2 the compensated distillation gradient is (0.004040, −0.272532, 0.268492). The main correction is not A: it nearly matches the teacher at that temperature. It concerns B and C. B must rise and C fall; labels alone did not distinguish them. With α=0.5 the total gradient is (−0.104487, −0.083012, 0.187499): A and B increase, C decreases. These are logit directions for this input, not evidence of improved accuracy on other inputs.
Why multiply by T² without promising invariance
When T becomes large relative to logit differences, both distributions approach uniform. Their difference decreases approximately as 1/T; the softmax derivative introduces another 1/T. The uncompensated gradient is thus of order 1/T². Multiplying loss by T² prevents teacher influence from becoming negligible merely because of this scaling. It does not make learning identical at every temperature: at finite T, both distributions and signal directions change.

First inspect the left panel: as T rises, signals collapse towards zero. On the right they remain distinct even at T=100. Then inspect A: its gradient changes sign between T=2 and T=4. This directly counters the idea that T² cancels every temperature effect. Temperature determines which differences are compared; the factor controls part of the signal scale. These roles are related but different.
The technical limit: comparing centred logits
Adding the same constant to all logits leaves softmax unchanged because the factor cancels between numerator and denominator. Absolute score level is therefore not identifiable from probabilities. Subtract each vector’s own mean, v̄ or z̄, before comparison. With K classes, the high-temperature expansion gives p_i ≈ 1/K + (v_i−v̄)/(KT), and similarly for q. T merely exceeding one is insufficient: it must be large relative to score spread.
For our vectors the compensated gradient at T=100 is about (−0.109414, −0.441091, 0.550505), close to the displayed limit. This connects distillation to a quadratic comparison of centred logits. Centring matters: vectors differing only by a constant describe identical probabilities and should not be penalised as different behaviours. It does not follow that arbitrarily high temperature is better: the mathematical limit describes the objective, not model generalisation.
What the foundational work showed, and what this example does not
Hinton, Vinyals and Dean describe this approach in their 2015 preprint Distilling the Knowledge in a Neural Network. We read the method, derivation and MNIST experiments: a two-layer 1200-unit teacher, an 800-unit student, and specific data and regularisation. They report 67 teacher errors, 146 baseline-student errors and 74 with distillation at T=20. These are historical author results, not our reproduction or a prediction for industrial images. The missing-digit section includes test-set bias adjustments, distinct from independent validation selection.
Our program trains no networks. It computes probabilities, KL and gradients for the stated vectors, then compares the analytic derivative with finite differences: increase and decrease one logit by 0.0001, evaluate loss and divide the difference by 0.0002. The check allows maximum error 10⁻⁷. It also checks zero gradient sum and invariance when 100 is added to every teacher logit. There is no randomness and no seed. These verify the mechanism, not accuracy.
The comparison needed for a domain model
To establish value for brackets, compare the same student trained with labels alone and with labels plus distillation. Declare data, budget, initialisations and protocol; select T and α on validation, keeping test separate. If images share the same part or batch, random image splitting can make test overly similar to training. Measure A↔B errors, novel cases and miscalibrated probabilities as well as overall accuracy.
A smaller model does not automatically imply lower device latency. Architecture, runtime, numerical precision and end-to-end measurements are needed. Training cost changes too: the teacher must produce transfer outputs, possibly cached; with K classes the loss operations per example grow linearly with K, while network cost is separate. Quantisation and distillation address different issues and can be combined only with joint evaluation. We performed none of these hardware measurements.
The answer: transfer differences, with a teacher to verify
A large model can teach how plausibility is distributed across alternatives, not just which class to choose. In our example distillation raises B and lowers C where the label treated them equally. Temperature and teacher weight determine how that information influences the student; T² compensates a scaling effect without guaranteeing equivalence or improvement. Practical value depends on teacher quality, data coverage and student capacity. The numerical demonstration explains the mechanism; domain validation remains to be done.
References and reproducible material
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, Distilling the Knowledge in a Neural Network, arXiv:1503.02531v1, 9 March 2015. Foundational reference, sections 2, 2.1 and 3 read; not presented as the 2026 state of the art. This is an AI-assisted educational monograph, not peer-reviewed original research. It does not establish an EL-AI proprietary model or an available application.
Hinton, Vinyals & Dean (2015) — Distilling the Knowledge in a Neural Network.
import math
def softmax(values, T):
maximum = max(values)
exps = [math.exp((x-maximum)/T) for x in values]
return [x/sum(exps) for x in exps]
v, z, T = [4., 3., 0.], [2., 0., 0.], 2.
p, q = softmax(v, T), softmax(z, T)
gradient = [T*(s-t) for s,t in zip(q,p)]
print([round(x, 6) for x in p])
print([round(x, 6) for x in gradient])
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 28 September 2026.

