The number that decides whether a document proceeds automatically
A business uses a classifier to distinguish two document types. Each input receives a category and a number: “90% confidence”. Sending only cases below 80% to human review might seem reasonable. But an overconfident model passes incorrect documents with the same apparent certainty as correct ones. This is not a cosmetic problem: the number shapes work, errors and expectations.
The question is precise: can we make confidence more trustworthy without retraining the classifier? We construct an example that reports 90% but gets only 60 of 100 documents right. We calculate a temperature that corrects this gap, then show why it adds no correct classifications, why it can worsen in another context and how a perfect average can conceal two poorly calibrated groups. The data are constructed to explain the mechanism, not collected from customers.
Three quantities that should not be confused
The predicted category is a discrete decision: document A or B. Accuracy is the fraction of correct decisions in a labelled set. Confidence is a value assigned to the chosen category before the outcome is known. A classification with 90% confidence has not been proved correct: we are reading a prediction about correctness that needs verification on comparable cases.
A desirable meaning of 90% is that, among many cases assigned this value, about nine in ten are correct. This is confidence calibration. It does not promise exactly nine successes in every consecutive block of ten or identify the incorrect case. The frequency concerns a population and context; a finite sample estimates it. Here we measure correctness of the winning class, not truthfulness of a generative answer verbally claiming certainty.
From internal scores to probabilities
Following the calculation requires probabilities, averages and natural logarithms. Readers unfamiliar with differentiation can treat the derivation as justification of the final rule. Suppose the model produces two real scores, called logits: z₀ and z₁. They are not probabilities and may be negative. Softmax transforms them into two positive numbers summing to one. A larger score difference produces a stronger preference.
T is a dimensionless parameter called temperature; it measures no physical heat. T = 1 leaves the model unchanged; increasing T moves both probabilities towards 0.5. Confidence q is the larger probability. Our winning score always exceeds the losing score by m = ln 9, approximately 2.197225. The winning class changes across documents, but the margin stays constant to isolate the reported-confidence problem.
Divide numerator and denominator of softmax by the exponential of the winning score to obtain this expression. At T = 1, exp(−ln 9) = 1/9, hence q = 9/10. Only the difference matters, not absolute scores: adding the same constant to both changes nothing. Division by a positive number also preserves their order. The winner remains the winner even when its probability decreases.
Choosing temperature with a verifiable objective
Construct one hundred calibration records with sixty correct and forty incorrect classifications. Their synthetic labels are determined by the script: we did not infer on one hundred real documents. How much should confidence decrease? Use mean negative log-likelihood, abbreviated NLL. It penalises the probability assigned to the true label: a confident mistake costs more than a hesitant one. In the binary case a correct record assigns q to truth, an incorrect record 1 − q.
a is the correct fraction and q the common confidence. Natural logarithms express loss in nats, a dimensionless information unit. At q = 0.9 we obtain L = 0.984250 nats per record. This is not a 98% error rate: accuracy remains 60%. The question has changed from counting correct decisions to assessing probabilities. To minimise loss, differentiate with respect to q and find the zero slope.
Between zero and one, the first derivative has a positive denominator. When q exceeds a, reducing it lowers loss; when q is smaller, increasing it helps. Positive second derivative makes q = a the unique minimum. Thus, for this homogeneous set, optimal confidence equals observed frequency. We have not proved that the sample frequency is the exact probability in the outside world.
Solve q(T) = a to obtain the formula. Temperature above one softens confidence and reduces loss from about 0.984 to 0.673. The sixty correct classifications remain sixty. We correct how we read the model, not its output categories. This closed form works because all margins are identical. With different margins, optimise the sum of losses instead of simply substituting average accuracy into this equation.
There is a structural limit: positive margin and temperature keep winning confidence above 0.5. If fewer than half the answers were correct in this same example, no finite temperature could bring q to that frequency. For a = 0.5, the limit T → infinity is required. A systematically reversed classifier needs its model or class mapping reconsidered; one confidence parameter cannot solve every error.
Evaluation must use data that do not choose T
Fitting T and celebrating results on the same hundred records describes only calibration-set fit. We therefore construct a second set with distinct identifiers and seventy correct answers. Keep T = 5.419023 without refitting. Its loss changes from 0.764528 to 0.632465 nats and confidence becomes 60% against 70% accuracy. It is closer to the correct frequency but does not match it.
This separation is logical, not statistical evidence of generalisation: we chose the second set’s counts. The script does not simulate random sampling or justify empirical confidence intervals. A real project needs representative data, separation from training and calibration, reliable labels and a protocol fixed before inspecting the test. Repeatedly selecting T after seeing evaluation results turns that test into another development set.
| Set | Correct | NLL T=1 | NLL T* |
|---|---|---|---|
| Calibration | 60% | 0.984250 | 0.673012 |
| Holdout | 70% | 0.764528 | 0.632465 |
| Shift | 95% | 0.215222 | 0.531099 |
A context change can reverse the benefit
The third row is a constructed counterexample: identical margins, but ninety-five correct classifications in one hundred. It might represent an easier stream, without simulating any specific cause. Old confidence of 90% is five percentage points from the observed frequency; calibrated confidence of 60% is thirty-five points away. NLL worsens from 0.215222 to 0.531099 nats. Lowering confidence is not always prudent; it can become systematic underestimation.

Read the graph by following the 60% curve: loss decreases to the dashed line, then rises towards nearly uniform prediction. On the 95% curve, that point is already in the underconfident region. The figure compares no trained networks; it compares three constructed label distributions around identical scores. It explains why yesterday’s calibration parameter needs checking when documents, suppliers, languages or classification rules change.
A perfect average can hide two problems
So far each set has been summarised by one number. A common indicator, Expected Calibration Error or ECE, groups confidences into bins and compares accuracy with mean confidence within each bin. It then sums absolute gaps weighted by record counts. The finite-data quantity is an estimate that also depends on binning. It is not universal certification of reliability.
N is the total and n_b the number of records in bin b; empty bins contribute nothing. All calibrated confidences equal 0.6 and occupy one bin. On the first set ECE is therefore |0.6 − 0.6| = 0, apart from numerical rounding. Now divide those hundred cases into equal groups: A has 45 correct out of 50, B only 15 out of 50. Total correct remains sixty and every case still reports 60%.
Group A is underestimated by thirty points and B overestimated by thirty. Overall ECE is zero because confidence binning mixes the errors before taking the absolute value. The weighted average of separately computed group ECEs would be 0.30. No common T can assign 90% to one group and 30% to the other when given the same margin: the chosen transformation lacks both flexibility and useful information.
Groups might represent scan types or document families, but here they are only synthetic labels. This counterexample proves no defect in a particular product or language. It motivates a validation question: is probability informative where it will be used? Splitting everything into tiny groups does not automatically help because frequencies become unstable. Groups need process-based justification, adequate sample sizes and evaluation outside the data used to choose corrections.
The operational threshold changes even when the classifier does not improve
Return to the initial rule: automate if q ≥ 0.8. Before calibration all hundred documents pass; afterwards none do. Coverage, the automatically handled fraction, falls from 100% to 0%. We cannot announce 100% accuracy on accepted cases: the set is empty and its correct fraction undefined. We exposed a confidence problem without constructing a sustainable production workflow.
A real decision must connect probability with error costs, review capacity and deferral options. This experiment estimates none of those costs and recommends no deployment threshold. It shows why comparing confidence percentages alone is insufficient. Distinguish classifier accuracy, probability quality, coverage and whole-process performance, including what happens to documents sent for review.
Alternatives, published evidence and reproducibility
Temperature uses one parameter and preserves class order within each input. An affine logit transformation can add a shift; class-specific maps or more flexible transformations can change decisions. They gain expressive power but must be estimated from data: more freedom does not imply better out-of-sample reliability. Grouping probabilities into bins likewise trades detail against the observations available for each estimate.
Guo, Pleiss, Sun and Weinberger study calibration on images and documents with separate datasets in their ICML 2017 paper. Their comparison shows frequent, not universal, benefits from temperature scaling. We cite methodological evidence without replicating their networks or benchmarks.
In our attachment, make_data constructs labels and logits; confidence applies the transformation; nll scores probabilities; run freezes T before evaluating the other sets. Assertions check distinct identifiers, unchanged classes and directions of loss changes. No seed is needed because nothing is random. Floating-point arithmetic makes theoretically zero ECE appear near 10⁻¹⁶ in JSON; this is numerical rounding, not a scientific discovery.
What does that 90% actually mean?
It means nine correct cases in ten only as a probabilistic interpretation requiring verification on comparable cases in the deployment context. Temperature can correct overconfidence without changing the classifier; our calculation shows exactly how and at what coverage cost. A good average does not guarantee every subgroup, and a correction useful in one stream may worsen another. For a vertical model, ask not just “how confident is it?” but “where have we checked that these numbers mean what they appear to mean?”.
Bibliography
from experiment import run
r = run()
print(round(r['temperature'], 6))
for m in r['metrics']:
print(m['dataset'], round(m['confidence'], 2),
round(m['accuracy'], 2), round(m['nll_nats'], 6))
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.

