ELAI S.r.l.

Two models have different perplexity: are we measuring the same thing?

A worked example shows how tokenization can reverse a ranking: probability, bits per byte, context and limits of model comparisons.

Two models have different perplexity: are we measuring the same thing?

A better number can hide a different unit

We are choosing a language model to read maintenance reports. Two candidates are evaluated on the same text: one has perplexity 2, the other about 3.17. Knowing that lower perplexity is usually preferable, we might choose the first. But what does the denominator count? If one model splits a word into three fragments and another treats it as one element, they make different numbers of predictions. Direct comparison can reward the way text is divided rather than the probability assigned to the text itself.

Abstract. We derive perplexity from token probabilities and show a ranking reversal with three synthetic probabilistic models. We then normalize total surprise by the bytes of the same text, distinguishing that adjustment from a universal capability evaluation. We examine document aggregation, Unicode, context and the difference between tokenization probability and string probability. We neither train nor run a neural network: conditional probabilities are explicit teaching data, and Python recomputes every numerical result.

Tokens and probabilities: what is predicted

A tokenizer turns text into a sequence of elements called tokens. They may represent words, fragments, spaces or other vocabulary units. An autoregressive model assigns a probability to the next token using preceding tokens. To evaluate known text, we inspect the probabilities assigned to the observed tokens, not those it would select as its most likely answer. This article considers only this causal evaluation.

The probability of a token path is the product of conditional probabilities. “Conditional” means each value may depend on its prefix; token independence is not assumed. To avoid very small products, sum negative logarithms. Base-2 logarithms measure the quantity in bits, often interpreted as surprise: an event of probability 1/2 contributes one bit, and probability 1/4 contributes two.

P_path = ∏ᵢ p(tᵢ | t₁,…,tᵢ₋₁) L_bits = −Σᵢ log₂ p(tᵢ | t₁,…,tᵢ₋₁) PPL = 2^(L_bits / N)

N is the number of scored tokens, and L_bits is total surprise. Perplexity, abbreviated PPL, exponentiates average surprise per token. Natural logarithms and the natural exponential give the same result when used consistently. PPL is not a percentage of correct answers. Zero probability for an observed token produces infinite surprise; numerical approximations must be stated rather than silently repaired.

The running phrase: “motore caldo”

We keep the same Italian phrase in every language version because translating it would change the experiment. “motore caldo” contains twelve UTF-8 bytes: six letters, one space and five letters. Synthetic model A divides it into [mo, to, re, space, cal, do], six tokens. It assigns probability 1/2 to each observed token given its prefix. Path probability is (1/2)⁶=1/64, total surprise is six bits and PPL=2^(6/6)=2.

Model B instead uses three tokens, [motore, space, caldo], with probability 1/4 each. We obtain (1/4)³=1/64 and still six total bits. But the average is now two bits per token: PPL=2^(6/3)=4. The value doubled without changing the probability of the path representing the phrase. We have not shown that A understands the domain better; we changed the number of units being averaged.

Model C uses B’s three tokens but probabilities [1/2, 1/4, 1/4]. Their product is 1/32, twice A’s probability for the sequence. Surprise falls to five bits. Yet PPL=2^(5/3)≈3.1748, still above 2. A naive ranking selects A even though C assigns higher probability to the phrase’s token path. This is the opening example. Other vocabulary probabilities are not simulated; they may receive the remaining mass needed to complete each distribution.

ModelTokensPath probabilityTotal bitsPPLBits/byte
A61/64620.5
B31/64640.5
C31/3253.17480.4167

Using a common denominator: text bytes

To remove this particular unit mismatch, divide total surprise by B, the UTF-8 byte count of the text actually scored. Call the result BPB, bits per byte. Here B is a count, not the model name in the table. The denominator must cover exactly the content contributing to the numerator:

BPB = L_bits / B = N log₂(PPL) / B

A and B both yield 6/12=0.5 bits per byte. C yields 5/12≈0.4167. Comparison now agrees with path probabilities: A and B tie for this phrase, while C assigns more probability. The second form explains why PPL alone is insufficient: N and B are also needed. A different normalization creates no new evidence; it makes the unit of an existing measurement explicit.

Synthetic probabilities; no network was run. On the left A appears better because PPL is lower; on the right C has fewer bits per byte on the same text. Panel scales differ and must not be compared against each other.
Synthetic probabilities; no network was run. On the left A appears better because PPL is lower; on the right C has fewer bits per byte on the same text. Panel scales differ and must not be compared against each other.

A byte is not a letter, and one language is not another

ASCII is convenient because each character occupies one byte, but UTF-8 does not behave that way for all text. The program checks that precomposed “é” takes two bytes, visually similar “e” plus combining accent takes three, and Chinese “热” takes three. Unicode NFC normalization brings the first two representations to the same form, but applying it is a preprocessing choice to fix before evaluation, identically for every candidate.

Comparing models on the same byte-for-byte corpus removes denominator ambiguity. It does not make an Italian corpus equivalent to its Chinese translation: bytes, linguistic constructions and content distributions change. Lower BPB in one language does not alone prove greater competence in it. Our four editorial versions translate the reasoning while deliberately keeping the experimental string identical.

How to aggregate documents without changing the question

Suppose one document has 12 bytes and 6 bits of surprise, and another 120 bytes and 12 bits. Their BPBs are 0.5 and 0.1. The simple mean is 0.3, but the corpus contains 132 bytes and 18 bits: corpus BPB is 18/132≈0.13636. The simple mean weights documents equally; the ratio of sums weights bytes equally. Both are definable statistics, but answer different questions. The former must not be presented as the latter.

The same caution applies to PPL: corpus perplexity is not the arithmetic mean of document perplexities. Sum scored-token losses, divide by their count, then exponentiate. If the business objective weights each case equally, case-level aggregation may be intentional, but it must be stated alongside the corpus measure. The denominator is an evaluation choice, not a reporting detail.

Context, masks and boundaries are part of the measurement

A model scores a token given context, not in isolation. Resetting prior text at every block changes the prediction problem. Hugging Face’s perplexity guide discusses disjoint and sliding windows and code avoiding duplicate scoring. We read the definition, method and listing; we did not run its GPT-2 benchmark or adopt its numbers as our results.

Our example scores every path probability from a conventional initial context. It excludes an end-of-sequence token: we score the textual continuation, not the decision to stop there. Real comparisons must align starts, endings, prompts, templates, document separators and excluded tokens. If only answers are scored, loss cannot be divided by all prompt-plus-answer bytes; that would artificially lower the result.

Giving two tokenizers the same 1,000-token window also does not guarantee equal textual context: they may cover different portions of a report. State what information each model receives. The synthetic case avoids this issue: the phrase is short and every probability is defined on its full prefix. For a real corpus, this remains an experimental choice to document, not something bits per byte automatically solve.

The subtle caveat: a path and a string need not have the same probability

We deliberately said “path probability”. If different token sequences decode to the same text, the text’s total probability must sum contributions from allowed sequences under consistent boundary and termination conventions. A deterministic tokenizer may select a canonical path, while the generative model assigns probability to other paths producing the same characters. Scoring only the canonical path does not perform that sum.

An abstract example exposes the distinction: two complete, mutually exclusive paths both decoding to “ab” have probabilities 0.02 and 0.08. String probability is 0.10, while an evaluator selecting the first records 0.02. We did not implement marginalization over every path of a real tokenizer. Our BPB values therefore represent normalized loss of the selected path, not a promise of exact textual likelihood independent of all tokenization choices.

What to evaluate in a domain-specific model

On maintenance reports held out of training, language loss can measure how predictable the model finds domain text. It does not automatically measure correct fault identification, unit preservation, procedural compliance or invented causes. A model may give high probability to frequent sentences that are useless for a decision. Selection therefore needs task tests alongside language metrics, with relevant errors and a separate protocol. We report no unperformed application results.

Answer: clarify the unit before ranking

Different perplexities need not indicate different language quality: they may average surprise over different units. Here A scores 2 and C about 3.17, yet C assigns greater probability to the same phrase’s sequence. Byte normalization exposes this without automatically resolving context, preprocessing, alternative paths or task utility. Ask which loss was measured, on what content and under what conditions. Only then does comparing numbers make sense.

The listing reproduces the three scores. The archive includes explicit tokens, probabilities, phrase-reconstruction checks, aggregation, Unicode examples and plot data. Everything is deterministic, with no seed. Scoring traverses N probabilities once, costing O(N); it does not include real-model inference cost. The figure uses saved JSON numbers. The following documentation supports the definition and methodological caution; numerical examples are independent educational analysis.

Source and code

Hugging Face Transformers — Perplexity of fixed-length models.

from math import log2
text = "motore caldo"
for name, probs in [("A", [.5]*6), ("B", [.25]*3), ("C", [.5,.25,.25])]:
    bits = -sum(log2(p) for p in probs)
    ppl = 2**(bits/len(probs))
    bpb = bits/len(text.encode("utf8"))
    print(name, bits, ppl, bpb)

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 3 October 2026.