Knowing the words does not establish the new context
A specialized model routes technical support requests, uses sector vocabulary correctly and passes historical tests. After a service change, its priorities become less reliable. Did the request mix change, or did the meaning of urgency change? These explanations call for different interventions. Before adding documents or retraining, identify which relationship moved.
Abstract. Two synthetic ticket categories explain importance weighting: changing old-example weights to represent a new population. It works when input frequencies change but the input-to-answer relationship stays stable. A second world has identical inputs and a different rule: weights stay the same, while evaluation and even classifier ranking become wrong. Results are exact enumerations, not measured language-model performance.
Two distributions, two questions
Let x be ticket type A or B, and y its label: one urgent, zero not urgent. p(x) describes type frequency. Conditional probability p(y = 1 | x) describes urgency within that type; the bar means given that type. More B tickets does not itself mean B tickets became more urgent. The joint distribution combines both: p(x,y) = p(x)p(y|x).
In source world s, 80% of tickets are A and 20% B; 90% of A and 60% of B are urgent. In target world t, A drops to 20% and B rises to 80%, initially retaining those 90% and 60% conditional rates. This is covariate shift: p(x) changes, p(y|x) does not. s and t identify source and target, not two measurements of one sample.
A historical 16% error becomes 34%
Classifier U always predicts urgent. Its simplicity isolates evaluation rather than proposing a useful system. Zero-one loss counts a mistake when prediction differs from the label. U errs on 10% of A and 40% of B. Source risk, the population-average error, is 0.8 × 0.1 + 0.2 × 0.4 = 0.16. Target risk is 0.2 × 0.1 + 0.8 × 0.4 = 0.34. The classifier is unchanged; its harder cases became more common.
Can old labels recover new risk? Multiply each type’s contribution by its target-to-source frequency ratio: w(A) = 0.25 and w(B) = 4. A counts one quarter, B fourfold. This does not invent labels; it changes how strongly old cases represent today’s target population.
Here f is the classifier and L is one for an error, zero otherwise; sums run over types and labels. The second line is the decisive assumption. Replacing the target conditional distribution with the source one rewrites target risk as weighted source risk. If that substitution is false, the last line is not target risk. Source probability must also be positive wherever target probability is positive: weighting cannot supply an unseen category.
| Type | Label y | Count per 1000 | Weight |
|---|---|---|---|
| A | 0 | 80 | 0.25 |
| A | 1 | 720 | 0.25 |
| B | 0 | 80 | 4 |
| B | 1 | 120 | 4 |
The table is an exact synthetic fixture, not a random sample. U misses 80 nonurgent A and 80 nonurgent B: 160 of 1000. Weighted errors are 80 × 0.25 + 80 × 4 = 340; total weight remains 1000. Weighted risk is 0.34, exactly target risk under frequency change alone. Code uses rational fractions and saves counts; no seed or training is needed.
Same inputs, different rule: where weighting fails
Keep target inputs at 20% A and 80% B, but change the service rule: now only 10% of A is urgent, while B remains 60%. This synthetic conditional change isolates the problem without claiming a real operational cause. U’s A error rises from 10% to 90%. New risk is 0.2 × 0.9 + 0.8 × 0.4 = 0.50. Old-data weighting still gives 0.34 because it contains no new label information.
Compare U with V, which predicts nonurgent for A and urgent for B. In the first target world V errs on 90% of A and 40% of B: risk 0.50, worse than U. After the rule change it errs on only 10% of A: risk 0.34, better than U. Weighted historical data retain the old ranking. This can reverse model selection, not merely miscalibrate a score. Both classifiers are fixed rules, not trained on test data.

Monitoring p(x) alone cannot distinguish the two target worlds: both contain 20% A and 80% B. This is an information limit, not something a more sophisticated detector on the same inputs fixes. New labels, verifiable outcomes or reliable rule-change information are needed. Delayed labels delay evaluation; labels collected only for disputed cases introduce a selection process rather than automatically representing all tickets.
Even in the favorable case, a few examples can carry too much weight
Weights are exact because we constructed the distributions. In practice they must be estimated, and source-rare categories may get large weights, amplifying label mistakes. A descriptive concentration measure is n_eff = (Σ_i w_i)² / Σ_i w_i². With 800 weights of 0.25 and 200 of 4 it is 1000² / 3250 ≈ 307.69. No tickets were deleted, and this alone is not a confidence interval: it indicates uneven influence.
Capping weights may reduce instability but breaks the exact identity derived above; it is a tradeoff to evaluate, not a free improvement. Better vocabulary representations may help distinguish previously confused categories, but do not establish stable business rules. Document updates, label revisions and population corrections affect different components. Confusing them makes success or failure hard to explain.
What a vertical-model project should take away
Before claiming domain adaptation, state what changed and what was verified. Weighting exactly corrects 16% to 34% only under conditional stability; when the rule changes, true error is 50% and new information is needed. Sector vocabulary does not replace this diagnosis. Useful evaluation reports population, period, label definition and transfer assumptions alongside the score. We measured no EL-AI model; we made a condition of specialization projects auditable.
Sources and experimental scope
The 2007 paper studies weighted cross-validation under stable conditionals and, theoretically, known density ratios. Section 4.1 fits a linear model to a sinc target using 150 samples, ten-fold validation and 1000 repetitions. The authors report improved risk estimation over unweighted validation in that protocol. We did not reproduce it: our discrete example verifies the identity for two fixed classifiers and violates its assumption in a second case. This is neither new research nor a general language-network result.
The snippet computes U’s three errors; the archive adds V, the joint table, total weight, weight concentration and assertions. Computation scales with type-label cells without numerical optimization. A future real-ticket test should define operational labels, separate model selection from evaluation and check new-case coverage. That test was not performed here.
from fractions import Fraction as F
source = [F(4,5), F(1,5)]
target = [F(1,5), F(4,5)]
old_error = [F(1,10), F(2,5)]
new_error = [F(9,10), F(2,5)]
weights = [t/s for t,s in zip(target,source)]
print('source', float(sum(s*e for s,e in zip(source,old_error))))
print('weighted', float(sum(s*w*e for s,w,e in zip(source,weights,old_error))))
print('changed rule', float(sum(t*e for t,e in zip(target,new_error))))
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.

