ELAI S.r.l.

The model knows the sector: why does it fail when requests change?

Changing request frequencies differs from changing the rule. An exact example shows when data weighting corrects evaluation and when it fails.

The model knows the sector: why does it fail when requests change?

Knowing the words does not establish the new context

A specialized model routes technical support requests, uses sector vocabulary correctly and passes historical tests. After a service change, its priorities become less reliable. Did the request mix change, or did the meaning of urgency change? These explanations call for different interventions. Before adding documents or retraining, identify which relationship moved.

Abstract. Two synthetic ticket categories explain importance weighting: changing old-example weights to represent a new population. It works when input frequencies change but the input-to-answer relationship stays stable. A second world has identical inputs and a different rule: weights stay the same, while evaluation and even classifier ranking become wrong. Results are exact enumerations, not measured language-model performance.

Two distributions, two questions

Let x be ticket type A or B, and y its label: one urgent, zero not urgent. p(x) describes type frequency. Conditional probability p(y = 1 | x) describes urgency within that type; the bar means given that type. More B tickets does not itself mean B tickets became more urgent. The joint distribution combines both: p(x,y) = p(x)p(y|x).

In source world s, 80% of tickets are A and 20% B; 90% of A and 60% of B are urgent. In target world t, A drops to 20% and B rises to 80%, initially retaining those 90% and 60% conditional rates. This is covariate shift: p(x) changes, p(y|x) does not. s and t identify source and target, not two measurements of one sample.

A historical 16% error becomes 34%

Classifier U always predicts urgent. Its simplicity isolates evaluation rather than proposing a useful system. Zero-one loss counts a mistake when prediction differs from the label. U errs on 10% of A and 40% of B. Source risk, the population-average error, is 0.8 × 0.1 + 0.2 × 0.4 = 0.16. Target risk is 0.2 × 0.1 + 0.8 × 0.4 = 0.34. The classifier is unchanged; its harder cases became more common.

Can old labels recover new risk? Multiply each type’s contribution by its target-to-source frequency ratio: w(A) = 0.25 and w(B) = 4. A counts one quarter, B fourfold. This does not invent labels; it changes how strongly old cases represent today’s target population.

R_t(f) = Σ_x Σ_y p_t(x)p_t(y|x) L(f(x),y) p_t(y|x) = p_s(y|x), w(x) = p_t(x)/p_s(x) R_t(f) = Σ_x Σ_y p_s(x)p_s(y|x) w(x)L(f(x),y)

Here f is the classifier and L is one for an error, zero otherwise; sums run over types and labels. The second line is the decisive assumption. Replacing the target conditional distribution with the source one rewrites target risk as weighted source risk. If that substitution is false, the last line is not target risk. Source probability must also be positive wherever target probability is positive: weighting cannot supply an unseen category.

TypeLabel yCount per 1000Weight
A0800.25
A17200.25
B0804
B11204

The table is an exact synthetic fixture, not a random sample. U misses 80 nonurgent A and 80 nonurgent B: 160 of 1000. Weighted errors are 80 × 0.25 + 80 × 4 = 340; total weight remains 1000. Weighted risk is 0.34, exactly target risk under frequency change alone. Code uses rational fractions and saves counts; no seed or training is needed.

Same inputs, different rule: where weighting fails

Keep target inputs at 20% A and 80% B, but change the service rule: now only 10% of A is urgent, while B remains 60%. This synthetic conditional change isolates the problem without claiming a real operational cause. U’s A error rises from 10% to 90%. New risk is 0.2 × 0.9 + 0.8 × 0.4 = 0.50. Old-data weighting still gives 0.34 because it contains no new label information.

Compare U with V, which predicts nonurgent for A and urgent for B. In the first target world V errs on 90% of A and 40% of B: risk 0.50, worse than U. After the rule change it errs on only 10% of A: risk 0.34, better than U. Weighted historical data retain the old ranking. This can reverse model selection, not merely miscalibrate a score. Both classifiers are fixed rules, not trained on test data.

Exact errors of U and V. Weighted historical error matches the target under frequency change alone; changing the rule reverses preference. Entirely synthetic data.
Exact errors of U and V. Weighted historical error matches the target under frequency change alone; changing the rule reverses preference. Entirely synthetic data.

Monitoring p(x) alone cannot distinguish the two target worlds: both contain 20% A and 80% B. This is an information limit, not something a more sophisticated detector on the same inputs fixes. New labels, verifiable outcomes or reliable rule-change information are needed. Delayed labels delay evaluation; labels collected only for disputed cases introduce a selection process rather than automatically representing all tickets.

Even in the favorable case, a few examples can carry too much weight

Weights are exact because we constructed the distributions. In practice they must be estimated, and source-rare categories may get large weights, amplifying label mistakes. A descriptive concentration measure is n_eff = (Σ_i w_i)² / Σ_i w_i². With 800 weights of 0.25 and 200 of 4 it is 1000² / 3250 ≈ 307.69. No tickets were deleted, and this alone is not a confidence interval: it indicates uneven influence.

Capping weights may reduce instability but breaks the exact identity derived above; it is a tradeoff to evaluate, not a free improvement. Better vocabulary representations may help distinguish previously confused categories, but do not establish stable business rules. Document updates, label revisions and population corrections affect different components. Confusing them makes success or failure hard to explain.

What a vertical-model project should take away

Before claiming domain adaptation, state what changed and what was verified. Weighting exactly corrects 16% to 34% only under conditional stability; when the rule changes, true error is 50% and new information is needed. Sector vocabulary does not replace this diagnosis. Useful evaluation reports population, period, label definition and transfer assumptions alongside the score. We measured no EL-AI model; we made a condition of specialization projects auditable.

Sources and experimental scope

Sugiyama, Krauledat, Müller (2007), Covariate Shift Adaptation by Importance Weighted Cross Validation, JMLR 8:985–1005, sections 2–4.1.

The 2007 paper studies weighted cross-validation under stable conditionals and, theoretically, known density ratios. Section 4.1 fits a linear model to a sinc target using 150 samples, ten-fold validation and 1000 repetitions. The authors report improved risk estimation over unweighted validation in that protocol. We did not reproduce it: our discrete example verifies the identity for two fixed classifiers and violates its assumption in a second case. This is neither new research nor a general language-network result.

The snippet computes U’s three errors; the archive adds V, the joint table, total weight, weight concentration and assertions. Computation scales with type-label cells without numerical optimization. A future real-ticket test should define operational labels, separate model selection from evaluation and check new-case coverage. That test was not performed here.

from fractions import Fraction as F
source = [F(4,5), F(1,5)]
target = [F(1,5), F(4,5)]
old_error = [F(1,10), F(2,5)]
new_error = [F(9,10), F(2,5)]
weights = [t/s for t,s in zip(target,source)]
print('source', float(sum(s*e for s,e in zip(source,old_error))))
print('weighted', float(sum(s*w*e for s,w,e in zip(source,weights,old_error))))
print('changed rule', float(sum(t*e for t,e in zip(target,new_error))))

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.