One case, two routes, one more question
An AI agent must route a business case: continue ordinary processing or send it for manual review. It can query an additional record source, but every query consumes time and resources. Searching again sounds cautious; stopping sounds risky. The useful question is more precise: can the next answer change the choice enough to justify its cost? This article works through that calculation without assuming that more tool calls automatically make a better agent.
Abstract. We construct a synthetic decision problem with two actions, a hidden exception and imperfect tools. We derive the intervention threshold, update probabilities after an answer, and compare expected cost before and after a query. The first useful check takes cost from 12 to 8.88 units, including the query price. A less discriminating check has zero decision value in the same case. Finally, a strategy that decides whether to repeat the check according to the first answer reaches 8.448 units under stronger assumptions. These are executed educational calculations, not measurements of an EL-AI product or results from a language model.
Before the numbers: what the agent can decide
Percentages, weighted averages and the distinction between a fact and our knowledge of it are enough to follow the argument. Let H indicate an exception in the case: H is 1 if it exists and 0 otherwise. A check neither creates nor resolves the exception; it only provides evidence. The available actions are S, standard processing, and R, review. This separation matters: we are evaluating an information request, not a tool that directly modifies the case or executes a transaction.
Assign synthetic costs in a common arbitrary unit. Choosing S when an exception exists incurs a loss of 100; without an exception it costs zero. R always costs 12 and, within the model, prevents further loss. These are not euros, measured times or company rates. Assuming that review resolves the case simplifies the arithmetic: real review can make mistakes and would need a different cost table. The check price, fixed at 2, is also an educational choice and must use the same unit.
| Action | H = 0 | H = 1 |
|---|---|---|
| S | 0 | 100 |
| R | 12 | 12 |
Now introduce p, the probability that the case contains an exception before checking. Our running case uses p = 0.20: a hypothetical population with 20% exceptions. This is not the model declaring confidence, nor a percentage extracted from an assistant’s sentence. It is an assigned parameter in our experiment. In an application it would need estimation and validation on relevant cases, avoiding unverified transfer of a frequency observed in a different population. Assume it is known for now so we can isolate the decision problem.
What does deciding without more information cost?
The first formula asks a simple question: which action has the lower average cost with the information available? S has expected cost 100p because loss 100 occurs with probability p. R still costs 12. Call the smaller amount L₀. The zero subscript reminds us that no check has been purchased yet. The average describes repeated cases under the same assumptions; it does not guarantee the cost of an individual case.
Threshold p* is 0.12. Below 12%, S is preferable; above it, R is preferable; at the threshold the actions are equivalent in the model. With p = 20% we would therefore send the case for review. This gives us a concrete baseline for any proposed search. Without it, statements such as “the tool is accurate” or “the agent collected substantial evidence” do not establish that querying improved the outcome. We need a choice to improve and a consequence against which to measure it.
The tool answers: how our knowledge changes
Tool A returns a positive signal, +, or a negative signal, −. Positive means “evidence of an exception”, not “an exception is certainly present”. Define sensitivity s = P(+ | H = 1) = 0.8: 80% of actual exceptions produce a positive signal. Define false-positive rate f = P(+ | H = 0) = 0.1: 10% of ordinary cases nevertheless raise an alarm. The vertical bar means “conditional on”. These values are synthetic assumptions too, not claimed performance of an existing archive or service.
Before updating probability, we need to know how often each answer occurs. A positive may arise from a detected exception or a false alarm. Add these possibilities, which cannot both hold for the same case. Calling the probability of a positive answer z gives 0.24. Among 1,000 hypothetical cases distributed exactly according to the model, 160 would be true positives and 80 false positives: 240 positives altogether. This is an expected count constructed from the parameters, not an actually collected sample.
The final two lines apply Bayes’ rule: among all cases compatible with the answer, what fraction actually contains the exception? After a positive, 160 of the 240 expected cases are exceptions: p₊ = 2/3, about 66.67%. After a negative, 40 exceptions remain among 760 cases: p₋ = 1/19, about 5.26%. Only the negative branch crosses the 12% threshold. Thus + still leads to R, whereas − can lead to S. A negative signal does not certify absence of an exception; it changes the cost balance.
Information value is assessed before knowing the answer
When querying the tool, we do not know which branch will occur. We must weight the cost of the best decision in each branch by its probability. Let L₁ denote cost after observing the signal but before adding the query price. Gross value V is the reduction relative to L₀. “Gross” prevents a practical misunderstanding: information may help yet be too expensive. The total cost of the querying strategy is c + L₁, with c = 2.
The result means that, under ideal repetition in the assumed population, average cost falls from 12 to 8.88 units: net benefit 3.12. It does not mean every query saves 3.12. In the positive branch we pay 2 and still send the case for review. The overall gain comes from avoiding many reviews in the negative branch while accepting the model’s residual losses. If the price exceeded 5.12, checking would no longer be worthwhile in this one-query comparison. At exactly that price the strategies tie.
Why is gross value nonnegative? Because the agent could ignore the signal and retain its original choice. Formally, for any fixed action a, the minimum cost conditional on the signal cannot exceed the cost of a in that branch. Averaging and choosing a to be the original optimal action gives L₁ ≤ L₀. The argument requires that actions remain available, information can be ignored, and the check does not directly alter the case. A service that locks the case, discloses data or consumes a deadline is not described by informational benefit alone.
Knowing more can have zero value
Compare A with tool B, with s = 0.6 and f = 0.4. At the same initial p, a positive occurs with probability 0.44 and raises exception probability to 3/11, about 27.27%; a negative occurs with probability 0.56 and lowers it to 1/7, about 14.29%. The check distinguishes something: both probabilities differ from each other and from the initial value. Yet both remain above 12%. Either answer leads to R and cost 12. Thus L₁ remains 12 and V is zero.
This separates reducing uncertainty from improving a decision. Mutual information measures an average reduction in probabilistic uncertainty in bits; it does not know our cost table. For B in the running case, the calculation gives about 0.01864 bits but no decision benefit. Understanding entropy is not necessary to see why: the evidence shifts the estimate without crossing the operational threshold. B can have value with different initial probabilities or costs. We are not declaring it universally useless.
To reconstruct that number, define h(x) = −x log₂(x) − (1−x) log₂(1−x), with zero contributions at x = 0 and x = 1. Answer uncertainty is h(z); knowing H would leave the average p h(s) + (1−p) h(f). Their difference I(H;Z) is mutual information. For B it is h(0.44) − h(0.6), since h(0.4) = h(0.6). This detail shows that zero decision value is not caused by constructing a tool entirely independent of the problem.

In the plot, look at curve height at the point describing the case, not merely the globally highest curve. At p = 0.20, A lies above the cost line and B does not. Near the endpoints, when the state is already almost certain, little remains to gain. Changing error loss or review cost would change the curves. Tool reliability and the decision situation work together: a single tool ranking detached from the task would miss that dependence.
An upper bound: knowing the state perfectly
Before designing more expensive tools, ask what a perfect answer would be worth. Knowing H, we would choose S for ordinary cases and R for exceptions. Residual cost would be 0.8 × 0 + 0.2 × 12 = 2.4. Perfect information would therefore be worth 12 − 2.4 = 9.6 units. It is not worth 12, because knowing about the exception does not remove review cost. No purely informational tool about the same state can improve on this bound under our fixed actions and costs. Paying more than 9.6 would make no sense even before examining accuracy.
A second question: value depends on the first answer
We have so far allowed one check. Now allow at most two, both with A’s performance and cost 2. Add a strong assumption: responses are conditionally independent given H. Once exception presence or absence is fixed, knowing the first signal does not change the probabilities of the second. This does not mean responses are independent without knowing H. We might imagine two checks with separate error mechanisms, but this is only a probabilistic construction: two archives copying the same source do not automatically satisfy it.
After a first positive, we start from p = 2/3. Even a negative second check would only lower probability to 4/13, about 30.77%, still above threshold. A positive would raise it further. Both outcomes lead to R: the second check has zero gross value, so we stop. After a first negative, we start from 1/19. A second positive leads to 4/13 and therefore R; a second negative leads to 1/82, about 1.22%, and therefore S. In this branch the check really can change the action.
Work through the negative branch rather than hiding the benefit inside an algorithm. Without another check we would spend 100/19 = 5.26316 units on average. With the check, subsequent action cost, excluding the new query, falls to 2.69474; adding 2 gives 4.69474. The gain conditional on the first negative is about 0.56842. That branch occurs with probability 0.76, so the additional expected saving from the start is 0.432. The full strategy costs 8.88 − 0.432 = 8.448. Rounded numbers explain the arithmetic; the program retains exact fractions.
Write a general rule for a budget of n remaining queries. Wₙ(p) is the minimum achievable cost from the current point; with zero queries it is L₀. With at least one query, compare stopping now with paying c and continuing optimally after each answer. The sum over signals y includes + and −, and pᵧ is the updated probability. This is a finite-horizon recursion: it solves the specified small model, not every planning problem faced by a real agent.
This also exposes a limit of “ask only when the next answer pays for itself in isolation”. With multiple steps available, a first observation may prepare a more useful later check: the correct criterion must consider continuation, as Wₙ does. Our first check already pays for itself, but that coincidence is not a universal principle. Naive enumeration of two answers over h steps grows as 2ʰ; caching equivalent states avoids repeated calculations without generally removing planning difficulty.
The same answer repeated is not a second piece of evidence
If the second call returns exactly the same cached datum as the first, its outcome is already determined conditional on what we know. Exception probability does not change; additional information value is zero. Reapplying Bayes’ rule as if it were an independent test would count the same evidence twice. This also applies when the answer is rephrased. Before estimating another tool’s value, the agent must distinguish a new observation from a new presentation of the old observation.
Reproduce the calculation and understand what is missing
The attached package evaluates averages with Fraction, Python’s rational representation, and checks the main results using assertions. There is no random sampling, hence no seed is needed. The graph uses 1,001 p values from zero to one; decimal conversion is for export and display. Function risk chooses the smaller cost, branches computes answer probabilities and updated probabilities, voi subtracts expected costs, and value implements the recursion. The final snippet imports these functions from experiment.py: run it in the extracted archive directory, not as a program without local dependencies.
The main limitation is knowledge of parameters, not arithmetic. We need reliable probabilities for the current context, a coherent cost table and a joint error model when chaining tools. A probability confidently stated by a language model does not itself meet those conditions. Waiting cost may also depend on case state; review capacity may be limited; the state may change during search. Such cases require an expanded model, not continued use of the same numbers with precision they do not possess.
To transfer the reasoning to a business project, a proposed evaluation could compare three policies on cases separate from parameter-estimation data: no query, one fixed query and adaptive querying. Outcomes, costs, calls and source dependencies should be recorded, separating probability quality from decision quality. That evaluation has not been run here: no business archives, readers or operators were involved. The model is an application hypothesis for designing and evaluating agents; it does not document an already available EL-AI feature.
Answering the initial question
Ask when the expected improvement in future decisions exceeds the request’s total cost. In our example a useful answer is worth 5.12 units before its price; a less discriminating answer has zero value in the same situation; a second independent check pays only after the first negative. The aim is not to make an agent reluctant to ask. It is to connect each question to a choice that could change, and that choice to explicit consequences. Then we can explain why the agent continues or stops instead of measuring intelligence by query count.
Sources and nature of this contribution
Conceptual reference: David L. Poole and Alan K. Mackworth, Artificial Intelligence: Foundations of Computational Agents, third edition, 2023, sections 12.3 and 12.4. We consulted sequential-decision formalism, policy optimization and properties of information value. This is a textbook, not a new preprint or agent benchmark. This article’s administrative example, parameters, numerical derivation, plot and code were constructed for this explanation and do not reproduce the book’s examples. The article is an AI-assisted educational monograph, not original peer-reviewed research.
Poole & Mackworth (2023), §12.3 — Sequential Decisions.
Poole & Mackworth (2023), §12.4 — The Value of Information and Control.
from fractions import Fraction as F
from experiment import risk, branches, voi, value, STRONG, WEAK
p = F(1, 5)
print('baseline', float(risk(p)))
for name, tool in [('A', STRONG), ('B', WEAK)]:
print(name, [(float(w), float(q)) for w, q in branches(p, *tool)])
print('gross value', float(voi(p, *tool)))
print('one query', float(value(p, 1)))
print('adaptive two queries', float(value(p, 2)))
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 9 October 2026.

