The cost is not taking the image but giving it meaning
An agricultural business collects many inspection images but can ask experts to label only some of them. Labelling assigns a category or other verified information needed to train and evaluate a model. Will choosing images where AI is most uncertain use that time well? Perhaps. Yet two consecutive photographs of one leaf may both be difficult while describing almost the same problem, leaving other imaging conditions unexplored.
Our question is narrow: with a budget of two new labels, how can we avoid spending both in an already represented region? Six imaginary images become points on a plane. We compare uncertainty selection with geometric coverage, derive a guarantee and expose its limits. Coordinates and probabilities are synthetic: they contain no agronomic measurements, diagnose no plants and demonstrate no EL-AI project result.
A map of images, not of the field
An embedding is a numerical representation: a network can turn an image into a vector. If that representation suits the task, images similar in relevant features may have nearby vectors. The “if” matters: proximity might reflect lighting and background instead of the property to recognise. We use two abstract components in arbitrary units to make the reasoning visible. They are neither geographical coordinates nor a projection of real images.
L is an already labelled image at the origin. A and B are near L; C and D form a second region; E and F a third. Each candidate also receives synthetic confidence for binary classification. Highest uncertainty here means winning confidence nearest 0.5. All confidences are at least 0.5, so this selects the two smallest values, A and B. The rule ignores similarity between the two selections.
| Point | x₁ | x₂ | Confidence |
|---|---|---|---|
| L | 0 | 0 | — |
| A | 0.2 | 0 | 0.5 |
| B | 0.3 | 0.1 | 0.51 |
| C | 4 | 0 | 0.8 |
| D | 4.2 | 0.2 | 0.78 |
| E | 0 | 4 | 0.9 |
| F | 0.2 | 4.1 | 0.88 |
What covering a set means
Connect each point to its nearest labelled representation. The longest segment identifies the worst represented region. Coverage radius is that length: minimising it seeks a set leaving no candidate too far from a chosen example. It does not reveal neighbours’ labels or guarantee a shared class. It is a geometric surrogate, useful only when the geometry is relevant.
U contains the six candidates and S the two to label. Distance d is Euclidean. The minimum chooses each point’s nearest centre; the maximum chooses the worst point. Selected images have zero distance to themselves, and L remains available without consuming budget. These details prevent forgetting existing labels or optimising average distance while claiming a worst-case guarantee.
With S = {A,B}, F remains √[(0.2−0.3)² + (4.1−0.1)²] = √16.01, about 4.00125, from nearest centre B. This is the final radius. Two labels near the origin cover one region closely and nearly ignore the others. It is a constructed example, not proof that uncertainty sampling is always wrong: if an important class boundary lies there, those labels could be valuable.
Choose the least represented point at each step
Farthest-first is greedy: it makes the locally preferred choice without enumerating all future combinations. Starting with L, compute each candidate’s nearest-centre distance, select the largest and add it to the centres. Update distances and repeat until the budget is spent. Updating matters: after selecting D, nearby C must not retain its old priority.
Initially D is √17.68 ≈ 4.20476 from L, farther than any other candidate, so it is selected. C becomes about 0.28284 from D, while F remains about 4.10488 from L and becomes the second choice. With {L,D,F}, the largest residual distance is B to L, about 0.31623; C to D and E to F are smaller. Radius thus falls from uncertainty selection’s 4.00125 to geometric selection’s 0.31623.

How good is this selection? A bounded guarantee
We can compare every pair: six candidates produce 6 × 5 / 2 = 15 combinations. The script enumerates them and finds {C,E} among the optima, with radius 0.31623. Greedy reaches the optimum here, but need not always do so. The general guarantee is weaker: with a metric distance and fixed initial centres, greedy radius is at most twice optimal radius. Two concerns distances, not classification errors or labels saved.
Let b be the budget and r the final greedy radius. Take the b selected points plus a farthest point after the last selection. These b + 1 points are pairwise at least r apart: each new point was farthest from available centres, and adding centres never increases maximum distance. Each is also at least r from initial centres. Suppose r exceeds 2r_opt. None can then be covered by initial centres within r_opt. The b new optimal centres must cover all b + 1 points, so two share a centre. The triangle inequality bounds their distance by 2r_opt, contradicting their separation of at least r. This proves the formula.
The proof also prevents misuse. An arbitrary similarity need not satisfy the triangle inequality; squared Euclidean distance, for example, does not have that property in the same form. The proof knows nothing about labels, noise or plant diseases. Turning coverage into an accuracy promise requires additional assumptions about models and representation–class relationships, which we have not verified.
When diversity chases an isolated point
Add O = (7,7), with synthetic confidence 0.99. It could represent a different background, acquisition error or a genuinely important rare condition; geometry cannot decide which. Greedy selects O and D, leaving radius about 4.10488. Enumerating the 21 available pairs finds {B,O} among best coverings, with radius 4.00125. This is a concrete nonoptimal greedy result that still satisfies the proven bound.
| With O | New centres | Maximum radius | Mean distance |
|---|---|---|---|
| Greedy | O,D | 4.10488 | 1.27199 |
| min max | B,O | 4.00125 | 2.23669 |
| min mean | D,E | 7.35391 | 1.19666 |
The table averages all seven candidates, including selected ones contributing zero. Minimising this mean selects {D,E}: it represents most points well but leaves O far away. Minimising the maximum sacrifices average coverage to protect the worst case. No winner exists independently of the objective. If O is a capture fault, chasing it may waste budget; if it records an important rare condition, ignoring it may be the mistake. Data and purpose need checking, not automatic deletion of every anomaly.
Distance depends on how we represent images
Another test multiplies the first coordinate by 0.1. On the original pool without O, greedy now selects F then D: the final set is unchanged but order changes. New radius is about 0.20100. This is not an improvement over 0.31623: the measuring rule changed, so values are not directly comparable. Real embedding weights, normalisation and networks may amplify colour, texture or context. Standardising components alone does not guarantee agronomically meaningful distance.
Budget also needs definition. Here two images cost exactly two units. A class label and pixel-level outline can require very different effort; ambiguous images may need an expert. Optimising two items is not optimising twenty minutes. Unequal costs change the combinatorial problem, and the previous guarantee cannot automatically transfer to a distance-divided-by-cost heuristic.
Code, computational cost and field validation
The download includes coordinates, confidences, greedy selection and pair enumeration. The final snippet calls run and prints scenario, method, selections and radius; JSON also stores mean distances. No seed is needed: data and rules are deterministic, with alphabetical tie-breaking. No network is trained or real agricultural image acquired. The greedy/optimal radius assertion checks these finite cases; the proof explains the general result under its assumptions.
For N candidates, dimension d and b new labels, caching current minimum distances updates N distances per selection: O(Nbd), plus initial comparisons with existing centres. It uses O(N) extra memory for minima, without a full N × N matrix. Our readable script recomputes distances to all centres, costing O(Nb(|L| + b)d). Enumerating all pairs and evaluating N points costs O(N³d). These are operation counts, not device latency or memory measurements.
To establish agricultural utility, one must actually label data, train the same model under comparable budgets and include random selection. Tests should separate plots, seasons or acquisition sequences when those correlations matter: near-duplicate frames across training and test can make any strategy look effective. Multiple repetitions and recorded expert effort are needed. This is a proposed, unexecuted experiment; we report no accuracy gains, savings or EL-AI results.
Research and conclusion: coverage is one question, not the whole answer
Sener and Savarese, ICLR 2018, connect representative subset selection with active learning, analysing theoretical assumptions and testing robust optimisation with image networks. Here we study elementary geometry, without reproducing their networks, benchmarks or robust optimisation.
Research extends beyond k-center. Cohen-Addad and colleagues, ICML 2026, study selection and weighting using low-rank structure and loss sensitivity. Their objectives and assumptions differ from geometric coverage here; this does not make farthest-point selection universally optimal.
The initial question now has a precise answer: with a small budget, uncertainty alone can purchase nearly the same geometric information twice. Updating coverage after each selection avoids that specific redundancy and has a provable metric guarantee. It does not decide whether representations matter, rare points deserve attention or new labels will improve the model. The useful outcome is to separate and measure these decisions while keeping the agricultural problem at the centre.
References and reproducibility
Cohen-Addad et al. — ICML 2026, PMLR 306:21154–21175.
Cohen-Addad et al. — full methods and experiments, arXiv 2606.16045 v1.
from experiment import run
r = run()
for case in r['results']:
print(case['scenario'])
for m in case['methods']:
print(m['method'], m['selected'], round(m['radius'], 6))
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.

