ELAI S.r.l.

When a document tries to command an agent: separating data, instructions and permissions

A local experiment explains how to control AI agent tools: provenance, recipients, information flows and the limits of guarantees.

When a document tries to command an agent: separating data, instructions and permissions

Reading a document does not authorise it

Imagine an AI agent asked to read an invoice and add a note for the finance department. The document also contains a sentence asking it to change destination or delete a record. A person can usually recognise the invoice as material to examine, not a supervisor assigning new tasks. A system mixing instructions and retrieved text in one conversation needs to make this distinction an architectural property.

The central question is whether we can limit the consequences of a hostile document even when a model misinterprets it. We build a small controller checking three conditions before a simulated write: permitted operation, authorised destination and recipient allowed to read all data used. We prove a permission-propagation property and test eight cases. No model is queried and no message is sent: this tests a software contract, not an LLM’s attack resistance.

The boundary between a proposal and an action

Here an agent combines a language model with software tools such as search and record updates. Prompt injection attempts to turn unauthorised content into instructions for the system. The text may be grammatical and arrive through a source we legitimately read; the problem is the authority assigned to it. Permission to consult a document is different from permission for that document to decide our actions.

A proposal is a candidate call produced by a model: tool name and arguments. Execution is when the system causes an effect. Our controller sits between them. Instead of judging whether a sentence sounds suspicious, it checks properties defined by the task and permissions. The only authorised effect in this exercise is an internal note to finance, a symbolic department identifier, containing only information that department may read.

Fix the assumptions too. The initial task, controller and document metadata are trusted; document content and linguistic outputs are not. Every write passes through the controller. An attacker cannot modify Python code or manufacture authorisation labels. These are assumptions of the teaching model, not guarantees of the dataclass used below. If the model also has an unrestricted shell, this example is not a sandbox that makes it safe.

Every value carries two additional facts

Text alone says neither who may read it nor where it came from. Associate each value v with allowed readers R(v) and sources S(v). Invoice document a is readable by finance and manager; restricted budget b only by manager. These names represent identities already resolved by the system, not words a document can invent to obtain access. The code uses frozenset, whose elements cannot be modified in place.

R(a) = {finance, manager} R(b) = {manager} R(c) = R(a) ∩ R(b) = {manager}

The ∩ symbol denotes intersection: retain only readers allowed by both inputs. If a summary contains invoice and budget information, finance cannot read the summary because it could not read the second document. Union would be wrong: it would extend budget access to someone authorised only for the invoice. Combining values must not become a shortcut for expanding permissions.

R(f(v₁,…,vₙ)) = ⋂ᵢ R(vᵢ) S(f(v₁,…,vₙ)) = ⋃ᵢ S(vᵢ)

For sources use union, ⋃: the result retains all origins it depends on. f represents a transformation that correctly declares every input, such as concatenating texts or producing a summary. Our controller uses R for access decisions and records S to explain them. Known provenance does not make content true: an identified source can still contain an error.

Intersection is conservative. If a summary reads a restricted document but ultimately repeats only a public sentence, this mechanism still retains the restriction. It has not proved that the secret did not influence the sentence choice. Relaxing the label requires a further rule, declassification, established by an authority other than input text. We do not implement it: transformations cannot expand R in this exercise.

Three conditions before writing

A candidate call contains operation o, destination d and body b. The first check permits only append_internal_note. The second requires d to be finance selected by the trusted task, not a destination extracted from the document. The third requires finance to belong to R(b). All must pass: an approved destination does not authorise every body, and readable content does not authorise every operation.

allow(o,d,b) = [o = append_internal_note] ∧ trusted_destination(d) ∧ [d = finance] ∧ [d ∈ R(b)]

∧ means “and”: one false condition blocks the call. trusted_destination is not a language classifier; in the script it is an attribute assigned only by the trusted context. In C4 the string finance comes from the document and is rejected despite matching the expected name. This is intentionally restrictive: retaining the original destination removes any need to accept destination proposals from consulted material.

The check must precede the effect. Only when allow is true does the script append to mock_log, an in-memory simulated write. We call no remote service and promise no distributed transactions. Production would need to prevent alternate tool paths and ensure the object, destination and permissions do not change between check and execution. A correct check on stale data does not protect the actual action.

A small proof: permissions do not expand

We can say more than “the eight cases work”, but only under the stated assumptions. Consider a chain of transformations always using intersection and recording every input. For a final value v, every allowed reader must be allowed by each source on which v depends. This holds for initial values because their metadata are assumed correct. We now show why every transformation preserves it.

Suppose result c depends on a and b. By definition, an identity r in R(c) belongs to both R(a) and R(b). If the property already holds for a and b, r may read all their respective sources. It may therefore read all sources of c, their union. Repeating this step gives an inductive proof. The final check r ∈ R(c) prevents explicit delivery to a reader excluded by one of c’s sources.

r ∈ R(c) ⇒ r ∈ R(a) ∧ r ∈ R(b) R(c) ⊆ R(a); R(c) ⊆ R(b)

The conclusion concerns explicit flow of labelled values. It does not prove the whole program vulnerability-free, metadata authentic or every information channel accounted for. Effectiveness depends exactly on the assumptions: one forgotten input, bypassed check or incorrect initial label falls outside the proof. Its value is identifying where system review must focus.

Eight candidate calls and the reason for each decision

The experiment constructs candidate calls directly rather than asking a model to generate them. C1 adds an invoice summary to finance and passes. C2 tries a different operation and is blocked. C3 proposes an outside destination; C4 proposes finance extracted from the document. Both fail the destination rule. We do not measure how many attempts an attacker might invent or the probability of a model following them.

C5 tries to insert the restricted budget directly; C6 combines it with the permitted summary; C7 paraphrases its amount in words. All fail the reader check: changing form does not change permissions. C8 invents a total of 999 euros instead of the document’s 120. It passes because operation, destination and read access are valid. This prevents presenting the policy as a truth-verification mechanism.

CaseOperationDestinationReadersAllowed
C11111
C20110
C31000
C41010
C51100
C61100
C71100
C81111

In the table, 1 means a passed check and 0 a failed check; the final column is their conjunction. The simulated log contains C1 and C8, while six calls are blocked. “Six out of eight” is not an effectiveness rate against prompt injection: cases are hand-selected, not a representative attack sample. The reproducible result is that code implements its declared policy, including its inability to recognise the invented total.

Check matrix calculated by the script. Each row is a synthetic call, not an attack observed on a model. C8 passes like C1 despite its false amount: authorisation and content accuracy are different properties.
Check matrix calculated by the script. Each row is a synthetic call, not an attack observed on a model. C8 passes like C1 despite its false amount: authorisation and content accuracy are different properties.

The fragile point: losing labels during a transformation

Add a counterexample outside the eight cases: a component takes the paraphrased budget, extracts its plain string and creates a new object with the invoice’s readers. The controller now allows it. It has not discovered harmless text; someone discarded the information needed for the decision. The script deliberately includes this broken component and verifies that its outcome differs from correct propagation.

Serialisation, caches, service boundaries and error handling therefore belong inside the protected boundary. A final-function check is insufficient if earlier stages rebuild everything from strings without provenance. Making a Python object “immutable” also falls short: arbitrary code can create another object. A real system must assign and preserve metadata through trusted components separated from the model’s ability to propose text.

What this experiment cannot guarantee

Implicit dependencies also exist. If whether a note appears depends on restricted information, an observer can learn something even when its body is a public constant. Tracking concatenated text alone is insufficient. Our example does not interpret arbitrary conditionals, control execution timing or error messages, or prove these channels absent. The proved property is deliberately limited to declared explicit flows.

We have not verified a model-generated plan either. If the initial task is misread and policy is too permissive, the controller can approve an unwanted action. Defining permissions is part of the problem, not a post-development formality. Our case is simple because we choose one operation and a fixed destination; open-ended workflows require explicit additional objects, roles and conditions.

A language filter and permission controller answer different questions. One tries to recognise problematic content; the other checks a concrete proposal against a contract. They can coexist, but our test attributes no unmeasured effectiveness to a filter. At the opposite extreme, blocking every write would prevent legitimate C1 too. Architecture comparisons need both violations and useful task completion, with documented costs and contexts.

The research connection and verification path

CaMeL, Debenedetti and colleagues’ preprint v2 dated 24 June 2025, separates planning from untrusted reading and applies flow policies. Its AgentDojo evaluation distinguishes utility and security; the authors discuss limitations and side channels. Our controller is not a reproduction.

The attached code is entirely local and uses Python’s standard library; Matplotlib is only for the figure. derived preserves reader intersection and source union; policy returns three booleans and their conjunction; run executes cases and asserts expected outcomes. The simulated log shows that only C1 and C8 cause an effect. The broken-component counterexample is separate and explicitly violates the assumptions.

We report no agent latency, token cost or attack-prevention percentage because none was measured. Small-set intersection is a different software cost from querying a model; a complete system adds identity resolution, permission retrieval, propagation across services and event recording. For EL-AI this is AI-assisted technical analysis, not an announcement of an available security feature or completed audit.

The answer: reading does not grant unlimited action

A document may propose instructions, but that must not grant new operations or destinations. In our restricted model, a fixed task, permissions travelling with data and checks before effects block specific classes of unauthorised calls. The guarantee ends where provenance is lost or channels are uncontrolled. It remains separate from summary correctness: an agent must both respect permissions and report grounded information. C8 shows why neither property replaces the other.

Bibliography

Edoardo Debenedetti et al. — Defeating Prompt Injections by Design, arXiv:2503.18813v2, 24 June 2025, preprint.

Google Research — CaMeL research artifact.

from experiment import run
r = run()
for c in r['cases']:
    print(c['case'], c['checks'], c['allowed'])
print(r['counts'])
print('Broken wrapper:', r['broken_wrapper']['allowed'])

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 4 October 2026.