ELAI S.r.l.

The agent remembers correctly, but the rule expired: putting time into memory

Relevance does not imply current validity. An executable example separates effective dates, system knowledge, late updates and conflicts.

The agent remembers correctly, but the rule expired: putting time into memory

A documented answer can be out of date

An agent answers an attachment-limit question with 100 MB, citing an authentic, highly relevant document. But the document announces next month’s rule; today the limit is 80 MB. This need not be model fabrication: memory may retain content while losing its temporal meaning. How should retrieval prevent same topic from being mistaken for valid for this question?

Abstract. Build memory with two axes: when a fact applies and when the system holds a particular version. A wholly synthetic upload limit changes, is learned late, and has a future announcement. Derive and execute temporal filtering in SQLite, distinguishing supported answers, missing information and conflicts. No language model or embedding engine is evaluated; we test the deterministic evidence-preparation layer for an agent.

Two dates answer two different questions

Validity time asks when a rule applies; knowledge time asks when this version was in the system. Suppose a limit changes from 50 to 80 MB on September 1, but ingestion occurs September 3. World and archive disagree on September 2. Reconstructing the rule for that day today gives 80; reconstructing what the agent could read then gives 50. Both are useful but different reconstructions.

Add a third event: on September 20 the system ingests an announcement of 100 MB starting October 1. The latest ingested document now concerns the future. Choosing the newest timestamp does not solve the current question either. Names, values and documents are synthetic, not EL-AI limits or features. MB is a policy value here, not measured upload performance.

Preserve versions without rewriting the past

Each row keeps ID, value, source, validity start/end and knowledge start/end. Half-open intervals include the start and exclude the end, so old and new rules meet September 1 without boundary overlap. We use daily dates at midnight UTC; real systems must choose precision and time zone first. 9999-12-31 represents an unknown end, not a promise of millennia-long validity.

When the update arrives September 3, retain the old snapshot and close its knowledge interval then. Add a corrected snapshot of the old rule ending September 1, plus the 80 MB rule. On September 20 likewise close the open-ended 80 MB snapshot, replace it with an October 1 end and add the future rule. This records changing knowledge: retrospective correction must not pretend earlier awareness.

The filter turns the question into four comparisons

Let t be the date being asked about and k the knowledge cutoff. A row qualifies if t falls within validity and k within knowledge. Both are today for a current answer; an audit may use different dates. Keeping them separate prevents silently using later updates to judge what the system could know earlier.

valid_from ≤ t < valid_to known_from ≤ k < known_to

Semantic retrieval can identify topic, but similarity does not replace temporal constraints. Assign synthetic relevance 0.97 to old text, 0.93 to current and 0.99 to future announcement. Score-only ranking selects 100 MB; filtering for September 29 first selects 80 MB. These are neither truth probabilities nor embedding-model outputs; they isolate a selection error independent of the semantic engine.

Valid date tKnowledge kMB valueInterpretation
2026-09-292026-09-2980current
2026-08-202026-09-2950historical
2026-09-022026-09-2980retrospective
2026-09-022026-09-0250known_then
2026-10-022026-09-29100announced_future
2025-12-012026-09-29—missing

known_then means the September 2 archive reported 50, not that the real limit was 50. The future row means 100 is announced from October 1, not that 100 MB is allowed today or guaranteed to happen. Missing data is not zero. A reliable answer preserves these qualifications in natural language instead of giving the model a bare number.

A conflict is not resolved by averaging

Insert another source claiming 90 MB during the same interval as 80. Assume equal scope and no documented precedence. Both pass temporal filtering. Collect distinct values: none means missing, one means consistent support in available material, more than one means conflict. Do not average to 85 or choose 90 merely because it was ingested later. Legitimate precedence needs explicit authority or supersession rules, not pressure to produce an answer.

supported does not mean certified truth. Available relevant eligible rows merely agree on a value. Two copies are not independent confirmation; an authoritative source may be wrong; an update may not yet have arrived. Deterministic checks must remain separate from evidence trust. An agent should qualify available-source claims and expose source, version and reference date.

From the logical rule to an executed query

SELECT id, value_mb, source FROM memory WHERE valid_from <= :t AND :t < valid_to AND known_from <= :k AND :k < known_to ORDER BY id

The query selects evidence, not prose. ORDER BY id makes output reproducible without ranking authority. Bind parameters rather than concatenate SQL text. The in-memory SQLite experiment starts from the same five synthetic rows, then separately adds conflict. A boundary check with September 3 knowledge returns 50 for August 31 and 80 for September 1, directly testing exclusive end boundaries.

A direct scan needs O(n) comparisons before sorting. Large archives can index entity keys and intervals, but selectivity and query plans determine actual cost; no performance measurements were made. Keys must identify product, environment and rule scope: mixing two services creates artificial conflict. Our single-service fixture isolates time. More documents do not resolve unmodeled scope.

Rule reconstructed with September 29 knowledge: 50 MB before September, 80 during September, 100 announced from October. Future segment is dashed; vertical line marks answer date. Synthetic data, not EL-AI site limits.
Rule reconstructed with September 29 knowledge: 50 MB before September, 80 during September, 100 announced from October. Future segment is dashed; vertical line marks answer date. Synthetic data, not EL-AI site limits.

Agent memory also includes copies and summaries

An update may fix the main document while leaving old vector-store chunks, an agent summary and a cached answer. If summaries lose source IDs, invalidation targets become unknown. Preserve derivation: which version produced each fragment or summary. Provenance is not truth; a transformation chain does not guarantee correct extraction but permits rechecking after input changes.

W3C’s PROV Primer describes entities, revisions, derivations and generation times. SQL Server documentation shows historical database queries with FOR SYSTEM_TIME AS OF. Neither replaces application-defined rule-effective time. Our prototype explicitly manages both intervals; SQLite does not maintain them automatically.

Retrieval must also cover the question. Taking only the closest document and rejecting it as future can yield missing despite a current rule in the archive. Filter before final selection or expand candidates when needed. This avoids that specific loss but does not guarantee ingestion completeness. For change-sensitive current questions, direct authoritative-source checks may be needed; local retrieval success does not prove global freshness.

The answer preserves date, source and knowledge limits

Give the agent a structured object with value or status, eligible sources, requested validity date and knowledge snapshot. Expose conflict alternatives; on absence seek evidence or state the limit. Generated prose still needs checking: correct filtering cannot prevent omission of from October 1. A final check could compare values and temporal qualifications against the deterministic object. This extension is proposed, not implemented here.

Useful memory stores more than what was said: it preserves when it applies and when the system learned it. Relevance selects topic, validity selects period, provenance makes reconstruction auditable. Our experiment distinguishes 80 MB today, 50 historically and 100 only in the future announcement, without asking a language model to guess which document wins.

Sources and reproducibility

W3C — PROV Model Primer, Working Group Note, 30 April 2013, sections 2.6–2.9.

Microsoft — Query data in a system-versioned temporal table, AS OF and historical queries.

The snippet shows one knowledge snapshot for a readable validity filter. The archive contains the complete two-axis version, five rows, six questions, conflict and boundary checks, executed in Python and SQLite. No data was sent to an external model. Results describe this deterministic fixture, not general agent reliability or an available EL-AI feature.

records = [
    dict(value=50, start='2026-01-01', end='2026-09-01'),
    dict(value=80, start='2026-09-01', end='2026-10-01'),
    dict(value=100, start='2026-10-01', end='9999-12-31'),
]
for date in ('2026-08-20', '2026-09-29', '2026-10-02'):
    values = {r['value'] for r in records if r['start'] <= date < r['end']}
    status = 'missing' if not values else ('conflict' if len(values)>1 else 'supported')
    print(date, status, sorted(values))
# One knowledge snapshot only. Full archive implements both temporal axes in SQLite.

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.