Skip to main content

Memory retrieval

Durable memory is not the same thing as prompt context. Memory is the complete, append-only record of what a RuntimeAgent has retained. Prompt context is a bounded, audience-appropriate selection from that record for one turn. Treating those as the same set makes every remembered fact appear in every conversation until an arbitrary recency limit evicts it. That is bounded storage access, not recall. This document owns the contract for selecting durable memory into model context. context-resolution.md owns the broader perspective and audience boundary; ../versioned-state.md owns canonical storage and revision history. prompt-budget.md (repository) owns the final whole-request admission boundary: retrieved evidence competes with history, tools, images and output reservation, and optional context may shrink again during later tool rounds.

Three independent axes

Every durable memory needs three concepts that must not be collapsed:
  • subject — who or what the claim concerns
  • source — who or what supplied the claim
  • visibility — which audiences may receive the claim
A preference such as “the CTO prefers tabs” illustrates why all three matter. The CTO may be its subject. The CTO, another actor, a repository policy, or a source message may supply it. The fact may be reusable across a workplace, or it may have been learned in a private conversation that does not permit wider use. The rule is:
Subject changes relevance. Visibility changes eligibility. Source explains provenance.
A subject match is a positive retrieval signal, never an authorization boundary. If another engineer asks whether to use tabs or spaces, a memory about the CTO’s preference may be relevant even though the current speaker is not its subject. Conversely, semantic relevance never makes a private memory eligible for a wider audience. Subject and source identity come from the perspective-scoped conversation graph, not from claims inside message text. Visibility comes from retained source and audience provenance. These are structured references, not prose conventions.

Store references, render labels

Memory bodies remain natural prose. Actor labels such as Victor (id U123) are rendered when context is assembled; they are not embedded as a storage format inside the body. Persisting labels inline would duplicate mutable display names across many rows, freeze one prompt presentation into canonical data, and leave nothing honest to write before identity resolution completes. A structured nullable subject reference represents “not resolved yet” directly and can be repaired without parsing or rewriting prose. This is the same separation used for replies: stored content is the canonical fact, while presentation is derived from current identity data. A display-name change therefore changes future rendering without rewriting memory history.

Hard eligibility precedes ranking

Retrieval begins inside one RuntimeAgent perspective. It must reject ineligible rows before computing relevance:
  1. the memory belongs to the current RuntimeAgent
  2. the entry is live and its selected revision is current
  3. source and audience provenance permit reuse in the current conversation
  4. any explicit policy restriction is satisfied
No ranking score may restore a row rejected here. In particular, a high semantic similarity score does not override a private-conversation boundary.

Hybrid candidate retrieval

Eligible memory is selected through a union of complementary candidate sources:
  • subject candidates — memories about the current actor, positively boosted but not exclusive
  • conversation candidates — memories tied to the current conversation or its durable context
  • lexical candidates — PostgreSQL full-text search for exact phrases, names, identifiers, acronyms, and other token-sensitive matches
  • semantic candidates — vector similarity for paraphrases and conceptual relevance
  • always-on candidates — explicit high-salience instructions or preferences whose contract requires them on every applicable turn
The union is deduplicated by memory entry and revision before reranking. Lexical and semantic retrieval are complementary: neither is a fallback for the other. Exact identifiers and negation-sensitive wording are often lexical strengths; paraphrases and topic similarity are semantic strengths.

Explainable reranking

A bounded candidate set is reranked from inspectable features such as:
  • semantic relevance to the current message and recent conversation
  • current-subject match
  • current-conversation match
  • explicit salience
  • confidence
  • recency decay
Subject mismatch carries no automatic penalty. A subject match adds evidence; its absence does not imply irrelevance. Weights are configuration and evaluation results, not architectural truth. Every selected memory should retain the component scores and the reason it was chosen, so retrieval misses and surprising inclusions can be diagnosed from durable state rather than guessed from a final prompt. The resolver injects the highest-ranked eligible memories under both a count and token budget. Rendered memory includes enough subject and provenance information for the model to distinguish “Victor prefers tabs” from “the current speaker previously told me this.”

One retrieval service

Automatic prompt assembly and an agent-facing recall(query) tool use the same retrieval service. They may request different budgets or query text, but they do not define separate eligibility rules or ranking algorithms. Tool results return stable memory entry and revision identifiers together with subject, source, visibility-safe provenance, and scores. Those identifiers are what later amend, supersede, or forget operations target.

Truthful confirmation

A reply may describe only durable operations that actually committed. This is a retrieval-adjacent contract rather than a prompt-style preference, because the failure it prevents is a lie about canonical state. Observed in production: after two remember calls an agent told the user it had “re-attributed” the earlier memory and that the earlier entry was “marked as superseded”. The database held two independent live revision-1 entries; the first was untouched. The agent also claimed a shared workspace had forced it to store the memory anonymously, which this runtime never does. Three mechanisms enforce it, in order of preference:
  1. Verbs that cannot misrepresent themselves. remember creates and only creates. Its argument schema has no entry id, no revision, and no supersede flag, so there is no request through which a model could ask it to modify an existing entry — and therefore no committed result that could honestly be described as a supersession. Correcting an entry is amend; retiring one is forget.
  2. An authoritative committed ledger. Every durable write records itself from the store’s return value, never from the model’s request. Before the reply is generated, the complete list is injected as a developer-role instruction naming each operation, its stable entry id, and its resulting revision — plus explicit prohibitions on claiming supersession, re-attribution, deletion, or anonymous storage when no such operation committed. Prevention is preferred over detection: a model handed a complete, authoritative account of what happened has no remaining reason to narrate something else.
  3. A narrow claim guard. A final reply asserting a durable mutation that did not commit fails the turn. It matches first-person past-tense assertions and deliberately does not attempt to parse negation, hypotheticals, or reported speech — prose is not a security boundary. It exists so that ignoring (2) is loud rather than silent.
Re-attribution and supersession claims become permissible only after a successful amend or forget, and a re-attribution claim specifically requires an amend that actually changed the subject. A body-only correction commits an amend and still leaves “I re-attributed it” false. A tool call that changed nothing — an unknown or already-forgotten entry id — returns a result saying so and contributes nothing to the ledger, so it cannot become licence to claim the write happened.

Bounded tool rounds

Recall must be able to precede a correction: the model cannot amend an entry without first obtaining its id. That requires more than one tool round, which is not the same as an unbounded agent loop. The bound is three rounds and it is structural rather than advisory. On the last permitted round the tool list is removed from the request, so the model is not asked to stop — it is unable to continue. Fan-out inside a single round is bounded separately. Every call in a round is validated before any is executed, so a malformed second call cannot leave a committed first write behind from a round the runtime then rejects. A durable write the reply never learns about is exactly the state the ledger exists to make impossible.

Embeddings are derived indexes

PostgreSQL remains the canonical memory store. A vector is a disposable retrieval index derived from immutable memory revision text, not a second source of truth. Embeddings attach to memory_entry_revisions.id, not to the mutable live entry. An amendment creates a new revision and a new embedding. Historical embeddings remain audit history and are excluded from current retrieval by revision state. Model and dimension metadata live with each embedding so changing embedding models does not rewrite canonical memory. The default implementation is PostgreSQL full-text retrieval. Optional pgvector and Qdrant adapters add semantic candidates through the same scoped interface. Embeddings are independently pluggable. See memory providers (repository) for configuration, rebuilds, and the adapter contract. Learned rerankers, clustering, and memory trees remain future work that requires measured quality improvements.

Evaluation contract

Retrieval changes are tested against cases that distinguish eligibility from relevance:
  • the current actor asks about their children; the subject boost retrieves the relevant fact
  • another actor asks about tabs versus spaces; the CTO preference remains retrievable with its subject and source intact
  • a coding question does not select an unrelated favorite-food memory
  • a food question does select that favorite-food memory
  • a private direct-message memory is never eligible in another actor’s turn, regardless of semantic similarity
  • an amended memory exposes only its current revision to live retrieval
  • an actor rename changes rendering without rewriting memory text
  • two actors with the same display name remain distinct through stable ids
Quality should be measured on a corpus with known queries before semantic weights are treated as reliable. Hybrid retrieval earns its complexity by improving measured recall and precision, not by making the schema look sophisticated.

Current implementation

Migration 0038_memory_audience_retrieval.sql adds source audience evidence, revision-scoped sharing grants, an index outbox, embedding generations, and retrieval audit records. The original entries, revisions, actor references, and API shapes remain intact. Memory privacy (repository) defines the enforced policy; memory providers (repository) describes setup and upgrades. Automatic context and recall share memory.Store.search. Authorization in SQL precedes candidate selection and all text is hydrated from live canonical rows. PostgreSQL full-text terms retrieve lexical candidates across the archive, rather than just the latest 100 memories. Nonempty queries require lexical or semantic relevance; subject and conversation matches boost those candidates. Empty queries rank by subject, conversation, and recency. Automatic broad questions such as “what’s new?” use that empty-query mode; the recall tool also supports it. Lexical and semantic ranks are fused with reciprocal-rank contributions. Subject and conversation matches add explicit boosts. Each source contributes at most 200 candidates. Mesh caps valid, distinct semantic hits before canonical hydration, even if a provider ignores the requested limit. Semantic search has a three-second deadline and a 10,000-revision eligible-set cap; unavailable, oversized, or rebuilding indexes degrade to authorized lexical results. These are conservative starting weights, not a claim of corpus-tuned retrieval quality. Explicit salience, confidence, learned reranking, and model-specific tokenization remain future work. Complete entries are packed into a configurable memory allowance (default 4096), counting the larger of prompt and recall JSON representations plus attribution overhead as a conservative text-token upper bound. The fixed section/tool header is outside this entry allowance. Oversized entries are skipped rather than truncated into changed claims. This allowance bounds memory context, not the entire model window; history, attachments, and tool results retain their separate existing limits. memory_retrieval_events records a query digest, destination audience, selected revision IDs, score components, conservative cost, and degradation category. It does not store query text or provider error bodies. The digest is a keyed HMAC under a subkey of MESH_AGENT_MASTER_KEY (internal/auditdigest), not a plain hash: recall queries are a few words about a person or a project, and a plain SHA-256 of them is reversed by hashing guesses. Without a master key no digest is recorded. Rows written before migration 20260925202721 held the plain hash; that migration cleared them. Prompt snapshots and tool observations retain the actual context under the existing operator trust boundary. When lexical matches are excluded with unknown source audiences, the audit records eligibility_withheld_unknown even if other memories were returned. A broad term can match several general facts while the specific answer is withheld. An earlier semantic failure category retains precedence; this diagnostic never authorizes disclosure or puts withheld text into the tool response. Legacy audience repair runs through forward migrations. Migration 0089 repairs current memory evidence skipped by 0085’s reliance on the optional, unpopulated runtime_agents.actor_id. It requires an inbound source with an exact connector identity and proven direct topology; other participants must be recorded outbound voices, with no contradictory inbound senders. For old Slack threads without stored topology, a later authenticated direct observation of the same connector/channel and counterpart can establish the fixed DM. Channel ID prefixes and speaker counts alone are insufficient. Missing proof leaves evidence unknown for operator review. Repair preserves revision bodies, source references, conversation restrictions, sharing grants, and historical transcript evidence. Every candidate is revalidated after optional provider work. Before final reply delivery, Mesh rechecks the audience and the used memory revisions. A change withholds the response. A turn can acknowledge its own successful amendment or forget operation; that operation does not remove earlier sources from the lineage of subsequent memories. New memories conservatively inherit all retrieved memories and loaded history message sources from the turn. Amendments retain earlier evidence. The source is connector-authenticated, never a model argument, and foreign or missing parent references reject the entire write. This can withhold an otherwise shareable summary when it has mixed sources; explicit review is preferable to silently losing provenance. remember, recall, amend, and forget retain their existing contracts and run budgets. A fifth tool, request_memory_sharing, can propose release of one revision in an authenticated private conversation. Only a matching authenticated source actor confirmation or an explicit authenticated operator review creates a grant. A model classification cannot create a grant.