Memory retrieval
Durable memory is not the same thing as prompt context. Memory is the complete, append-only record of what a RuntimeAgent has retained. Prompt context is a bounded, audience-appropriate selection from that record for one turn. Treating those as the same set makes every remembered fact appear in every conversation until an arbitrary recency limit evicts it. That is bounded storage access, not recall. This document owns the contract for selecting durable memory into model context.context-resolution.md owns the broader perspective and
audience boundary; ../versioned-state.md owns canonical
storage and revision history.
prompt-budget.md (repository) owns the final whole-request admission
boundary: retrieved evidence competes with history, tools, images and output
reservation, and optional context may shrink again during later tool rounds.
Three independent axes
Every durable memory needs three concepts that must not be collapsed:- subject — who or what the claim concerns
- source — who or what supplied the claim
- visibility — which audiences may receive the claim
Subject changes relevance. Visibility changes eligibility. Source explains provenance.A subject match is a positive retrieval signal, never an authorization boundary. If another engineer asks whether to use tabs or spaces, a memory about the CTO’s preference may be relevant even though the current speaker is not its subject. Conversely, semantic relevance never makes a private memory eligible for a wider audience. Subject and source identity come from the perspective-scoped conversation graph, not from claims inside message text. Visibility comes from retained source and audience provenance. These are structured references, not prose conventions.
Store references, render labels
Memory bodies remain natural prose. Actor labels such asVictor (id U123) are rendered when context is assembled; they are not embedded
as a storage format inside the body.
Persisting labels inline would duplicate mutable display names across many rows,
freeze one prompt presentation into canonical data, and leave nothing honest to
write before identity resolution completes. A structured nullable subject
reference represents “not resolved yet” directly and can be repaired without
parsing or rewriting prose.
This is the same separation used for replies: stored content is the canonical
fact, while presentation is derived from current identity data. A display-name
change therefore changes future rendering without rewriting memory history.
Hard eligibility precedes ranking
Retrieval begins inside one RuntimeAgent perspective. It must reject ineligible rows before computing relevance:- the memory belongs to the current RuntimeAgent
- the entry is live and its selected revision is current
- source and audience provenance permit reuse in the current conversation
- any explicit policy restriction is satisfied
Hybrid candidate retrieval
Eligible memory is selected through a union of complementary candidate sources:- subject candidates — memories about the current actor, positively boosted but not exclusive
- conversation candidates — memories tied to the current conversation or its durable context
- lexical candidates — PostgreSQL full-text search for exact phrases, names, identifiers, acronyms, and other token-sensitive matches
- semantic candidates — vector similarity for paraphrases and conceptual relevance
- always-on candidates — explicit high-salience instructions or preferences whose contract requires them on every applicable turn
Explainable reranking
A bounded candidate set is reranked from inspectable features such as:- semantic relevance to the current message and recent conversation
- current-subject match
- current-conversation match
- explicit salience
- confidence
- recency decay
One retrieval service
Automatic prompt assembly and an agent-facingrecall(query) tool use the same
retrieval service. They may request different budgets or query text, but they do
not define separate eligibility rules or ranking algorithms.
Tool results return stable memory entry and revision identifiers together with
subject, source, visibility-safe provenance, and scores. Those identifiers are
what later amend, supersede, or forget operations target.
Truthful confirmation
A reply may describe only durable operations that actually committed. This is a retrieval-adjacent contract rather than a prompt-style preference, because the failure it prevents is a lie about canonical state. Observed in production: after tworemember calls an agent told the user it had
“re-attributed” the earlier memory and that the earlier entry was “marked as
superseded”. The database held two independent live revision-1 entries; the first
was untouched. The agent also claimed a shared workspace had forced it to store
the memory anonymously, which this runtime never does.
Three mechanisms enforce it, in order of preference:
- Verbs that cannot misrepresent themselves.
remembercreates and only creates. Its argument schema has no entry id, no revision, and no supersede flag, so there is no request through which a model could ask it to modify an existing entry — and therefore no committed result that could honestly be described as a supersession. Correcting an entry isamend; retiring one isforget. - An authoritative committed ledger. Every durable write records itself from the store’s return value, never from the model’s request. Before the reply is generated, the complete list is injected as a developer-role instruction naming each operation, its stable entry id, and its resulting revision — plus explicit prohibitions on claiming supersession, re-attribution, deletion, or anonymous storage when no such operation committed. Prevention is preferred over detection: a model handed a complete, authoritative account of what happened has no remaining reason to narrate something else.
- A narrow claim guard. A final reply asserting a durable mutation that did not commit fails the turn. It matches first-person past-tense assertions and deliberately does not attempt to parse negation, hypotheticals, or reported speech — prose is not a security boundary. It exists so that ignoring (2) is loud rather than silent.
amend or forget, and a re-attribution claim specifically requires an amend
that actually changed the subject. A body-only correction commits an amend and
still leaves “I re-attributed it” false.
A tool call that changed nothing — an unknown or already-forgotten entry id —
returns a result saying so and contributes nothing to the ledger, so it cannot
become licence to claim the write happened.
Bounded tool rounds
Recall must be able to precede a correction: the model cannot amend an entry without first obtaining its id. That requires more than one tool round, which is not the same as an unbounded agent loop. The bound is three rounds and it is structural rather than advisory. On the last permitted round the tool list is removed from the request, so the model is not asked to stop — it is unable to continue. Fan-out inside a single round is bounded separately. Every call in a round is validated before any is executed, so a malformed second call cannot leave a committed first write behind from a round the runtime then rejects. A durable write the reply never learns about is exactly the state the ledger exists to make impossible.Embeddings are derived indexes
PostgreSQL remains the canonical memory store. A vector is a disposable retrieval index derived from immutable memory revision text, not a second source of truth. Embeddings attach tomemory_entry_revisions.id, not to the mutable live entry.
An amendment creates a new revision and a new embedding. Historical embeddings
remain audit history and are excluded from current retrieval by revision state.
Model and dimension metadata live with each embedding so changing embedding
models does not rewrite canonical memory.
The default implementation is PostgreSQL full-text retrieval. Optional pgvector
and Qdrant adapters add semantic candidates through the same scoped interface.
Embeddings are independently pluggable. See memory providers (repository)
for configuration, rebuilds, and the adapter contract. Learned rerankers, clustering,
and memory trees remain future work that requires measured quality improvements.
Evaluation contract
Retrieval changes are tested against cases that distinguish eligibility from relevance:- the current actor asks about their children; the subject boost retrieves the relevant fact
- another actor asks about tabs versus spaces; the CTO preference remains retrievable with its subject and source intact
- a coding question does not select an unrelated favorite-food memory
- a food question does select that favorite-food memory
- a private direct-message memory is never eligible in another actor’s turn, regardless of semantic similarity
- an amended memory exposes only its current revision to live retrieval
- an actor rename changes rendering without rewriting memory text
- two actors with the same display name remain distinct through stable ids
Current implementation
Migration0038_memory_audience_retrieval.sql adds source audience evidence,
revision-scoped sharing grants, an index outbox, embedding generations, and
retrieval audit records. The original entries, revisions, actor references, and
API shapes remain intact. Memory privacy (repository) defines the enforced
policy; memory providers (repository) describes setup and upgrades.
Automatic context and recall share memory.Store.search. Authorization in SQL
precedes candidate selection and all text is hydrated from live canonical rows.
PostgreSQL full-text terms retrieve lexical candidates across the archive, rather
than just the latest 100 memories. Nonempty queries require lexical or semantic
relevance; subject and conversation matches boost those candidates. Empty queries
rank by subject, conversation, and recency. Automatic broad questions such as
“what’s new?” use that empty-query mode; the recall tool also supports it.
Lexical and semantic ranks are fused with reciprocal-rank contributions. Subject
and conversation matches add explicit boosts. Each source contributes at most 200
candidates. Mesh caps valid, distinct semantic hits before canonical hydration,
even if a provider ignores the requested limit. Semantic search has a three-second
deadline and a 10,000-revision eligible-set cap; unavailable, oversized, or rebuilding
indexes degrade to authorized lexical results. These are conservative starting
weights, not a claim of corpus-tuned retrieval quality. Explicit salience,
confidence, learned reranking, and model-specific tokenization remain future work.
Complete entries are packed into a configurable memory allowance (default 4096),
counting the larger of prompt and recall JSON representations plus attribution
overhead as a conservative text-token upper bound. The fixed section/tool header
is outside this entry allowance. Oversized entries are skipped rather than truncated into changed claims.
This allowance bounds memory context, not the entire model window; history,
attachments, and tool results retain their separate existing limits.
memory_retrieval_events records a query digest, destination audience, selected
revision IDs, score components, conservative cost, and degradation category.
It does not store query text or provider error bodies. The digest is a keyed
HMAC under a subkey of MESH_AGENT_MASTER_KEY (internal/auditdigest), not a
plain hash: recall queries are a few words about a person or a project, and a
plain SHA-256 of them is reversed by hashing guesses. Without a master key no
digest is recorded. Rows written before migration 20260925202721 held the plain hash;
that migration cleared them. Prompt snapshots and tool
observations retain the actual context under the existing operator trust boundary.
When lexical matches are excluded with unknown source audiences, the audit records
eligibility_withheld_unknown even if other memories were returned. A broad term
can match several general facts while the specific answer is withheld. An earlier
semantic failure category retains precedence; this diagnostic never authorizes
disclosure or puts withheld text into the tool response.
Legacy audience repair runs through forward migrations. Migration 0089 repairs
current memory evidence skipped by 0085’s reliance on the optional, unpopulated
runtime_agents.actor_id. It requires an inbound source with an exact connector
identity and proven direct topology; other participants must be recorded outbound
voices, with no contradictory inbound senders. For old Slack threads without stored
topology, a later authenticated direct observation of the same connector/channel
and counterpart can establish the fixed DM. Channel ID prefixes and speaker counts
alone are insufficient. Missing proof leaves evidence unknown for operator review.
Repair preserves revision bodies, source references, conversation restrictions,
sharing grants, and historical transcript evidence.
Every candidate is revalidated after optional provider work. Before final reply
delivery, Mesh rechecks the audience and the used memory revisions. A change
withholds the response. A turn can acknowledge its own successful amendment or
forget operation; that operation does not remove earlier sources from the lineage
of subsequent memories.
New memories conservatively inherit all retrieved memories and loaded history
message sources from the turn. Amendments retain earlier evidence. The source is
connector-authenticated, never a model argument, and foreign or missing parent
references reject the entire write. This can withhold an otherwise shareable
summary when it has mixed sources; explicit review is preferable to silently
losing provenance.
remember, recall, amend, and forget retain their existing contracts and run
budgets. A fifth tool, request_memory_sharing, can propose release of one revision
in an authenticated private conversation. Only a matching authenticated source actor confirmation or
an explicit authenticated operator review creates a grant. A model classification
cannot create a grant.