> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Memory retrieval

# Memory retrieval

Durable memory is not the same thing as prompt context.

Memory is the complete, append-only record of what a RuntimeAgent has retained.
Prompt context is a bounded, audience-appropriate selection from that record for
one turn. Treating those as the same set makes every remembered fact appear in
every conversation until an arbitrary recency limit evicts it. That is bounded
storage access, not recall.

This document owns the contract for selecting durable memory into model context.
[`context-resolution.md`](/runtime/context-resolution) owns the broader perspective and
audience boundary; [`../versioned-state.md`](/versioned-state) owns canonical
storage and revision history.

[`prompt-budget.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/prompt-budget.md) owns the final whole-request admission
boundary: retrieved evidence competes with history, tools, images and output
reservation, and optional context may shrink again during later tool rounds.

## Three independent axes

Every durable memory needs three concepts that must not be collapsed:

* **subject** — who or what the claim concerns
* **source** — who or what supplied the claim
* **visibility** — which audiences may receive the claim

A preference such as “the CTO prefers tabs” illustrates why all three matter.
The CTO may be its subject. The CTO, another actor, a repository policy, or a
source message may supply it. The fact may be reusable across a workplace, or it
may have been learned in a private conversation that does not permit wider use.

The rule is:

> **Subject changes relevance. Visibility changes eligibility. Source explains
> provenance.**

A subject match is a positive retrieval signal, never an authorization boundary.
If another engineer asks whether to use tabs or spaces, a memory about the CTO’s
preference may be relevant even though the current speaker is not its subject.
Conversely, semantic relevance never makes a private memory eligible for a wider
audience.

Subject and source identity come from the perspective-scoped conversation graph,
not from claims inside message text. Visibility comes from retained source and
audience provenance. These are structured references, not prose conventions.

## Store references, render labels

Memory bodies remain natural prose. Actor labels such as
`Victor (id U123)` are rendered when context is assembled; they are not embedded
as a storage format inside the body.

Persisting labels inline would duplicate mutable display names across many rows,
freeze one prompt presentation into canonical data, and leave nothing honest to
write before identity resolution completes. A structured nullable subject
reference represents “not resolved yet” directly and can be repaired without
parsing or rewriting prose.

This is the same separation used for replies: stored content is the canonical
fact, while presentation is derived from current identity data. A display-name
change therefore changes future rendering without rewriting memory history.

## Hard eligibility precedes ranking

Retrieval begins inside one RuntimeAgent perspective. It must reject ineligible
rows before computing relevance:

1. the memory belongs to the current RuntimeAgent
2. the entry is live and its selected revision is current
3. source and audience provenance permit reuse in the current conversation
4. any explicit policy restriction is satisfied

No ranking score may restore a row rejected here. In particular, a high semantic
similarity score does not override a private-conversation boundary.

## Hybrid candidate retrieval

Eligible memory is selected through a union of complementary candidate sources:

* **subject candidates** — memories about the current actor, positively boosted
  but not exclusive
* **conversation candidates** — memories tied to the current conversation or its
  durable context
* **lexical candidates** — PostgreSQL full-text search for exact phrases, names,
  identifiers, acronyms, and other token-sensitive matches
* **semantic candidates** — vector similarity for paraphrases and conceptual
  relevance
* **always-on candidates** — explicit high-salience instructions or preferences
  whose contract requires them on every applicable turn

The union is deduplicated by memory entry and revision before reranking. Lexical
and semantic retrieval are complementary: neither is a fallback for the other.
Exact identifiers and negation-sensitive wording are often lexical strengths;
paraphrases and topic similarity are semantic strengths.

## Explainable reranking

A bounded candidate set is reranked from inspectable features such as:

* semantic relevance to the current message and recent conversation
* current-subject match
* current-conversation match
* explicit salience
* confidence
* recency decay

Subject mismatch carries no automatic penalty. A subject match adds evidence;
its absence does not imply irrelevance.

Weights are configuration and evaluation results, not architectural truth. Every
selected memory should retain the component scores and the reason it was chosen,
so retrieval misses and surprising inclusions can be diagnosed from durable
state rather than guessed from a final prompt.

The resolver injects the highest-ranked eligible memories under both a count and
token budget. Rendered memory includes enough subject and provenance information
for the model to distinguish “Victor prefers tabs” from “the current speaker
previously told me this.”

## One retrieval service

Automatic prompt assembly and an agent-facing `recall(query)` tool use the same
retrieval service. They may request different budgets or query text, but they do
not define separate eligibility rules or ranking algorithms.

Tool results return stable memory entry and revision identifiers together with
subject, source, visibility-safe provenance, and scores. Those identifiers are
what later amend, supersede, or forget operations target.

## Truthful confirmation

A reply may describe only durable operations that actually committed.

This is a retrieval-adjacent contract rather than a prompt-style preference,
because the failure it prevents is a lie about canonical state. Observed in
production: after two `remember` calls an agent told the user it had
"re-attributed" the earlier memory and that the earlier entry was "marked as
superseded". The database held two independent live revision-1 entries; the first
was untouched. The agent also claimed a shared workspace had forced it to store
the memory anonymously, which this runtime never does.

Three mechanisms enforce it, in order of preference:

1. **Verbs that cannot misrepresent themselves.** `remember` creates and only
   creates. Its argument schema has no entry id, no revision, and no supersede
   flag, so there is no request through which a model could ask it to modify an
   existing entry — and therefore no committed result that could honestly be
   described as a supersession. Correcting an entry is `amend`; retiring one is
   `forget`.
2. **An authoritative committed ledger.** Every durable write records itself from
   the *store's* return value, never from the model's request. Before the reply is
   generated, the complete list is injected as a developer-role instruction naming
   each operation, its stable entry id, and its resulting revision — plus explicit
   prohibitions on claiming supersession, re-attribution, deletion, or anonymous
   storage when no such operation committed. Prevention is preferred over
   detection: a model handed a complete, authoritative account of what happened
   has no remaining reason to narrate something else.
3. **A narrow claim guard.** A final reply asserting a durable mutation that did
   not commit fails the turn. It matches first-person past-tense assertions and
   deliberately does not attempt to parse negation, hypotheticals, or reported
   speech — prose is not a security boundary. It exists so that ignoring (2) is
   loud rather than silent.

Re-attribution and supersession claims become permissible only after a successful
`amend` or `forget`, and a re-attribution claim specifically requires an `amend`
that actually changed the subject. A body-only correction commits an amend and
still leaves "I re-attributed it" false.

A tool call that changed nothing — an unknown or already-forgotten entry id —
returns a result saying so and contributes nothing to the ledger, so it cannot
become licence to claim the write happened.

## Bounded tool rounds

Recall must be able to precede a correction: the model cannot amend an entry
without first obtaining its id. That requires more than one tool round, which is
not the same as an unbounded agent loop.

The bound is three rounds and it is structural rather than advisory. On the last
permitted round the tool list is removed from the request, so the model is not
asked to stop — it is unable to continue. Fan-out inside a single round is bounded
separately.

Every call in a round is validated before any is executed, so a malformed second
call cannot leave a committed first write behind from a round the runtime then
rejects. A durable write the reply never learns about is exactly the state the
ledger exists to make impossible.

## Embeddings are derived indexes

PostgreSQL remains the canonical memory store. A vector is a disposable retrieval
index derived from immutable memory revision text, not a second source of truth.

Embeddings attach to `memory_entry_revisions.id`, not to the mutable live entry.
An amendment creates a new revision and a new embedding. Historical embeddings
remain audit history and are excluded from current retrieval by revision state.
Model and dimension metadata live with each embedding so changing embedding
models does not rewrite canonical memory.

The default implementation is PostgreSQL full-text retrieval. Optional pgvector
and Qdrant adapters add semantic candidates through the same scoped interface.
Embeddings are independently pluggable. See [memory providers (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/memory-backends.md)
for configuration, rebuilds, and the adapter contract. Learned rerankers, clustering,
and memory trees remain future work that requires measured quality improvements.

## Evaluation contract

Retrieval changes are tested against cases that distinguish eligibility from
relevance:

* the current actor asks about their children; the subject boost retrieves the
  relevant fact
* another actor asks about tabs versus spaces; the CTO preference remains
  retrievable with its subject and source intact
* a coding question does not select an unrelated favorite-food memory
* a food question does select that favorite-food memory
* a private direct-message memory is never eligible in another actor’s turn,
  regardless of semantic similarity
* an amended memory exposes only its current revision to live retrieval
* an actor rename changes rendering without rewriting memory text
* two actors with the same display name remain distinct through stable ids

Quality should be measured on a corpus with known queries before semantic weights
are treated as reliable. Hybrid retrieval earns its complexity by improving
measured recall and precision, not by making the schema look sophisticated.

## Current implementation

Migration `0038_memory_audience_retrieval.sql` adds source audience evidence,
revision-scoped sharing grants, an index outbox, embedding generations, and
retrieval audit records. The original entries, revisions, actor references, and
API shapes remain intact. [Memory privacy (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/memory-privacy.md) defines the enforced
policy; [memory providers (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/memory-backends.md) describes setup and upgrades.

Automatic context and `recall` share `memory.Store.search`. Authorization in SQL
precedes candidate selection and all text is hydrated from live canonical rows.
PostgreSQL full-text terms retrieve lexical candidates across the archive, rather
than just the latest 100 memories. Nonempty queries require lexical or semantic
relevance; subject and conversation matches boost those candidates. Empty queries
rank by subject, conversation, and recency. Automatic broad questions such as
"what's new?" use that empty-query mode; the `recall` tool also supports it.

Lexical and semantic ranks are fused with reciprocal-rank contributions. Subject
and conversation matches add explicit boosts. Each source contributes at most 200
candidates. Mesh caps valid, distinct semantic hits before canonical hydration,
even if a provider ignores the requested limit. Semantic search has a three-second
deadline and a 10,000-revision eligible-set cap; unavailable, oversized, or rebuilding
indexes degrade to authorized lexical results. These are conservative starting
weights, not a claim of corpus-tuned retrieval quality. Explicit salience,
confidence, learned reranking, and model-specific tokenization remain future work.

Complete entries are packed into a configurable memory allowance (default 4096),
counting the larger of prompt and recall JSON representations plus attribution
overhead as a conservative text-token upper bound. The fixed section/tool header
is outside this entry allowance. Oversized entries are skipped rather than truncated into changed claims.
This allowance bounds memory context, not the entire model window; history,
attachments, and tool results retain their separate existing limits.

`memory_retrieval_events` records a query digest, destination audience, selected
revision IDs, score components, conservative cost, and degradation category.
It does not store query text or provider error bodies. The digest is a keyed
HMAC under a subkey of `MESH_AGENT_MASTER_KEY` (`internal/auditdigest`), not a
plain hash: recall queries are a few words about a person or a project, and a
plain SHA-256 of them is reversed by hashing guesses. Without a master key no
digest is recorded. Rows written before migration `20260925202721` held the plain hash;
that migration cleared them. Prompt snapshots and tool
observations retain the actual context under the existing operator trust boundary.
When lexical matches are excluded with unknown source audiences, the audit records
`eligibility_withheld_unknown` even if other memories were returned. A broad term
can match several general facts while the specific answer is withheld. An earlier
semantic failure category retains precedence; this diagnostic never authorizes
disclosure or puts withheld text into the tool response.

Legacy audience repair runs through forward migrations. Migration 0089 repairs
current memory evidence skipped by 0085's reliance on the optional, unpopulated
`runtime_agents.actor_id`. It requires an inbound source with an exact connector
identity and proven direct topology; other participants must be recorded outbound
voices, with no contradictory inbound senders. For old Slack threads without stored
topology, a later authenticated direct observation of the same connector/channel
and counterpart can establish the fixed DM. Channel ID prefixes and speaker counts
alone are insufficient. Missing proof leaves evidence unknown for operator review.
Repair preserves revision bodies, source references, conversation restrictions,
sharing grants, and historical transcript evidence.

Every candidate is revalidated after optional provider work. Before final reply
delivery, Mesh rechecks the audience and the used memory revisions. A change
withholds the response. A turn can acknowledge its own successful amendment or
forget operation; that operation does not remove earlier sources from the lineage
of subsequent memories.

New memories conservatively inherit all retrieved memories and loaded history
message sources from the turn. Amendments retain earlier evidence. The source is
connector-authenticated, never a model argument, and foreign or missing parent
references reject the entire write. This can withhold an otherwise shareable
summary when it has mixed sources; explicit review is preferable to silently
losing provenance.

`remember`, `recall`, `amend`, and `forget` retain their existing contracts and run
budgets. A fifth tool, `request_memory_sharing`, can propose release of one revision
in an authenticated private conversation. Only a matching authenticated source actor confirmation or
an explicit authenticated operator review creates a grant. A model classification
cannot create a grant.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.