Conversation history search
An agent that persists every message but can only read the last 30 of them in one room does not have recall. It has a window. Mesh already stores the complete, append-only transcript of every conversation a RuntimeAgent has ever participated in (message_events, per
data-model.md). The runtime can now search it. This
document specifies the retrieval surface that closes that gap, and — more
importantly — specifies it as a consumer of the existing disclosure model
rather than as a second, weaker path to the same rows.
Status: phases 1–3 are built — authorized lexical search, the search_history
tool, and retrieval audit rows. Semantic candidates (phase 4) are not. See What is built for the map from this
document to code.
What exists today
internal/graph/history.go is the only reader of prior messages on the turn
path. Its scope is one conversation:
- the key is
(runtime_agent_id, connector type, connector name, external_thread_id)— the perspective is the first term of the key, not a filter applied afterward - the window is
DefaultHistoryWindow = 30, capped atMaxHistoryWindow = 500, because history is a per-turn token and latency cost - dropping the oldest messages is the crude form of compaction
search_history now
searches authorized messages outside that window.
What does already exist is the hard part: authorization.
ListRecentConversationTurns admits a row only if
- structured source evidence exists for it at all (
message_source_evidence, migration0038_memory_audience_retrieval.sql) — unknown provenance fails closed, and mesh_audience_allows(message audience, destination audience)holds, andmesh_messages_allowed(...)holds — the recursive lineage walk from migration0063_message_disclosure_lineage.sql, so an earlier agent reply cannot launder a private memory it was built from.
The one-sentence contract
Conversation search is a new candidate source and a new ranking problem. It is not a new authorization boundary. Eligibility is decided in SQL by the same predicates that already gate history and memory, before anything is ranked.This mirrors
memory-retrieval.md’s “hard eligibility
precedes ranking,” and for the same reason: a relevance score that can resurrect
an ineligible row is not a relevance score, it is a bypass.
Eligibility
The policy: a conversation in a public channel is retrievable by anyone, from anywhere in the workspace. A private conversation — a DM, a private channel — is retrievable only in a conversation whose readers could already read it; in practice, a DM is retrievable only in that same person’s DM with the agent. The existing audience predicate implements most of that already:- a direct message is normalized to
memberswithreaders = [that actor's external id](event.DefaultAudience), andmesh_audience_allowsadmits amembersorigin only into amembersdestination in the sameconnector_type:connector_namenamespace whose reader set is a subset of the origin’s. Another actor’s DM therefore cannot reach this turn — the case that matters most. publicorigins are admissible anywhere.workspaceorigins reach workspace or member destinations in the same namespace, and never an externally shared conversation.
mesh_audience_allows answers “may this text be disclosed to
this destination audience?” It does not answer “was the requester in that
room?” A workspace-visible channel the requester never joined is currently
disclosable to them. That is defensible for a synthesized fact — workspace
visibility is exactly the claim that it was not private — but search returns
raw transcript lines from a named room at a named time, which is a
meaningfully different artifact. Pulling verbatim text out of a room someone was
never in is surprising even when it was technically workspace-visible.
Decision: no participation filter. An earlier draft of this document
recommended a search-only term requiring the requester to have spoken in a room
before its transcript could surface. The product decision went the other way,
deliberately: a public channel is fair game. If a conversation took place in a
workspace-visible channel, any member of that workspace may retrieve it through
the agent, from a DM or from a channel, whether or not they were ever in the
room — exactly as they could by opening the channel and scrolling. The audience
predicate is therefore the whole rule, unchanged and never loosened:
Verified identity linking adds one narrow permission path:
single-reader private evidence can follow the same verified person into their
other direct conversation. Both identities require live proof in this perspective.
Original actor IDs remain intact;
speaker=requester resolves their current
association so it includes messages written before linking. Revocation separates
the actors again and is rechecked before sending retrieved history. Groups,
workspace scopes, external tool grants and explicit conversation restrictions
retain their existing permissions.
The requester still matters to ranking: a room they spoke in scores a small
boost (requester_spoke), because “where were we talking about X” should
prefer rooms they were in. A boost never makes an ineligible row eligible.
Retrieval
Candidates come from the same union shape memory retrieval uses, because the failure modes are identical:- lexical — PostgreSQL full-text search over
message_events.body. This is the load-bearing one for search: people look for names, identifiers, error strings, and exact phrases, which is where FTS is strong and embeddings are weak. - semantic — optional, through the existing pluggable provider interface
(
internal/memory/providers, pgvector and Qdrant adapters). Reuse the adapter contract; do not grow a second embedding stack. - structural — filters, not scores: actor, conversation, connector, time range.
- The first design was an expression GIN index on
to_tsvector('english', message_events.body). The planner declined to use it: the perspective filter is selective, so it drove from the agent’s conversations and re-parsed every body to test the match (~100 ms per 10k messages), and ranking re-parsed every candidate again. - What shipped is the sidecar (migration
0090_conversation_search.sql):message_search_documents(message_event_id, runtime_agent_id, tsv), GIN ontsv, btree onruntime_agent_id, kept current by anAFTER INSERT OR UPDATE OF bodytrigger onmessage_eventsso no write path can forget to index. The planner combines the two indexes in one BitmapAnd: on 300k rows, under 1 ms for a rare term and ~9 ms for a term in a quarter of an agent’s messages. It never rewrites or read-locksmessage_events;ON DELETE CASCADEmeans a deleted message is no longer findable. - The sidecar’s
runtime_agent_idis a derived copy used to narrow the scan. The query re-derives the perspective throughconnectorsas well, so the sidecar narrows the search and never vouches for a row. - an index is a derived, disposable artifact. Canonical text is always hydrated
from
message_eventsafter authorization, exactly as memory hydrates from live canonical rows — an index row is never the thing quoted. - Query words are OR’d (
plainto_tsquerywith its top-level&rewritten to|), andts_rank_cdrewards matching more of them, closer together. Models search with the words they remember, and requiring every one of them is how “Acme migration decision” misses “we decided to move Acme on Friday”. - Embedding work, when phase 4 arrives, should reuse the memory indexer’s
outbox pattern (
internal/memory/indexer.go) so indexing is asynchronous, resumable, and per-perspective.
The tool
Onequery-class tool per tools.md, running in the control plane
with no shell and no arbitrary network reach:
message_ids precedes hits in the result so a ledger observation truncated
mid-result still yields the full list to a resumed run.
Notes that are contracts, not preferences:
- no perspective or audience parameter. Both are bound from the run, like
HistoryReader’s baked-inagentID. A tool that accepted them would be a tool that could be called without them. - no actor id the model invented. Actor filters resolve through the perspective-scoped graph; an unresolvable filter returns nothing rather than widening.
- results return stable
message_event_ids, because those are what provenance and later citation need.
Provenance is the part that is easy to get wrong
Retrieved transcript lines must flow into the turn’s evidence exactly as loaded history does, throughmemoryEvidence.addHistory
(internal/turn/memory_context.go). Then:
memory_provenance.history_message_idson the outbound reply includes every searched message the reply saw, somesh_messages_allowedcan walk it later;- a memory the model forms from a search result conservatively inherits every source’s audience, which may withhold an otherwise shareable summary — the documented and correct trade;
- the pre-delivery recheck revalidates the audience and the used rows before the reply ships.
mesh_messages_allowed (501 input messages, 5,000 lineage nodes, 500
history_message_ids), and overflow denies the whole read. A generous
search limit can therefore make a later reply undeliverable. Bound
search_history results well under those caps — start at 10, hard cap 50 — and
treat the cap interaction as a test case, not a footnote.
As built: per call the limit defaults to 8 and caps at 20 hits (each hit up to
five messages with its neighbors), and per turn the tool stops adding once the
turn’s evidence holds 400 distinct messages, shrinking its request as it
approaches that ceiling and returning search_budget_exhausted at it. The
message being answered (persisted before the turn, and the best match for its
own question) and everything already in the loaded window are excluded.
Costs
- prompt budget. Search results compete with history, memory, tools, and
output reservation under
prompt-budget.md(repository). Snippets, not whole messages; a bounded allowance like memory’s. - index size and write amplification on the largest table in the schema.
- latency. A search round precedes the answer. The three-round tool bound
already covers
search_history→ answer, andsearch_history→remember→ answer. - retention and deletion. If a deployment ever honors “delete my messages,” the index and any embeddings are in scope. Cheap to build in now, expensive to retrofit.
Evaluation contract
Tests that distinguish eligibility from relevance, in the style of memory-retrieval’s list:- an actor asks about a decision in their own DM history; it is found
- another actor asks the same question; the first actor’s DM is never a candidate, at any similarity score
- a public-channel exchange both actors participated in is found by both
- a workspace-visible channel the requester never joined is found (public is fair game — see the decision above)
- a DM’s history never surfaces in a public channel, even for the DM’s own participant
- a private channel’s history reaches the private channel and never a public one
- a message with no source evidence (pre-
0038backfill gaps) is excluded - an agent reply built on a private memory is excluded from a wider audience via the lineage walk, even though its own text looks innocuous
- search results are present in
history_message_idson the resulting reply - a
limitnear the provenance caps does not produce an undeliverable reply - a semantic-provider outage degrades to lexical results, never to unfiltered results
- two actors with the same display name stay distinct through stable ids
Phasing
- Authorized lexical search, no tool. Built.
search_historytool wired through the existing registry, evidence ledger, and provenance capture. Built.- Retrieval audit rows in
conversation_search_events(migration0093): a keyed query digest (HMAC under a subkey ofMESH_AGENT_MASTER_KEY, so short keyword queries cannot be recovered by hashing guesses; empty without a master key) — never query text — filters, destination audience, selected message ids with score components, and exclusion counts by the first rule that removed each match (already_in_context,no_evidence,audience,lineage,below_cutoff), computed with the search’s own predicates — including linked requester identities and the identity-aware audience check — and with lineage counted over the search’s own ranked pool, so the counts hold past the 1,000-match counting ceiling. Every search writes one, including searches that find nothing, and a failed write fails the search (as memory retrieval does). Each row links to its tool call and is shown on the run trace assearch_audit. Built. - Optional semantic candidates through the existing provider adapters, with degradation. Not built.
- Compaction and summarization of old windows, which this doc deliberately does not solve.
What is built
The integration tests were checked by mutation: replacing
mesh_audience_allows
with true fails every privacy test with a named disclosure, and replacing
mesh_messages_allowed with true fails exactly the two lineage tests (the
laundered reply and the laundered neighbor), whose own audiences are public.