Skip to main content

Conversation history search

An agent that persists every message but can only read the last 30 of them in one room does not have recall. It has a window. Mesh already stores the complete, append-only transcript of every conversation a RuntimeAgent has ever participated in (message_events, per data-model.md). The runtime can now search it. This document specifies the retrieval surface that closes that gap, and — more importantly — specifies it as a consumer of the existing disclosure model rather than as a second, weaker path to the same rows. Status: phases 1–3 are built — authorized lexical search, the search_history tool, and retrieval audit rows. Semantic candidates (phase 4) are not. See What is built for the map from this document to code.

What exists today

internal/graph/history.go is the only reader of prior messages on the turn path. Its scope is one conversation:
  • the key is (runtime_agent_id, connector type, connector name, external_thread_id) — the perspective is the first term of the key, not a filter applied afterward
  • the window is DefaultHistoryWindow = 30, capped at MaxHistoryWindow = 500, because history is a per-turn token and latency cost
  • dropping the oldest messages is the crude form of compaction
That describes the history window before search shipped. search_history now searches authorized messages outside that window. What does already exist is the hard part: authorization. ListRecentConversationTurns admits a row only if
  • structured source evidence exists for it at all (message_source_evidence, migration 0038_memory_audience_retrieval.sql) — unknown provenance fails closed, and
  • mesh_audience_allows(message audience, destination audience) holds, and
  • mesh_messages_allowed(...) holds — the recursive lineage walk from migration 0063_message_disclosure_lineage.sql, so an earlier agent reply cannot launder a private memory it was built from.
That predicate set is the disclosure rule. Search does not get its own.

The one-sentence contract

Conversation search is a new candidate source and a new ranking problem. It is not a new authorization boundary. Eligibility is decided in SQL by the same predicates that already gate history and memory, before anything is ranked.
This mirrors memory-retrieval.md’s “hard eligibility precedes ranking,” and for the same reason: a relevance score that can resurrect an ineligible row is not a relevance score, it is a bypass.

Eligibility

The policy: a conversation in a public channel is retrievable by anyone, from anywhere in the workspace. A private conversation — a DM, a private channel — is retrievable only in a conversation whose readers could already read it; in practice, a DM is retrievable only in that same person’s DM with the agent. The existing audience predicate implements most of that already:
  • a direct message is normalized to members with readers = [that actor's external id] (event.DefaultAudience), and mesh_audience_allows admits a members origin only into a members destination in the same connector_type:connector_name namespace whose reader set is a subset of the origin’s. Another actor’s DM therefore cannot reach this turn — the case that matters most.
  • public origins are admissible anywhere.
  • workspace origins reach workspace or member destinations in the same namespace, and never an externally shared conversation.
One question had to be settled deliberately rather than discovered later. mesh_audience_allows answers “may this text be disclosed to this destination audience?” It does not answer “was the requester in that room?” A workspace-visible channel the requester never joined is currently disclosable to them. That is defensible for a synthesized fact — workspace visibility is exactly the claim that it was not private — but search returns raw transcript lines from a named room at a named time, which is a meaningfully different artifact. Pulling verbatim text out of a room someone was never in is surprising even when it was technically workspace-visible. Decision: no participation filter. An earlier draft of this document recommended a search-only term requiring the requester to have spoken in a room before its transcript could surface. The product decision went the other way, deliberately: a public channel is fair game. If a conversation took place in a workspace-visible channel, any member of that workspace may retrieve it through the agent, from a DM or from a channel, whether or not they were ever in the room — exactly as they could by opening the channel and scrolling. The audience predicate is therefore the whole rule, unchanged and never loosened:
What that means concretely for a Slack install, asked from each kind of room: Verified identity linking adds one narrow permission path: single-reader private evidence can follow the same verified person into their other direct conversation. Both identities require live proof in this perspective. Original actor IDs remain intact; speaker=requester resolves their current association so it includes messages written before linking. Revocation separates the actors again and is rechecked before sending retrieved history. Groups, workspace scopes, external tool grants and explicit conversation restrictions retain their existing permissions. The requester still matters to ranking: a room they spoke in scores a small boost (requester_spoke), because “where were we talking about X” should prefer rooms they were in. A boost never makes an ineligible row eligible.

Retrieval

Candidates come from the same union shape memory retrieval uses, because the failure modes are identical:
  • lexical — PostgreSQL full-text search over message_events.body. This is the load-bearing one for search: people look for names, identifiers, error strings, and exact phrases, which is where FTS is strong and embeddings are weak.
  • semantic — optional, through the existing pluggable provider interface (internal/memory/providers, pgvector and Qdrant adapters). Reuse the adapter contract; do not grow a second embedding stack.
  • structural — filters, not scores: actor, conversation, connector, time range.
Ranking features stay inspectable: lexical rank, recency decay, same-actor match, same-organization or same-conversation match, and a demotion for the current conversation’s recent window, which the turn already loaded — a search result that repeats the message three lines above it in the prompt is wasted budget. Index mechanics — measured, as this section used to insist:
  • The first design was an expression GIN index on to_tsvector('english', message_events.body). The planner declined to use it: the perspective filter is selective, so it drove from the agent’s conversations and re-parsed every body to test the match (~100 ms per 10k messages), and ranking re-parsed every candidate again.
  • What shipped is the sidecar (migration 0090_conversation_search.sql): message_search_documents(message_event_id, runtime_agent_id, tsv), GIN on tsv, btree on runtime_agent_id, kept current by an AFTER INSERT OR UPDATE OF body trigger on message_events so no write path can forget to index. The planner combines the two indexes in one BitmapAnd: on 300k rows, under 1 ms for a rare term and ~9 ms for a term in a quarter of an agent’s messages. It never rewrites or read-locks message_events; ON DELETE CASCADE means a deleted message is no longer findable.
  • The sidecar’s runtime_agent_id is a derived copy used to narrow the scan. The query re-derives the perspective through connectors as well, so the sidecar narrows the search and never vouches for a row.
  • an index is a derived, disposable artifact. Canonical text is always hydrated from message_events after authorization, exactly as memory hydrates from live canonical rows — an index row is never the thing quoted.
  • Query words are OR’d (plainto_tsquery with its top-level & rewritten to |), and ts_rank_cd rewards matching more of them, closer together. Models search with the words they remember, and requiring every one of them is how “Acme migration decision” misses “we decided to move Acme on Friday”.
  • Embedding work, when phase 4 arrives, should reuse the memory indexer’s outbox pattern (internal/memory/indexer.go) so indexing is asynchronous, resumable, and per-perspective.
Degradation matches memory: an unavailable or rebuilding semantic index falls back to authorized lexical results and records the degradation category. It never falls back to unauthorized results.

The tool

One query-class tool per tools.md, running in the control plane with no shell and no arbitrary network reach:
Each hit carries up to two eligible messages on either side, so a result reads as an exchange rather than an answer with its question missing. Neighbors pass the same predicates as hits; an ineligible neighbor is skipped, not shown. message_ids precedes hits in the result so a ledger observation truncated mid-result still yields the full list to a resumed run. Notes that are contracts, not preferences:
  • no perspective or audience parameter. Both are bound from the run, like HistoryReader’s baked-in agentID. A tool that accepted them would be a tool that could be called without them.
  • no actor id the model invented. Actor filters resolve through the perspective-scoped graph; an unresolvable filter returns nothing rather than widening.
  • results return stable message_event_ids, because those are what provenance and later citation need.

Provenance is the part that is easy to get wrong

Retrieved transcript lines must flow into the turn’s evidence exactly as loaded history does, through memoryEvidence.addHistory (internal/turn/memory_context.go). Then:
  • memory_provenance.history_message_ids on the outbound reply includes every searched message the reply saw, so mesh_messages_allowed can walk it later;
  • a memory the model forms from a search result conservatively inherits every source’s audience, which may withhold an otherwise shareable summary — the documented and correct trade;
  • the pre-delivery recheck revalidates the audience and the used rows before the reply ships.
A search path that returned text outside this ledger would be a hole in the disclosure lineage, not a missing nicety. The provenance arrays have hard caps in mesh_messages_allowed (501 input messages, 5,000 lineage nodes, 500 history_message_ids), and overflow denies the whole read. A generous search limit can therefore make a later reply undeliverable. Bound search_history results well under those caps — start at 10, hard cap 50 — and treat the cap interaction as a test case, not a footnote. As built: per call the limit defaults to 8 and caps at 20 hits (each hit up to five messages with its neighbors), and per turn the tool stops adding once the turn’s evidence holds 400 distinct messages, shrinking its request as it approaches that ceiling and returning search_budget_exhausted at it. The message being answered (persisted before the turn, and the best match for its own question) and everything already in the loaded window are excluded.

Costs

  • prompt budget. Search results compete with history, memory, tools, and output reservation under prompt-budget.md (repository). Snippets, not whole messages; a bounded allowance like memory’s.
  • index size and write amplification on the largest table in the schema.
  • latency. A search round precedes the answer. The three-round tool bound already covers search_history → answer, and search_history → remember → answer.
  • retention and deletion. If a deployment ever honors “delete my messages,” the index and any embeddings are in scope. Cheap to build in now, expensive to retrofit.

Evaluation contract

Tests that distinguish eligibility from relevance, in the style of memory-retrieval’s list:
  • an actor asks about a decision in their own DM history; it is found
  • another actor asks the same question; the first actor’s DM is never a candidate, at any similarity score
  • a public-channel exchange both actors participated in is found by both
  • a workspace-visible channel the requester never joined is found (public is fair game — see the decision above)
  • a DM’s history never surfaces in a public channel, even for the DM’s own participant
  • a private channel’s history reaches the private channel and never a public one
  • a message with no source evidence (pre-0038 backfill gaps) is excluded
  • an agent reply built on a private memory is excluded from a wider audience via the lineage walk, even though its own text looks innocuous
  • search results are present in history_message_ids on the resulting reply
  • a limit near the provenance caps does not produce an undeliverable reply
  • a semantic-provider outage degrades to lexical results, never to unfiltered results
  • two actors with the same display name stay distinct through stable ids
Quality is measured on a corpus with known queries before any weight is called tuned.

Phasing

  1. Authorized lexical search, no tool. Built.
  2. search_history tool wired through the existing registry, evidence ledger, and provenance capture. Built.
  3. Retrieval audit rows in conversation_search_events (migration 0093): a keyed query digest (HMAC under a subkey of MESH_AGENT_MASTER_KEY, so short keyword queries cannot be recovered by hashing guesses; empty without a master key) — never query text — filters, destination audience, selected message ids with score components, and exclusion counts by the first rule that removed each match (already_in_context, no_evidence, audience, lineage, below_cutoff), computed with the search’s own predicates — including linked requester identities and the identity-aware audience check — and with lineage counted over the search’s own ranked pool, so the counts hold past the 1,000-match counting ceiling. Every search writes one, including searches that find nothing, and a failed write fails the search (as memory retrieval does). Each row links to its tool call and is shown on the run trace as search_audit. Built.
  4. Optional semantic candidates through the existing provider adapters, with degradation. Not built.
  5. Compaction and summarization of old windows, which this doc deliberately does not solve.
Steps 1–3 are the useful product. Step 4 is an improvement that must be earned with measurements.

What is built

The integration tests were checked by mutation: replacing mesh_audience_allows with true fails every privacy test with a named disclosure, and replacing mesh_messages_allowed with true fails exactly the two lineage tests (the laundered reply and the laundered neighbor), whose own audiences are public.