Skip to main content

Persistence

Mesh persists one RuntimeAgent’s conversation graph as the canonical record of who that agent observed saying what, where, and to whom. There is no single deployment-global conversational worldview shared by every RuntimeAgent. See perspective.md for the normative isolation boundary and data-model.md for the perspective-scoped entity model.

Perspective-scoped by construction

Every graph write belongs to exactly one RuntimeAgent perspective. The same external Slack event may therefore be represented once in Lyra’s graph and once in Daedalus’s graph. Those rows may carry different derived semantics: for example, a message may be engage for Lyra and observe for Daedalus. That is not duplicate canonical state accidentally waiting to be normalized. It is two agents observing the same external event through different perspectives. The owning perspective is mechanically recoverable from the write path and from every graph read: connectors, actors, and connector_identities carry a NOT NULL runtime_agent_id, and connectors is unique on (runtime_agent_id, type, name). See data-model.md for which tables carry the column and why the other two do not. The conversation namespace used to be the only scoping mechanism. It was not sufficient: it resolves from per-install configuration, two installs could resolve to the same value, and the shared connectors row that resulted merged both agents’ conversations, identities, and actors.

Append-only by construction

message_events is never updated. Every inbound message and every outbound agent reply is a new row inside one perspective. That is what makes attribution auditable: we can always answer “why did this agent reply, and to whom” by replaying that agent’s events. The spine the events hang off — connectors, conversations, actors, connector identities — is upserted, because those are identities rather than history. Those identities are scoped to the owning RuntimeAgent perspective; co-resident agents do not share them merely because the control plane can correlate their external accounts.

Actor opacity

An Actor is a participant as known inside one perspective. Persistence must not classify the participant as human, agent, or service. Actor.kind is therefore not part of the model, and the column is gone. In particular, an outbound message is attributable to this RuntimeAgent’s own Actor/connector identity without requiring that Actor to carry kind = 'agent'.

Write path

internal/graph.Store is constructed with a RuntimeAgent id and cannot be used without one (ErrNoPerspective). RecordMessage takes an event.NormalizedMessage and performs, in order:
  1. UpsertConnector(runtime_agent_id, type, name) — idempotent on that triple. Two agents whose installs resolve to the same namespace get two rows.
  2. UpsertConversation(runtime_agent_id, connector_id, external_thread_id, …) — idempotent on (connector_id, external_thread_id). The statement re-resolves the connector under the perspective key, so a foreign connector id writes nothing.
  3. Resolve the connector identity by (runtime_agent_id, connector_id, external_id) and, when promotion rules require it, create the Actor inside this perspective. A participant known by another RuntimeAgent does not satisfy this lookup.
  4. InsertMessageEvent(…) — inserts this perspective’s immutable observation. The conversation and the actor are both re-resolved under the perspective key inside the statement, which is what stands in for an ownership column on the largest table in the system.
  5. UpsertConversationParticipant(conversation_id, actor_id, created_at) — records the explicit actor-to-conversation edge, under the same check. The stored event timestamp is used, so replaying a connector delivery cannot move participation history forward.
graph.TxStore wraps the same logic in a single transaction, so a partially written graph (a conversation without its message, an actor without its identity) is never visible to readers.

Idempotency

Connectors redeliver. Idempotency is evaluated inside the perspective receiving the event. message_events has UNIQUE (conversation_id, external_message_id), and the insert uses:
The SET is deliberately a no-op: it exists only so the already-stored row is RETURNINGed instead of the statement erroring. A replay therefore returns the original event id and cannot rewrite history. An identical external message received by another RuntimeAgent belongs to a different perspective/conversation identity and is not a duplicate of this row.

Outbound replies are events too

graph.OutboundMessage derives the outbound normalized event from the inbound trigger: same perspective, same connector, same conversation, direction = 'outbound', and parent_external_message_id = the message that triggered the turn. The outbound author is this RuntimeAgent’s own connector identity represented as an Actor in its graph. No Actor classification is required. Another RuntimeAgent that later observes the Slack reply resolves that external identity as an ordinary Actor in its perspective; it receives no hidden link back to the sender’s RuntimeAgent row.

Ordering guarantees in the runtime

app.Runtime.HandleSlack persists the inbound event before calling the model. If the turn fails we still have a durable, attributable record that the message arrived, which also makes the turn retryable. Losing history is worse than losing a reply. The outbound event is persisted after delivery, because we only learn the external message id from the send. If that write fails the error is surfaced but the send result is still returned, so the caller does not retry a send that already happened. These guarantees apply within one RuntimeAgent perspective. A failure in one perspective must not cause another RuntimeAgent to reuse or inherit its partially written graph state.

Context reads

Graph history reads are perspective-scoped for the same reason writes are. Context assembly must never query deployment-global actor/conversation history and filter it down afterward. The owning RuntimeAgent is part of the read key. See runtime/context-resolution.md. This matters especially for actor-centric history: Lyra’s Victor and Daedalus’s Victor may be distinct Actor rows with different memory and conversation histories even when the control plane can correlate their external identities.

A note on jsonb and the simple query protocol

Mesh runs pgx in simple query protocol mode so it can sit behind PgBouncer in transaction pooling mode. In that mode a Go []byte parameter is encoded as a bytea literal, which Postgres rejects for a jsonb column (invalid input syntax for type json). Queries on the hot path therefore avoid binding jsonb as a parameter and let the column default to '{}'. Anything that needs to write real metadata must cast explicitly ($n::jsonb) and pass string, not []byte.

Authoring a migration

Add a new file to db/migrations/ named <version>_<name>.sql. The version is a UTC timestamp prefix, generated once at authoring time:
The runner (internal/migrate) applies any migration whose version the ledger does not already record, in numeric version order, and skips the ones it does — the apply-by-ledger-presence model Rails, Django, and goose all use. Numeric order is the order a batch applies in, not a precondition on which files may apply: a pending file that sorts below the highest applied version is applied, not refused. Independent migrations authored on parallel branches both apply in either merge order, with no renumber (MESH-366). The runner is still forward-only in the sense that there are no down-migrations; it just no longer mistakes “lower version” for “already superseded.”
This is a deliberate change from the original total-order guard, which refused any below-max pending file. That guard turned a benign parallel-authoring case — a new 4-digit file (0098) sorting below a timestamp-prefixed applied max — into a fleet-wide Fly boot outage (MESH-364): every redeploying instance hit the refusal, exit(1)’d, and crash-looped. Independence is now the default.

Why timestamps, not a sequential counter

Mesh used a zero-padded counter (0001, 0002, …) through 0095. A counter forces a total order on work that is really a partial order: two branches developed in parallel each grab the same next integer and collide the moment they both merge, forcing a hand-renumber (see PR #415, renumbered 0091→0095). A timestamp is stamped per branch at authoring time, so parallel work almost never collides. Two files that do land on the same second are a hard error, not a silent tiebreak: the prefix is the version, so both carry the same version and the loader rejects them as a duplicate. scripts/new-migration.sh handles this for you by advancing to the next free second, but if you hand-name a file you must give it a prefix no other migration uses. Legacy 00xx files and new timestamp files coexist in one directory: the runner orders by numeric value, so 95 < 20260925170000 orders correctly even though "0095" is lexically smaller only by length. Do not renumber the existing 00xx migrations — their strings are live ledger keys. (New 00xx files are separately banned by CI; timestamps only — see MESH-365. Ordering below a timestamp max is no longer fatal, but a fresh 4-digit prefix is still rejected at the CI gate before it can land.)

The one thing the clock does not buy you

Timestamps give an order, not a dependency guarantee. If migration B truly requires the schema migration A builds, wall-clock skew between two machines is not a safe way to order them. Put dependent changes in the same file / same PR. Independent changes (a new column here, a new index there — the 99% case) need no ordering relationship at all, which is exactly why timestamps are safe for them.

Running the persistence tests

The graph tests include a real Postgres round trip, skipped unless a database is configured:
CI provisions a postgres:16 service and applies the migrations, so the SQL is exercised on every run. Perspective-isolation coverage exists and must stay:
  • internal/graph proves two Stores recording the same external event produce distinct connector, conversation, actor, and message-event rows.
  • internal/admin/inspection_integration_test.go runs two RuntimeAgents through the real write path in the same namespace with a shared external user id, and proves no actor, conversation, identity, history row, or engagement verdict crosses — including that a write handed a foreign conversation or actor id inserts nothing.
  • internal/migrate proves the ownership backfill derives correctly — including for a connector written under an install’s earlier resolved namespace — and refuses ambiguous input, leaving no DDL and no ledger row behind.
  • internal/app proves each agent’s recorder, history reader, and engagement reader are wired to that agent’s own perspective.
A fixture that gives the two agents different namespaces is not testing the boundary — that shape was already isolated by accident before the ownership key existed.

Regenerating the query layer

internal/store is generated by sqlc from db/queries/*.sql against db/migrations/*.sql. Never edit it by hand:
This is why removal of legacy fields such as Actor.kind should land as a real schema/query regeneration change rather than by hand-editing generated Go.