Persistence
Mesh persists one RuntimeAgent’s conversation graph as the canonical record of who that agent observed saying what, where, and to whom. There is no single deployment-global conversational worldview shared by every RuntimeAgent. Seeperspective.md for the normative isolation boundary and
data-model.md for the perspective-scoped entity model.
Perspective-scoped by construction
Every graph write belongs to exactly one RuntimeAgent perspective. The same external Slack event may therefore be represented once in Lyra’s graph and once in Daedalus’s graph. Those rows may carry different derived semantics: for example, a message may beengage for Lyra and observe for Daedalus.
That is not duplicate canonical state accidentally waiting to be normalized. It
is two agents observing the same external event through different perspectives.
The owning perspective is mechanically recoverable from the write path and from
every graph read: connectors, actors, and connector_identities carry a
NOT NULL runtime_agent_id, and connectors is unique on
(runtime_agent_id, type, name). See data-model.md
for which tables carry the column and why the other two do not.
The conversation namespace used to be the only scoping mechanism. It was not
sufficient: it resolves from per-install configuration, two installs could resolve
to the same value, and the shared connectors row that resulted merged both
agents’ conversations, identities, and actors.
Append-only by construction
message_events is never updated. Every inbound message and every outbound
agent reply is a new row inside one perspective. That is what makes attribution
auditable: we can always answer “why did this agent reply, and to whom” by
replaying that agent’s events.
The spine the events hang off — connectors, conversations, actors, connector
identities — is upserted, because those are identities rather than history.
Those identities are scoped to the owning RuntimeAgent perspective; co-resident
agents do not share them merely because the control plane can correlate their
external accounts.
Actor opacity
An Actor is a participant as known inside one perspective. Persistence must not classify the participant as human, agent, or service.Actor.kind is therefore not part of the model, and the column is gone. In
particular, an outbound message is attributable to this RuntimeAgent’s own
Actor/connector identity without requiring that Actor to carry kind = 'agent'.
Write path
internal/graph.Store is constructed with a RuntimeAgent id and cannot be used
without one (ErrNoPerspective). RecordMessage takes an
event.NormalizedMessage and performs, in order:
UpsertConnector(runtime_agent_id, type, name)— idempotent on that triple. Two agents whose installs resolve to the same namespace get two rows.UpsertConversation(runtime_agent_id, connector_id, external_thread_id, …)— idempotent on(connector_id, external_thread_id). The statement re-resolves the connector under the perspective key, so a foreign connector id writes nothing.- Resolve the connector identity by
(runtime_agent_id, connector_id, external_id)and, when promotion rules require it, create the Actor inside this perspective. A participant known by another RuntimeAgent does not satisfy this lookup. InsertMessageEvent(…)— inserts this perspective’s immutable observation. The conversation and the actor are both re-resolved under the perspective key inside the statement, which is what stands in for an ownership column on the largest table in the system.UpsertConversationParticipant(conversation_id, actor_id, created_at)— records the explicit actor-to-conversation edge, under the same check. The stored event timestamp is used, so replaying a connector delivery cannot move participation history forward.
graph.TxStore wraps the same logic in a single transaction, so a partially
written graph (a conversation without its message, an actor without its identity)
is never visible to readers.
Idempotency
Connectors redeliver. Idempotency is evaluated inside the perspective receiving the event.message_events has UNIQUE (conversation_id, external_message_id), and the insert uses:
SET is deliberately a no-op: it exists only so the already-stored row is
RETURNINGed instead of the statement erroring. A replay therefore returns the
original event id and cannot rewrite history.
An identical external message received by another RuntimeAgent belongs to a
different perspective/conversation identity and is not a duplicate of this row.
Outbound replies are events too
graph.OutboundMessage derives the outbound normalized event from the inbound
trigger: same perspective, same connector, same conversation, direction = 'outbound', and parent_external_message_id = the message that triggered the
turn.
The outbound author is this RuntimeAgent’s own connector identity represented as
an Actor in its graph. No Actor classification is required. Another RuntimeAgent
that later observes the Slack reply resolves that external identity as an
ordinary Actor in its perspective; it receives no hidden link back to the
sender’s RuntimeAgent row.
Ordering guarantees in the runtime
app.Runtime.HandleSlack persists the inbound event before calling the
model. If the turn fails we still have a durable, attributable record that the
message arrived, which also makes the turn retryable. Losing history is worse
than losing a reply.
The outbound event is persisted after delivery, because we only learn the
external message id from the send. If that write fails the error is surfaced but
the send result is still returned, so the caller does not retry a send that
already happened.
These guarantees apply within one RuntimeAgent perspective. A failure in one
perspective must not cause another RuntimeAgent to reuse or inherit its partially
written graph state.
Context reads
Graph history reads are perspective-scoped for the same reason writes are. Context assembly must never query deployment-global actor/conversation history and filter it down afterward. The owning RuntimeAgent is part of the read key. Seeruntime/context-resolution.md.
This matters especially for actor-centric history: Lyra’s Victor and Daedalus’s
Victor may be distinct Actor rows with different memory and conversation
histories even when the control plane can correlate their external identities.
A note on jsonb and the simple query protocol
Mesh runs pgx in simple query protocol mode so it can sit behind PgBouncer in transaction pooling mode. In that mode a Go[]byte parameter is encoded as a
bytea literal, which Postgres rejects for a jsonb column
(invalid input syntax for type json).
Queries on the hot path therefore avoid binding jsonb as a parameter and let
the column default to '{}'. Anything that needs to write real metadata must
cast explicitly ($n::jsonb) and pass string, not []byte.
Authoring a migration
Add a new file todb/migrations/ named <version>_<name>.sql. The version is a
UTC timestamp prefix, generated once at authoring time:
internal/migrate) applies any migration whose version the ledger
does not already record, in numeric version order, and skips the ones it does
— the apply-by-ledger-presence model Rails, Django, and goose all use. Numeric
order is the order a batch applies in, not a precondition on which files
may apply: a pending file that sorts below the highest applied version is
applied, not refused. Independent migrations authored on parallel branches both
apply in either merge order, with no renumber (MESH-366). The runner is still
forward-only in the sense that there are no down-migrations; it just no longer
mistakes “lower version” for “already superseded.”
This is a deliberate change from the original total-order guard, which refused any below-max pending file. That guard turned a benign parallel-authoring case — a new 4-digit file (0098) sorting below a timestamp-prefixed applied max — into a fleet-wide Fly boot outage (MESH-364): every redeploying instance hit the refusal,exit(1)’d, and crash-looped. Independence is now the default.
Why timestamps, not a sequential counter
Mesh used a zero-padded counter (0001, 0002, …) through 0095. A counter
forces a total order on work that is really a partial order: two branches
developed in parallel each grab the same next integer and collide the moment
they both merge, forcing a hand-renumber (see PR #415, renumbered 0091→0095).
A timestamp is stamped per branch at authoring time, so parallel work almost
never collides. Two files that do land on the same second are a hard error,
not a silent tiebreak: the prefix is the version, so both carry the same version
and the loader rejects them as a duplicate. scripts/new-migration.sh handles
this for you by advancing to the next free second, but if you hand-name a file
you must give it a prefix no other migration uses.
Legacy 00xx files and new timestamp files coexist in one directory: the runner
orders by numeric value, so 95 < 20260925170000 orders correctly even though
"0095" is lexically smaller only by length. Do not renumber the existing
00xx migrations — their strings are live ledger keys. (New 00xx files are
separately banned by CI; timestamps only — see MESH-365. Ordering below a
timestamp max is no longer fatal, but a fresh 4-digit prefix is still rejected
at the CI gate before it can land.)
The one thing the clock does not buy you
Timestamps give an order, not a dependency guarantee. If migration B truly requires the schema migration A builds, wall-clock skew between two machines is not a safe way to order them. Put dependent changes in the same file / same PR. Independent changes (a new column here, a new index there — the 99% case) need no ordering relationship at all, which is exactly why timestamps are safe for them.Running the persistence tests
The graph tests include a real Postgres round trip, skipped unless a database is configured:postgres:16 service and applies the migrations, so the SQL is
exercised on every run.
Perspective-isolation coverage exists and must stay:
internal/graphproves two Stores recording the same external event produce distinct connector, conversation, actor, and message-event rows.internal/admin/inspection_integration_test.goruns two RuntimeAgents through the real write path in the same namespace with a shared external user id, and proves no actor, conversation, identity, history row, or engagement verdict crosses — including that a write handed a foreign conversation or actor id inserts nothing.internal/migrateproves the ownership backfill derives correctly — including for a connector written under an install’s earlier resolved namespace — and refuses ambiguous input, leaving no DDL and no ledger row behind.internal/appproves each agent’s recorder, history reader, and engagement reader are wired to that agent’s own perspective.
Regenerating the query layer
internal/store is generated by sqlc from db/queries/*.sql
against db/migrations/*.sql. Never edit it by hand:
Actor.kind should land as a real
schema/query regeneration change rather than by hand-editing generated Go.