> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Persistence

# Persistence

Mesh persists one RuntimeAgent's conversation graph as the canonical record of
*who that agent observed saying what, where, and to whom*. There is no single
deployment-global conversational worldview shared by every RuntimeAgent.

See [`perspective.md`](/perspective) for the normative isolation boundary and
[`data-model.md`](/data-model) for the perspective-scoped entity model.

## Perspective-scoped by construction

Every graph write belongs to exactly one RuntimeAgent perspective.

The same external Slack event may therefore be represented once in Lyra's graph
and once in Daedalus's graph. Those rows may carry different derived semantics:
for example, a message may be `engage` for Lyra and `observe` for Daedalus.
That is not duplicate canonical state accidentally waiting to be normalized. It
is two agents observing the same external event through different perspectives.

The owning perspective is mechanically recoverable from the write path and from
every graph read: `connectors`, `actors`, and `connector_identities` carry a
`NOT NULL runtime_agent_id`, and `connectors` is unique on
`(runtime_agent_id, type, name)`. See [`data-model.md`](/data-model#how-the-physical-schema-expresses-it)
for which tables carry the column and why the other two do not.

The conversation namespace used to be the only scoping mechanism. It was not
sufficient: it resolves from per-install configuration, two installs could resolve
to the same value, and the shared `connectors` row that resulted merged both
agents' conversations, identities, and actors.

## Append-only by construction

`message_events` is never updated. Every inbound message and every outbound
agent reply is a new row inside one perspective. That is what makes attribution
auditable: we can always answer "why did this agent reply, and to whom" by
replaying that agent's events.

The spine the events hang off — connectors, conversations, actors, connector
identities — *is* upserted, because those are identities rather than history.
Those identities are scoped to the owning RuntimeAgent perspective; co-resident
agents do not share them merely because the control plane can correlate their
external accounts.

## Actor opacity

An Actor is a participant as known inside one perspective. Persistence must not
classify the participant as human, agent, or service.

`Actor.kind` is therefore not part of the model, and the column is gone. In
particular, an outbound message is attributable to this RuntimeAgent's own
Actor/connector identity without requiring that Actor to carry `kind = 'agent'`.

## Write path

`internal/graph.Store` is constructed with a RuntimeAgent id and cannot be used
without one (`ErrNoPerspective`). `RecordMessage` takes an
`event.NormalizedMessage` and performs, in order:

1. `UpsertConnector(runtime_agent_id, type, name)` — idempotent on that triple.
   Two agents whose installs resolve to the same namespace get two rows.
2. `UpsertConversation(runtime_agent_id, connector_id, external_thread_id, …)` —
   idempotent on `(connector_id, external_thread_id)`. The statement re-resolves
   the connector under the perspective key, so a foreign connector id writes
   nothing.
3. Resolve the connector identity by
   `(runtime_agent_id, connector_id, external_id)` and, when promotion rules
   require it, create the Actor inside this perspective. A participant known by
   another RuntimeAgent does not satisfy this lookup.
4. `InsertMessageEvent(…)` — inserts this perspective's immutable observation.
   The conversation and the actor are both re-resolved under the perspective key
   inside the statement, which is what stands in for an ownership column on the
   largest table in the system.
5. `UpsertConversationParticipant(conversation_id, actor_id, created_at)` —
   records the explicit actor-to-conversation edge, under the same check. The
   stored event timestamp is used, so replaying a connector delivery cannot move
   participation history forward.

`graph.TxStore` wraps the same logic in a single transaction, so a partially
written graph (a conversation without its message, an actor without its identity)
is never visible to readers.

## Idempotency

Connectors redeliver. Idempotency is evaluated inside the perspective receiving
the event. `message_events` has `UNIQUE (conversation_id,
external_message_id)`, and the insert uses:

```sql theme={null}
ON CONFLICT (conversation_id, external_message_id) DO UPDATE
    SET body = message_events.body
```

The `SET` is deliberately a no-op: it exists only so the already-stored row is
`RETURNING`ed instead of the statement erroring. A replay therefore returns the
original event id and **cannot rewrite history**.

An identical external message received by another RuntimeAgent belongs to a
different perspective/conversation identity and is not a duplicate of this row.

## Outbound replies are events too

`graph.OutboundMessage` derives the outbound normalized event from the inbound
trigger: same perspective, same connector, same conversation, `direction =
'outbound'`, and `parent_external_message_id` = the message that triggered the
turn.

The outbound author is this RuntimeAgent's own connector identity represented as
an Actor in its graph. No Actor classification is required. Another RuntimeAgent
that later observes the Slack reply resolves that external identity as an
ordinary Actor in *its* perspective; it receives no hidden link back to the
sender's RuntimeAgent row.

## Ordering guarantees in the runtime

`app.Runtime.HandleSlack` persists the inbound event **before** calling the
model. If the turn fails we still have a durable, attributable record that the
message arrived, which also makes the turn retryable. Losing history is worse
than losing a reply.

The outbound event is persisted **after** delivery, because we only learn the
external message id from the send. If that write fails the error is surfaced but
the send result is still returned, so the caller does not retry a send that
already happened.

These guarantees apply within one RuntimeAgent perspective. A failure in one
perspective must not cause another RuntimeAgent to reuse or inherit its partially
written graph state.

## Context reads

Graph history reads are perspective-scoped for the same reason writes are.
Context assembly must never query deployment-global actor/conversation history
and filter it down afterward. The owning RuntimeAgent is part of the read key.
See [`runtime/context-resolution.md`](/runtime/context-resolution).

This matters especially for actor-centric history: Lyra's Victor and Daedalus's
Victor may be distinct Actor rows with different memory and conversation
histories even when the control plane can correlate their external identities.

## A note on jsonb and the simple query protocol

Mesh runs pgx in simple query protocol mode so it can sit behind PgBouncer in
transaction pooling mode. In that mode a Go `[]byte` parameter is encoded as a
bytea literal, which Postgres rejects for a `jsonb` column
(`invalid input syntax for type json`).

Queries on the hot path therefore avoid binding `jsonb` as a parameter and let
the column default to `'{}'`. Anything that needs to write real metadata must
cast explicitly (`$n::jsonb`) and pass `string`, not `[]byte`.

## Authoring a migration

Add a new file to `db/migrations/` named `<version>_<name>.sql`. The version is a
**UTC timestamp prefix**, generated once at authoring time:

```bash theme={null}
./scripts/new-migration.sh add_retract_idle
# -> db/migrations/20260925170000_add_retract_idle.sql
```

The runner (`internal/migrate`) applies **any migration whose version the ledger
does not already record**, in numeric version order, and skips the ones it does
— the apply-by-ledger-presence model Rails, Django, and goose all use. Numeric
order is the order a *batch* applies in, **not** a precondition on which files
may apply: a pending file that sorts *below* the highest applied version is
applied, not refused. Independent migrations authored on parallel branches both
apply in either merge order, with no renumber (MESH-366). The runner is still
forward-only in the sense that there are no down-migrations; it just no longer
mistakes "lower version" for "already superseded."

> This is a deliberate change from the original total-order guard, which refused
> any below-max pending file. That guard turned a benign parallel-authoring case
> — a new 4-digit file (`0098`) sorting below a timestamp-prefixed applied max —
> into a fleet-wide Fly boot outage (MESH-364): every redeploying instance hit
> the refusal, `exit(1)`'d, and crash-looped. Independence is now the default.

### Why timestamps, not a sequential counter

Mesh used a zero-padded counter (`0001`, `0002`, ...) through `0095`. A counter
forces a *total* order on work that is really a *partial* order: two branches
developed in parallel each grab the same next integer and collide the moment
they both merge, forcing a hand-renumber (see PR #415, renumbered `0091`→`0095`).
A timestamp is stamped per branch at authoring time, so parallel work almost
never collides. Two files that *do* land on the same second are a hard error,
not a silent tiebreak: the prefix is the version, so both carry the same version
and the loader rejects them as a duplicate. `scripts/new-migration.sh` handles
this for you by advancing to the next free second, but if you hand-name a file
you must give it a prefix no other migration uses.

Legacy `00xx` files and new timestamp files coexist in one directory: the runner
orders by numeric value, so `95 < 20260925170000` orders correctly even though
`"0095"` is lexically smaller only by length. Do **not** renumber the existing
`00xx` migrations — their strings are live ledger keys. (New `00xx` files are
separately banned by CI; timestamps only — see MESH-365. Ordering below a
timestamp max is no longer fatal, but a fresh 4-digit prefix is still rejected
at the CI gate before it can land.)

### The one thing the clock does not buy you

Timestamps give an *order*, not a *dependency guarantee*. If migration B truly
requires the schema migration A builds, wall-clock skew between two machines is
not a safe way to order them. Put dependent changes in the **same file / same
PR**. Independent changes (a new column here, a new index there — the 99% case)
need no ordering relationship at all, which is exactly why timestamps are safe
for them.

## Running the persistence tests

The graph tests include a real Postgres round trip, skipped unless a database is
configured:

```bash theme={null}
createdb mesh_test
for f in db/migrations/*.sql; do psql -d mesh_test -v ON_ERROR_STOP=1 -f "$f"; done
MESH_TEST_DATABASE_URL="postgres:///mesh_test" go test ./internal/graph/...
```

CI provisions a `postgres:16` service and applies the migrations, so the SQL is
exercised on every run.

Perspective-isolation coverage exists and must stay:

* `internal/graph` proves two Stores recording the same external event produce
  distinct connector, conversation, actor, and message-event rows.
* `internal/admin/inspection_integration_test.go` runs two RuntimeAgents through
  the real write path in the **same namespace with a shared external user id**,
  and proves no actor, conversation, identity, history row, or engagement verdict
  crosses — including that a write handed a foreign conversation or actor id
  inserts nothing.
* `internal/migrate` proves the ownership backfill derives correctly — including
  for a connector written under an install's *earlier* resolved namespace — and
  refuses ambiguous input, leaving no DDL and no ledger row behind.
* `internal/app` proves each agent's recorder, history reader, and engagement
  reader are wired to that agent's own perspective.

A fixture that gives the two agents different namespaces is not testing the
boundary — that shape was already isolated by accident before the ownership key
existed.

## Regenerating the query layer

`internal/store` is generated by [sqlc](https://sqlc.dev) from `db/queries/*.sql`
against `db/migrations/*.sql`. Never edit it by hand:

```bash theme={null}
sqlc generate
```

This is why removal of legacy fields such as `Actor.kind` should land as a real
schema/query regeneration change rather than by hand-editing generated Go.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.