> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Secrets and recovery

# Secrets and recovery

The master key, its custody, its rotation, and the drill that proves a backup
restores. This is the runbook for the two failure narratives
[`roadmap/tracks.md` §E4 (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap/tracks.md#e4--secrets-and-recovery) names: the
key leaks and there is no rotation path, or the key is lost and every ciphertext
is permanently opaque.

## What is encrypted, and with what

Every credential Mesh holds — Slack bot tokens and signing secrets, Telegram
tokens, model-provider keys, GitHub and Linear tokens, the error tracker's DSN,
memory-provider keys, agent-runtime tokens — is a row in the `secrets` table,
sealed with AES-256-GCM under one **master key**. The key comes from
`MESH_AGENT_MASTER_KEY` in the process environment and is never written to the
database; ciphertext and key must have different blast radii, or a database dump
is plaintext. `agent-provisioning.md` explains why this is the one credential an
operator holds personally.

Each row carries a **key version tag**. Rows written before rotation existed
carry the legacy tag `v1`, which says nothing about which key wrote them. Every
row written since carries the writing key's **fingerprint**: `sha256:` plus the
first 16 hex characters of the SHA-256 of the key material. A fingerprint names a
key without revealing it. The `secret_keys` row named `agent:credentials` — the
**descriptor** — carries the fingerprint of the key that last wrote or rotated
the table, and it is what boot reads first.

## Key custody

* **Where the key lives:** in the deployment platform's secret store, injected
  as `MESH_AGENT_MASTER_KEY` into the process environment. Nowhere in the
  repository, nowhere in the database, nowhere in a backup of the database.
* **Who can read it:** whoever can read the deployment's secrets. That set is
  the set of people who can read every credential Mesh holds; keep it small and
  written down for your deployment.
* **Where it is backed up:** the key must be held in at least one place that is
  not the deployment platform — a password manager entry or an offline record —
  because a lost key makes every backup worthless. The database backup and the
  key backup must never be the same artifact.
* **What the process does about a wrong key:** refuses to boot. If the
  descriptor names a fingerprint and the process was handed a different key,
  `mesh` exits with one line naming both fingerprints and the remedy, before the
  agent loader touches a single credential. A legacy `v1` descriptor cannot be
  judged that way; the loader's per-credential decrypt is the only test for it.

## The commands

Both are subcommands of the `mesh` binary, run against `DATABASE_URL` with no
other serving configuration, like `mesh migrate`. Neither prints a key or a
plaintext; fingerprints and counts are the whole output.

```sh theme={null}
DATABASE_URL=... MESH_AGENT_MASTER_KEY=... mesh verify-secrets
```

Decrypts every row with the key and prints, per owner type and kind, how many
opened and how many did not, plus whether the descriptor names this key. Exits
non-zero if anything failed to open or the descriptor names another key. It
writes nothing and takes no locks.

```sh theme={null}
DATABASE_URL=... MESH_AGENT_MASTER_KEY=<current> MESH_AGENT_MASTER_KEY_NEXT=<next> mesh rotate-key [--dry-run]
```

Re-encrypts every row from the current key to the next in **one transaction**,
then moves the descriptor. Properties worth knowing:

* **Atomic.** The table is never half-rotated. Any row that cannot be opened
  with the current key (or, see below, the next) stops the run and rolls
  everything back; a row it cannot read is a row it must not overwrite.
* **Locked.** The transaction takes an `EXCLUSIVE` lock on `secrets` and
  `secret_keys` for its duration: readers proceed (a live process keeps
  decrypting), writers wait. A concurrent dashboard save cannot insert a row
  the rotation never saw; it lands after the commit, under whichever key that
  process holds, where the boot check will see it.
* **Idempotent, without trusting tags.** A row tagged with the next key is
  skipped only after it actually opens with the next key; a mis-tagged or
  corrupt row is refused, not skipped. So a run interrupted by a lost
  connection is simply run again.
* **Handles the restart race.** A row written under the *next* key by a process
  that predates the rotation but tagged `v1` is restamped, not refused.
* **Refuses the no-op.** The same key as current and next is an error.
* `--dry-run` computes the full report and commits nothing.

## Runbook: rotating the master key

The serving process reads with one key and does not accept two at once, so a
rotation is a restart. The clean sequence is **stop, rotate, start with the next
key**: downtime is the stop-to-start interval, and the rotation itself is one
transaction over a small table, so that interval is dominated by the restart.

Rotating while the process is still serving is possible — the table lock makes
writers wait rather than slip past — but any credential saved after the
rotation commits and before the restart is encrypted under the old key, and the
process will then refuse to boot with the next key until `rotate-key` is run
again. Stopping first removes that step; the boot check makes forgetting it
safe rather than silent.

The master key also keys the query digests in two audit tables,
`memory_retrieval_events` and `conversation_search_events` (`internal/auditdigest`),
so a query cannot be recovered from them by hashing guesses. Those digests are
not re-encrypted by `rotate-key`, because they are one-way: rows written before
a rotation stay comparable only with each other, and rows after it start a new
digest space. Nothing in Mesh reads them to make a decision.

1. Generate the next key. Any high-entropy string works; 32 random bytes,
   base64-encoded, is the shape the code normalizes to.

2. Store the next key in the secret store **before** rotating, under a name that
   distinguishes it from the current one. A rotation whose next key is lost is
   the worst outcome in this document.

3. Stop the serving process (scale to zero, or stop the container). Skipping
   this is allowed; see above for what it costs.

4. Dry-run from an operator shell against the production database:

   ```sh theme={null}
   MESH_AGENT_MASTER_KEY=<current> MESH_AGENT_MASTER_KEY_NEXT=<next> mesh rotate-key --dry-run
   ```

   Read the per-kind counts. They are the inventory of what will move.

5. Rotate:

   ```sh theme={null}
   MESH_AGENT_MASTER_KEY=<current> MESH_AGENT_MASTER_KEY_NEXT=<next> mesh rotate-key
   ```

6. Update the deployment's `MESH_AGENT_MASTER_KEY` to the next key and start
   the process. Do not set `MESH_AGENT_MASTER_KEY_NEXT` on the serving process;
   the service never reads it.

7. Verify from the operator shell with the next key:

   ```sh theme={null}
   MESH_AGENT_MASTER_KEY=<next> mesh verify-secrets
   ```

8. Retire the old key from the secret store once verify is clean and the
   service is serving.

**If boot fails at step 6** with "MESH\_AGENT\_MASTER\_KEY is not the key the
secrets table is encrypted under": step 3 was skipped and a credential was saved
under the old key between steps 5 and 6. Run step 5 again with the same current
and next keys — it rotates the straggler and skips the rest — then start again.

**If the old key has leaked**, rotation is step one of the compromise runbook
below, not the whole of it.

## Runbook: backup and restore, with the drill

The database is the system of record for every credential, every memory, and
the whole conversation graph. The key is not in it. A backup therefore restores
only if the key that encrypted it is still held; the drill proves both halves
together.

**Backup.** Use the platform's point-in-time recovery where it exists, and take
a logical dump on the same schedule as a belt-and-braces artifact that can be
restored anywhere:

```sh theme={null}
pg_dump --no-owner "$DATABASE_URL" > mesh-$(date -u +%Y%m%dT%H%MZ).sql
```

Only the `secrets` column values are ciphertext, and the dump never contains the
master key. Everything else in the dump — conversation bodies, memories, prompt
snapshots, model responses — is plaintext. Write the dump with restrictive
permissions, keep it in encrypted backup storage, and give it the same access
controls as the production database; a dump is the database.

**Restore drill.** Quarterly, and after any change to the secrets layer:

1. Create a scratch database and restore the most recent dump into it.
2. Run `mesh verify-secrets` against the scratch database with the production
   key. Every row must open and the descriptor must match.
3. Run it again with a deliberately wrong key. Every row must fail and the
   descriptor must be reported as not matching. This proves the check is real.
4. Optionally rehearse rotation on the scratch copy (`rotate-key` to a throwaway
   key, `verify-secrets` with it, then confirm a boot with the old key against
   the scratch copy is refused).
5. Drop the scratch database. Record the date below.

| Drill date | Dump | Result |
| - | - | - |
| 2026-09-10 | `pg_dump` 16 of a migrated database seeded with 4 secrets across 4 kinds and 2 owner types, restored into a fresh database on the same server | `verify-secrets` with the production key: 4 of 4 opened, descriptor matched. Wrong key: 0 of 4, descriptor mismatch reported, exit 1. `rotate-key` on the restored copy: 4 rotated, descriptor moved; `verify-secrets` with the next key: ok. Boot with the old key against the rotated copy: refused at load with the wrong-key message naming both fingerprints. |

The next drill is due by 2026-12-10. The most recent date is also written in
the README so a reader who never opens this page still sees it.

## Runbook: credential compromise

What to do, in order, when a credential Mesh holds has leaked.

**If the master key leaked:** everything it protected is exposed. Rotate the
master key first (runbook above), because until then every credential you
re-issue is stored under a key the attacker holds. Then treat every credential
in the table as compromised and work down the inventory `rotate-key` printed:

1. Slack: regenerate each connector's signing secret and reinstall the app to
   rotate its bot token; paste both through the connector settings.
2. Telegram: revoke and re-issue each bot token with BotFather; paste it.
3. Model providers: revoke each key in the provider console; paste the new one
   per agent, then the deployment fallback if one is set.
4. GitHub and Linear: revoke the agent accounts' tokens in those systems and
   capture new ones on the agent's tools tab.
5. Error tracker and memory providers: rotate at the vendor, paste in settings.

**If one credential leaked:** revoke it at the issuer and paste the replacement.
The master key is unaffected and need not rotate.

**If the master key is lost:** there is no recovery for the ciphertext. Delete
the `secrets` rows, set a new key, and re-capture every credential through the
UI. `verify-secrets` with any key will list what existed, by kind, so the
re-capture is a checklist rather than a guess.

## What is not built

* **Dual-key read.** The serving process reads with one key; rotation is a
  restart. The tags are load-bearing enough that a process accepting a
  `MESH_AGENT_MASTER_KEY_PREVIOUS` for reads is a bounded follow-on once
  zero-restart rotation matters (more than one replica).
* **KMS or Vault as the key provider.** `secret_keys.provider` admits `kms` and
  `vault`; only `env` is implemented.
* **Automated drill scheduling.** The quarterly cadence is a calendar entry.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.