Secrets and recovery
The master key, its custody, its rotation, and the drill that proves a backup restores. This is the runbook for the two failure narrativesroadmap/tracks.md §E4 (repository) names: the
key leaks and there is no rotation path, or the key is lost and every ciphertext
is permanently opaque.
What is encrypted, and with what
Every credential Mesh holds — Slack bot tokens and signing secrets, Telegram tokens, model-provider keys, GitHub and Linear tokens, the error tracker’s DSN, memory-provider keys, agent-runtime tokens — is a row in thesecrets table,
sealed with AES-256-GCM under one master key. The key comes from
MESH_AGENT_MASTER_KEY in the process environment and is never written to the
database; ciphertext and key must have different blast radii, or a database dump
is plaintext. agent-provisioning.md explains why this is the one credential an
operator holds personally.
Each row carries a key version tag. Rows written before rotation existed
carry the legacy tag v1, which says nothing about which key wrote them. Every
row written since carries the writing key’s fingerprint: sha256: plus the
first 16 hex characters of the SHA-256 of the key material. A fingerprint names a
key without revealing it. The secret_keys row named agent:credentials — the
descriptor — carries the fingerprint of the key that last wrote or rotated
the table, and it is what boot reads first.
Key custody
- Where the key lives: in the deployment platform’s secret store, injected
as
MESH_AGENT_MASTER_KEYinto the process environment. Nowhere in the repository, nowhere in the database, nowhere in a backup of the database. - Who can read it: whoever can read the deployment’s secrets. That set is the set of people who can read every credential Mesh holds; keep it small and written down for your deployment.
- Where it is backed up: the key must be held in at least one place that is not the deployment platform — a password manager entry or an offline record — because a lost key makes every backup worthless. The database backup and the key backup must never be the same artifact.
- What the process does about a wrong key: refuses to boot. If the
descriptor names a fingerprint and the process was handed a different key,
meshexits with one line naming both fingerprints and the remedy, before the agent loader touches a single credential. A legacyv1descriptor cannot be judged that way; the loader’s per-credential decrypt is the only test for it.
The commands
Both are subcommands of themesh binary, run against DATABASE_URL with no
other serving configuration, like mesh migrate. Neither prints a key or a
plaintext; fingerprints and counts are the whole output.
- Atomic. The table is never half-rotated. Any row that cannot be opened with the current key (or, see below, the next) stops the run and rolls everything back; a row it cannot read is a row it must not overwrite.
- Locked. The transaction takes an
EXCLUSIVElock onsecretsandsecret_keysfor its duration: readers proceed (a live process keeps decrypting), writers wait. A concurrent dashboard save cannot insert a row the rotation never saw; it lands after the commit, under whichever key that process holds, where the boot check will see it. - Idempotent, without trusting tags. A row tagged with the next key is skipped only after it actually opens with the next key; a mis-tagged or corrupt row is refused, not skipped. So a run interrupted by a lost connection is simply run again.
- Handles the restart race. A row written under the next key by a process
that predates the rotation but tagged
v1is restamped, not refused. - Refuses the no-op. The same key as current and next is an error.
--dry-runcomputes the full report and commits nothing.
Runbook: rotating the master key
The serving process reads with one key and does not accept two at once, so a rotation is a restart. The clean sequence is stop, rotate, start with the next key: downtime is the stop-to-start interval, and the rotation itself is one transaction over a small table, so that interval is dominated by the restart. Rotating while the process is still serving is possible — the table lock makes writers wait rather than slip past — but any credential saved after the rotation commits and before the restart is encrypted under the old key, and the process will then refuse to boot with the next key untilrotate-key is run
again. Stopping first removes that step; the boot check makes forgetting it
safe rather than silent.
The master key also keys the query digests in two audit tables,
memory_retrieval_events and conversation_search_events (internal/auditdigest),
so a query cannot be recovered from them by hashing guesses. Those digests are
not re-encrypted by rotate-key, because they are one-way: rows written before
a rotation stay comparable only with each other, and rows after it start a new
digest space. Nothing in Mesh reads them to make a decision.
- Generate the next key. Any high-entropy string works; 32 random bytes, base64-encoded, is the shape the code normalizes to.
- Store the next key in the secret store before rotating, under a name that distinguishes it from the current one. A rotation whose next key is lost is the worst outcome in this document.
- Stop the serving process (scale to zero, or stop the container). Skipping this is allowed; see above for what it costs.
-
Dry-run from an operator shell against the production database:
Read the per-kind counts. They are the inventory of what will move.
-
Rotate:
-
Update the deployment’s
MESH_AGENT_MASTER_KEYto the next key and start the process. Do not setMESH_AGENT_MASTER_KEY_NEXTon the serving process; the service never reads it. -
Verify from the operator shell with the next key:
- Retire the old key from the secret store once verify is clean and the service is serving.
Runbook: backup and restore, with the drill
The database is the system of record for every credential, every memory, and the whole conversation graph. The key is not in it. A backup therefore restores only if the key that encrypted it is still held; the drill proves both halves together. Backup. Use the platform’s point-in-time recovery where it exists, and take a logical dump on the same schedule as a belt-and-braces artifact that can be restored anywhere:secrets column values are ciphertext, and the dump never contains the
master key. Everything else in the dump — conversation bodies, memories, prompt
snapshots, model responses — is plaintext. Write the dump with restrictive
permissions, keep it in encrypted backup storage, and give it the same access
controls as the production database; a dump is the database.
Restore drill. Quarterly, and after any change to the secrets layer:
- Create a scratch database and restore the most recent dump into it.
- Run
mesh verify-secretsagainst the scratch database with the production key. Every row must open and the descriptor must match. - Run it again with a deliberately wrong key. Every row must fail and the descriptor must be reported as not matching. This proves the check is real.
- Optionally rehearse rotation on the scratch copy (
rotate-keyto a throwaway key,verify-secretswith it, then confirm a boot with the old key against the scratch copy is refused). - Drop the scratch database. Record the date below.
The next drill is due by 2026-12-10. The most recent date is also written in
the README so a reader who never opens this page still sees it.
Runbook: credential compromise
What to do, in order, when a credential Mesh holds has leaked. If the master key leaked: everything it protected is exposed. Rotate the master key first (runbook above), because until then every credential you re-issue is stored under a key the attacker holds. Then treat every credential in the table as compromised and work down the inventoryrotate-key printed:
- Slack: regenerate each connector’s signing secret and reinstall the app to rotate its bot token; paste both through the connector settings.
- Telegram: revoke and re-issue each bot token with BotFather; paste it.
- Model providers: revoke each key in the provider console; paste the new one per agent, then the deployment fallback if one is set.
- GitHub and Linear: revoke the agent accounts’ tokens in those systems and capture new ones on the agent’s tools tab.
- Error tracker and memory providers: rotate at the vendor, paste in settings.
secrets rows, set a new key, and re-capture every credential through the
UI. verify-secrets with any key will list what existed, by kind, so the
re-capture is a checklist rather than a guess.
What is not built
- Dual-key read. The serving process reads with one key; rotation is a
restart. The tags are load-bearing enough that a process accepting a
MESH_AGENT_MASTER_KEY_PREVIOUSfor reads is a bounded follow-on once zero-restart rotation matters (more than one replica). - KMS or Vault as the key provider.
secret_keys.provideradmitskmsandvault; onlyenvis implemented. - Automated drill scheduling. The quarterly cadence is a calendar entry.