Skip to main content

Secrets and recovery

The master key, its custody, its rotation, and the drill that proves a backup restores. This is the runbook for the two failure narratives roadmap/tracks.md §E4 (repository) names: the key leaks and there is no rotation path, or the key is lost and every ciphertext is permanently opaque.

What is encrypted, and with what

Every credential Mesh holds — Slack bot tokens and signing secrets, Telegram tokens, model-provider keys, GitHub and Linear tokens, the error tracker’s DSN, memory-provider keys, agent-runtime tokens — is a row in the secrets table, sealed with AES-256-GCM under one master key. The key comes from MESH_AGENT_MASTER_KEY in the process environment and is never written to the database; ciphertext and key must have different blast radii, or a database dump is plaintext. agent-provisioning.md explains why this is the one credential an operator holds personally. Each row carries a key version tag. Rows written before rotation existed carry the legacy tag v1, which says nothing about which key wrote them. Every row written since carries the writing key’s fingerprint: sha256: plus the first 16 hex characters of the SHA-256 of the key material. A fingerprint names a key without revealing it. The secret_keys row named agent:credentials — the descriptor — carries the fingerprint of the key that last wrote or rotated the table, and it is what boot reads first.

Key custody

  • Where the key lives: in the deployment platform’s secret store, injected as MESH_AGENT_MASTER_KEY into the process environment. Nowhere in the repository, nowhere in the database, nowhere in a backup of the database.
  • Who can read it: whoever can read the deployment’s secrets. That set is the set of people who can read every credential Mesh holds; keep it small and written down for your deployment.
  • Where it is backed up: the key must be held in at least one place that is not the deployment platform — a password manager entry or an offline record — because a lost key makes every backup worthless. The database backup and the key backup must never be the same artifact.
  • What the process does about a wrong key: refuses to boot. If the descriptor names a fingerprint and the process was handed a different key, mesh exits with one line naming both fingerprints and the remedy, before the agent loader touches a single credential. A legacy v1 descriptor cannot be judged that way; the loader’s per-credential decrypt is the only test for it.

The commands

Both are subcommands of the mesh binary, run against DATABASE_URL with no other serving configuration, like mesh migrate. Neither prints a key or a plaintext; fingerprints and counts are the whole output.
Decrypts every row with the key and prints, per owner type and kind, how many opened and how many did not, plus whether the descriptor names this key. Exits non-zero if anything failed to open or the descriptor names another key. It writes nothing and takes no locks.
Re-encrypts every row from the current key to the next in one transaction, then moves the descriptor. Properties worth knowing:
  • Atomic. The table is never half-rotated. Any row that cannot be opened with the current key (or, see below, the next) stops the run and rolls everything back; a row it cannot read is a row it must not overwrite.
  • Locked. The transaction takes an EXCLUSIVE lock on secrets and secret_keys for its duration: readers proceed (a live process keeps decrypting), writers wait. A concurrent dashboard save cannot insert a row the rotation never saw; it lands after the commit, under whichever key that process holds, where the boot check will see it.
  • Idempotent, without trusting tags. A row tagged with the next key is skipped only after it actually opens with the next key; a mis-tagged or corrupt row is refused, not skipped. So a run interrupted by a lost connection is simply run again.
  • Handles the restart race. A row written under the next key by a process that predates the rotation but tagged v1 is restamped, not refused.
  • Refuses the no-op. The same key as current and next is an error.
  • --dry-run computes the full report and commits nothing.

Runbook: rotating the master key

The serving process reads with one key and does not accept two at once, so a rotation is a restart. The clean sequence is stop, rotate, start with the next key: downtime is the stop-to-start interval, and the rotation itself is one transaction over a small table, so that interval is dominated by the restart. Rotating while the process is still serving is possible — the table lock makes writers wait rather than slip past — but any credential saved after the rotation commits and before the restart is encrypted under the old key, and the process will then refuse to boot with the next key until rotate-key is run again. Stopping first removes that step; the boot check makes forgetting it safe rather than silent. The master key also keys the query digests in two audit tables, memory_retrieval_events and conversation_search_events (internal/auditdigest), so a query cannot be recovered from them by hashing guesses. Those digests are not re-encrypted by rotate-key, because they are one-way: rows written before a rotation stay comparable only with each other, and rows after it start a new digest space. Nothing in Mesh reads them to make a decision.
  1. Generate the next key. Any high-entropy string works; 32 random bytes, base64-encoded, is the shape the code normalizes to.
  2. Store the next key in the secret store before rotating, under a name that distinguishes it from the current one. A rotation whose next key is lost is the worst outcome in this document.
  3. Stop the serving process (scale to zero, or stop the container). Skipping this is allowed; see above for what it costs.
  4. Dry-run from an operator shell against the production database:
    Read the per-kind counts. They are the inventory of what will move.
  5. Rotate:
  6. Update the deployment’s MESH_AGENT_MASTER_KEY to the next key and start the process. Do not set MESH_AGENT_MASTER_KEY_NEXT on the serving process; the service never reads it.
  7. Verify from the operator shell with the next key:
  8. Retire the old key from the secret store once verify is clean and the service is serving.
If boot fails at step 6 with “MESH_AGENT_MASTER_KEY is not the key the secrets table is encrypted under”: step 3 was skipped and a credential was saved under the old key between steps 5 and 6. Run step 5 again with the same current and next keys — it rotates the straggler and skips the rest — then start again. If the old key has leaked, rotation is step one of the compromise runbook below, not the whole of it.

Runbook: backup and restore, with the drill

The database is the system of record for every credential, every memory, and the whole conversation graph. The key is not in it. A backup therefore restores only if the key that encrypted it is still held; the drill proves both halves together. Backup. Use the platform’s point-in-time recovery where it exists, and take a logical dump on the same schedule as a belt-and-braces artifact that can be restored anywhere:
Only the secrets column values are ciphertext, and the dump never contains the master key. Everything else in the dump — conversation bodies, memories, prompt snapshots, model responses — is plaintext. Write the dump with restrictive permissions, keep it in encrypted backup storage, and give it the same access controls as the production database; a dump is the database. Restore drill. Quarterly, and after any change to the secrets layer:
  1. Create a scratch database and restore the most recent dump into it.
  2. Run mesh verify-secrets against the scratch database with the production key. Every row must open and the descriptor must match.
  3. Run it again with a deliberately wrong key. Every row must fail and the descriptor must be reported as not matching. This proves the check is real.
  4. Optionally rehearse rotation on the scratch copy (rotate-key to a throwaway key, verify-secrets with it, then confirm a boot with the old key against the scratch copy is refused).
  5. Drop the scratch database. Record the date below.
The next drill is due by 2026-12-10. The most recent date is also written in the README so a reader who never opens this page still sees it.

Runbook: credential compromise

What to do, in order, when a credential Mesh holds has leaked. If the master key leaked: everything it protected is exposed. Rotate the master key first (runbook above), because until then every credential you re-issue is stored under a key the attacker holds. Then treat every credential in the table as compromised and work down the inventory rotate-key printed:
  1. Slack: regenerate each connector’s signing secret and reinstall the app to rotate its bot token; paste both through the connector settings.
  2. Telegram: revoke and re-issue each bot token with BotFather; paste it.
  3. Model providers: revoke each key in the provider console; paste the new one per agent, then the deployment fallback if one is set.
  4. GitHub and Linear: revoke the agent accounts’ tokens in those systems and capture new ones on the agent’s tools tab.
  5. Error tracker and memory providers: rotate at the vendor, paste in settings.
If one credential leaked: revoke it at the issuer and paste the replacement. The master key is unaffected and need not rotate. If the master key is lost: there is no recovery for the ciphertext. Delete the secrets rows, set a new key, and re-capture every credential through the UI. verify-secrets with any key will list what existed, by kind, so the re-capture is a checklist rather than a guess.

What is not built

  • Dual-key read. The serving process reads with one key; rotation is a restart. The tags are load-bearing enough that a process accepting a MESH_AGENT_MASTER_KEY_PREVIOUS for reads is a bounded follow-on once zero-restart rotation matters (more than one replica).
  • KMS or Vault as the key provider. secret_keys.provider admits kms and vault; only env is implemented.
  • Automated drill scheduling. The quarterly cadence is a calendar entry.