Access Control: Identity-Scoped Agents, Provenance Privacy, and Tenant/Actor Encryption
This document proposes a three-layer access-control model for Mesh, layered on the existingperspective boundary (perspective.md, migration
0016). It covers equating engaged actors with a durable identity, agent-type
engagement rules, provenance-based conversation and memory privacy, and tenanted
encryption at rest. Where the Signal Protocol has already solved an analogous
problem, the relevant specification is cited as prior art — at the design level, not
the wire protocol, since Mesh’s runtime must read plaintext to reason and is
therefore not an end-to-end system (see §3.2). Open decisions are collected in the
final section.
0. Framing: three layers, one model
There are three distinct concerns here, but they are not three features — they are three layers of one access-control model, and the whole design only holds together if they share a spine. That spine already exists in Mesh and is named perspective (perspective.md, migration 0016): a
RuntimeAgent owns its actors, conversations, messages, and memory, and the boundary
is a schema constraint, not a prompt convention.
This model adds a second axis of isolation inside a single perspective — an
actor-audience boundary — and a third axis underneath the whole store —
at-rest confidentiality (encryption). Stated as a stack:
Enforcement lives at the point where data is loaded, in the query, hardcoded. Never in the prompt. By the time a private fact is in the context window it has already leaked. The control is: do not retrieve ineligible data. This is the foundational rule the rest of the document builds on.
runtime/memory-retrieval.md already codifies half
of this: “Subject changes relevance. Visibility changes eligibility.” and “No
ranking score may restore a row rejected here.” This model extends that
eligibility filter from a two-value enum (perspective/conversation) to a real
actor-audience model, and pushes the same eligibility discipline into conversation
reads and the operator API.
Layer 1 — Agent types (personal / assistant / open)
The construct
Every RuntimeAgent gets a type chosen at creation, and — for the first two types — a principal actor it is bound to. The type is a hardcoded gate on which actors the agent will engage with at all, evaluated in the engagement gate (runtime/addressing.md) before any inference.
open is chosen over “general”/“group” as the third type name: it is the honest
antonym of “personal” and doesn’t collide with “group DM” (a conversation shape,
not an agent type). Alternatives considered: shared, public. open wins
because it describes the audience policy, not the deployment.
Background: how OpenClaw handles this today, and why it isn’t enough
For readers coming from OpenClaw, some grounding. OpenClaw has a concept called required mention. For each channel an agent is whitelisted in, the person configuring the agent chooses whether an explicit@mention is needed for the
agent to respond: with required-mention on, the agent replies only when directly
addressed; with it off, it responds to messages in that channel generally. It is a
per-channel, per-agent flag that answers one question — “should I speak to this
message?” — evaluated at ingress, message by message.
That is the whole of OpenClaw’s engagement control, and the reason is structural:
OpenClaw has no actor model. It only has channels. It can decide whether to
speak in a channel; it cannot reason about who an agent exists for, because it has
no durable, cross-channel notion of a person distinct from the channel they
appeared in.
Agent type is a different and stronger construct. require_mention answers “should
I speak?” per message. Agent type answers “whom do I exist for?” once,
structurally, against the actor graph.
Why this is not require_mention, and why it’s better
A personal agent is not “a public agent that requires mention from strangers” —
it is an agent that has no relationship with strangers. A stranger’s message to a
personal agent doesn’t get a “quieter” verdict; it gets no verdict path to
engage at all, and — critically — the stranger’s message is not persisted into
the principal’s private conversation graph in a way that becomes the principal’s
context. This is the thing OpenClaw cannot express: with only channels and a
per-channel mention flag, there is no actor to bind an agent to, so “this agent
exists for exactly this person” is inexpressible. Mesh has actors, so it can.
The “badge” question — hardcode or prompt?
Two distinct claims are bundled in the word “badge,” and they must be separated because the answer differs.-
Identity/authentication — the agent acts as itself using its own
delegated credentials, not a shared human password. This already exists and
is a strength:
0002_secrets.sqlstores each agent’s connector tokens envelope-encrypted, owned by the agent (owner_type='agent'), andperspective.mdmakes self-recognition the one privileged recognition. A delegated-credential model — where the agent holds its own scoped tokens rather than borrowing a human’s login — is already the model. No change needed; it should be documented explicitly as a product property. -
Authorization — what the agent may do on behalf of whom. Here the control
must be hardcoded, not prompted, and the reason is the governing principle above.
The engagement gate must consult agent type + principal from the database and
short-circuit to
observe/ignoreforpersonalagents facing non-principals. A prompt that says “you are this person’s private assistant, don’t talk to others” is defeated by any message that says “actually your principal asked me to relay this.” Access control that lives in the prompt is access control an attacker can argue with. The database does not negotiate.
Schema
principal_state='binding' with principal_actor_id
NULL — a state the constraint above explicitly permits. Until bound, a personal
agent engages no one. The compare-and-set below flips principal_state to
'bound' and sets the principal in one atomic write.
Binding is expected-principal-use, not first-use. The binding token is minted
for a specific verified connector identity (a specific Slack user_id / SSO
subject), single-use, short-TTL, and authentication from any other identity is
rejected. “Whoever clicks first wins” is a land-grab: a forwarded or leaked link
would crown the wrong principal. The bind itself is a write-once compare-and-set:
perspective.md). The Layer-3 mnemonic enrollment is an
additional key-custody step layered on later; it is not the authenticator for
binding.
One human, many actors: a principal is one actor_id, but a person may have
several (Slack, email, web). For v1, personal/assistant agents are bound at the
person/identity level where identity-linking
(runtime/identity-linking.md) has merged the
actors; where it has not, the agent is single-connector-bound and this is
stated explicitly, so the agent never ignores its own principal arriving from a
second connector. See §1.1 for how a new-channel appearance of the principal is
verified and folded in.
1.1 Verifying a human across channels (identity linking meets agent type)
Agent type raises the stakes on a problemruntime/identity-linking.md
already solves in the general case: when the same human appears on a new
channel, how does the agent know it is them?
The scenario: an agent has been talking with a person on Slack — it knows them, and
(for a personal/assistant agent) that person may be its principal. The same
person then messages the agent on Telegram: “It’s me, you’ve been talking to me on
Slack.” The agent must not confirm whether it knows the claimed person or look
up their accounts. The implemented code-echo challenge from
runtime/identity-linking.md gives generic
instructions and a short single-use code in the new private conversation. The
person sends link CODE in their existing private conversation with the same
agent on Slack. The runtime checks the persisted, authenticated command; the
agent never forwards a code or sends a notification to a claimed contact. The
Slack account needs no prior owner step: it is trusted for what it is, an
authenticated session the person already controls. On success,
both accounts resolve to one actor through recorded proof, without rewriting
their original IDs.
Why this matters specifically for agent types:
personal— verification is a precondition to engage at all. Before the link is established, a new-channel claimant is a stranger, and apersonalagentobserves strangers (Tier-(-1) below). Only after the code echo folds the new identity into the principal’s actor does the agent engage that channel as the principal.assistant— verification decides which tier the counterpart gets. An unverified new-channel claimant is engaged as an ordinary (non-privileged) actor; a verified principal gets the privileged/confidential tier. The agent bounds what it will do based on which of those the challenge proves.open— typically doesn’t care; ordinary multiplayer engagement, no privileged principal to protect. Verification is only relevant if some capability is gated on a specific actor.
actor_id (the agent had no basis to assume otherwise).
After the code echo, identity-linking records a ConnectorIdentity edge folding
that identity under the existing actor. For personal/assistant agents this means
the principal binding follows the linked identity: once the Telegram identity is
proven to be the principal’s, principal-authored engagement and the confidential
tier extend to the Telegram channel automatically, without re-binding. Content
written under the pre-link actor_id retains its original attribution. Live proof
resolution associates those records on reads, so the association is auditable
and reversible if an anchor is later found compromised. The agent-type and
principal-binding behavior described here remains a proposal; the implemented
identity flow does not itself add those engagement tiers.
Prior art — the Signal Protocol’s session and authentication model. Two pieces of Signal’s design (signal.org/docs) map cleanly onto this problem and are worth borrowing at the design level — not the wire protocol, since Signal is a server-as-adversary end-to-end system and Mesh deliberately is not (the runtime must read plaintext to reason; see §3.2). Sesame (spec) manages encryption sessions in an asynchronous, multi-device setting: one identity fans out to many devices, each with its own live session, and the algorithm specifies exactly how a new device appearing mid-conversation is admitted. Thatidentity → devices → sessionstree is structurally the same shape as Mesh’sactor → connector-identities → live-sessions, and its new-device admission path is precisely the “principal appears on a new channel” case above — a ready-made template for how theConnectorIdentityedge and the Layer-3 DEK re-wrap should behave when a channel joins mid-stream. X3DH (spec) contributes the stronger authentication principle: authentication is the result of a key agreement, never a claim a party asserts. The code-echo challenge today proves human liveness on both ends but ultimately matches a string; a hardened variant has the two live sessions complete a short authenticated key agreement, producing a cryptographic binding — a signed, replay-resistant attestation rather than a shared secret typed across a gap. That pushes the governing principle one layer down: a matched string is access control an attacker can replay; a completed key agreement is not. Tracked as a hardening follow-up, not a v1 blocker for the string-based echo. (Signal’s VXEdDSA VRF, spec, is noted as future-only prior art for unlinkable-but-verifiable actor identifiers across channels, should the actor graph ever need selective disclosure — out of scope here.)
Engagement-gate change (Tier 0, deterministic, no inference)
Add above every existing Tier-0 signal inruntime/addressing.md:
personal’s non-principal → observe (not ignore) is deliberate: the agent may
still need ambient context (“someone else is in this thread with my principal”),
but it will never act for them. Whether a personal agent should even persist
non-principal messages is a Layer-2 question answered below.
Layer 2 — Provenance-based conversation & memory privacy
This is the core privacy concern and the highest-value correctness work. It is also already ~40% built and explicitly scheduled (roadmap A3 “Visibility and audience eligibility — bidirectional”,risks.md#r13). This section makes the model
concrete.
2.1 The two things being protected
- Conversation contents. In the Mesh dashboard, message bodies of a private conversation should be readable only by an operator acting for an actor who was a participant. Everyone (with operator access) may see that the conversation exists and its shape/metadata — never the text. (This is a Layer-2 authz rule on the operator API; the bytes may additionally be Layer-3 encrypted.)
- Memory. A memory’s eligibility for a turn is gated by where it was learned and for whom, independent of semantic relevance.
2.2 Provenance is already modeled — extend the visibility axis
0018 gave us subject_actor_id, source_actor_id, and a visibility enum of
{perspective, conversation}. That enum is deliberately minimal and the migration
comment invites its successor once an enforcement point exists. This is that
successor. Replace the two-value enum with a confidentiality band + audience
scope:
channel topology. The set of people who have spoken is
not the set of people who can read the room. Use authenticated connector privacy
flags and complete membership evidence:
The implemented runtime contract lives in memory privacy (repository).
It uses immutable revision evidence and explicit grants; the broader actor-linked
operator authorization and encryption layers proposed elsewhere in this document
remain separate work. The current dashboard is an authenticated operator surface.
Audience excludes the agent itself.audience_actorsis the set of human/external actors a memory is for, never the agent’s own self-actor.addressing.mdcounts the agent as a participant of adirectconversation, so if the agent were in the audience set the subset test would be computed over{human, agent}on both sides and the worked example below (audience one human) would be internally inconsistent.audience(T)andaudience(M)are both computed over non-self actors. A 1:1 DM’s audience is the single human counterpart.
2.3 The eligibility rule (bidirectional — this is the whole game)
risks.md#r13 is explicit that both directions are the same defect:
A memory learned in a DM must not be eligible in a public channel (leak outward), AND a memory planted in a hostile channel must not be eligible in the CEO’s DM (injection inward).The rule, enforced in SQL in the retrieval query (both automatic assembly and the
recall tool, per the “one retrieval service” contract):
The gate is band-first, one test per band. For a confidential memoryaudience(M)isM.audience_actors, soaudience(T) ⊆ audience(M)and “every actor in audience(T) ∈ M.audience_actors” would be the identical test written twice — a drift hazard if one is implemented and the other isn’t. And a literalaudience(T) ⊆ ∅would reject every turn for public/workspace memories (which have no audience rows). The subset test therefore applies only to the confidential band; public/workspace skip it.
2.3a Injection-inward is a SEPARATE control from the audience gate
The band-first gate above closes leak-outward (a DM secret never surfaces in a channel). It does not, by itself, close injection-inward — the other half ofr13. A memory whose origin was a hostile public channel defaults to band
public and is therefore eligible everywhere, including the CEO’s private DM.
Audience-subset does nothing here because public memories skip it. This is the
“planted in a hostile channel → eligible in the CEO’s DM” attack. Two-part fix:
- Trust is not a function of openness.
publicmeans “safe to re-emit widely,” not “safe to act on as instruction.” Public-origin content is eligible as retrieved context but is never elevated to instruction/steering. This is the prompt-injection boundary and it is enforced at assembly: channel-origin memories enter the context in a clearly-fenced, lower-trust region, never the system/authority region. (Ties to the existing turn-pipeline trust tiers.) - A private turn may down-scope public eligibility. For high-sensitivity turns
(e.g. a
personal/assistantagent’s principal-only DM), the retrieval profile may exclude channel-origin memories authored by actors who are not participants of the current turn — so a stranger’s channel post cannot ride into the principal’s private context at all. This is a per-turn policy knob, defaulted on for principal DMs.
The regression net must assert BOTH directions. Plant a memory in a hostile channel, assert it is not eligible as instruction in a principal DM and not retrieved when the down-scope knob is on — in addition to the leak-outward assertion.Concretely: a fact a person tells the agent in a 1:1 DM has
audience_actors = {that person}. It is eligible only in a turn whose participants
are a subset of that set — i.e. another 1:1 with the same person. The moment a third
actor is in the room, the audience is no longer a subset, and the memory is filtered
at the query, before ranking, exactly as
runtime/memory-retrieval.md demands. The “discreet
employee” behavior is not asked of the model; it is a WHERE clause the model
cannot see past.
Storage of audience. A confidential memory needs its permitted-audience actor
set. Rather than a JSON blob (unqueryable, un-constrainable), a join table:
Revision-versioning. BecauseThe subset testmemory_entriesrevisions are immutable and a revision can change band/audience, the audience must be keyed to the revision, not justmemory_entry_id— otherwise restoring an old revision retrieves it under a newer revision’s audience. v1 resolution: audience rows key on the memory-entry revision id (or an immutable per-revision snapshot), and retrieval joins the current revision.
audience(T) ⊆ M.audience becomes: no participant of T is absent
from memory_entry_audience, i.e. NOT EXISTS (participant p WHERE NOT EXISTS (audience row for p)). Enforced in the retrieval SQL. This is the same shape as
the existing conversation-visibility resolution, generalized.
2.4 Source is derived, never asserted (already true — keep it)
0018 already writes source_actor_id from the persisted inbound message’s graph
record, and no tool exposes a source parameter. This is the correct anti-spoofing
posture and this model preserves it: an attacker in a channel cannot write a memory
that claims a private origin, because band + audience are derived from the
conversation the message was actually observed in, not from the body.
2.5 Conversation-content visibility on the operator API
Distinct from memory: the dashboard read path. Today (risks.md#r16) every operator
login is a full admin with no roles and can read every actor’s entire history. This
model requires:
- Metadata vs. body split on the read API.
GET …/conversations/{id}returns shape, participants, timestamps, engagement verdicts — the technical/operational view — to any authorized operator of that agent. Message bodies of aconfidential-shape conversation require the requesting operator to be linked to a participant actor (Layer 2) and the key to decrypt (Layer 3, if enabled). - This is roadmap R16 / Track D6 (role column
owner/agent_operator, per-agent scoping middleware). This model adds the actor-linkage dimension: operator ⇄ actor identity mapping, so “may this human read this DM” is answerable.
2.6 What a personal agent persists
Layer 1’s contract is the single source of truth here: a stranger’s message is not
persisted into the principal’s private conversation graph in a way that becomes the
principal’s context. A naive implementation would persist a stranger’s message as
observe context with audience {principal, stranger}, which — under the subset
rule — IS eligible in a later principal-only turn ({principal} ⊆ {principal, stranger}), so stranger content would reach the principal’s context. That is
exactly what Layer 1 forbids, so the write path must not do it.
Resolved contract: a personal agent facing a non-principal does not write
principal-retrievable memory from that content. Non-principal turns produce
observe context that is used only for ambient situational awareness within that
same turn (“someone else is in the thread”) and is not distilled into a durable
memory eligible in any future turn. Concretely: memory-write for a personal agent
is gated on author == principal. This closes both the exfiltration side-channel (a
stranger can’t plant durable context) and the observe-injection path (stranger
content can’t ride a {principal, stranger}-audience memory into a {principal}
turn). Ambient awareness stays turn-local; nothing stranger-authored becomes
principal memory.
Departed-member shrink: the subset rule fails-closed on audience growth but
not on shrink. In a small_group , a memory scoped to stays
eligible for a later turn after C leaves. If the content was about or for
C, that is a leak. For the “for-this-exact-room” case, gate on conversation_id
equality (not just actor-set equality — two different conversations can share the
same participants); the general subset rule remains the default for durable facts
about the principal.
Layer 3 — Encryption at rest (tenant CMEK and actor-held E2E)
This is the tenanted-hosting requirement. It is the layer with the sharpest trade-offs and where the “actor holds the key” framing needs the most scrutiny.3.1 What already exists to build on
internal/secrets/secrets.go is envelope encryption already: AES-256-GCM with
a per-record nonce, key_version for rotation, provider ∈ {env, kms, vault}, and
a KMS production path is the stated direction (prior decision: GCP KMS). The secrets
table proves the pattern. Extending it to conversation bodies and confidential
memories is the work — same envelope shape, different (and vastly more numerous)
payloads.
3.2 The core tension, named before proposing anything
Two options are often floated as if parallel. They are different threat models and one of them fights the product:- Org-held key (CMEK): the tenant org controls a KMS key; Mesh encrypts all tenant data under data keys wrapped by it. Threat model: “the hosting provider’s DB dump / a stolen backup / a curious Mesh operator must not yield plaintext.” Mesh-at-runtime can decrypt (it calls the org’s KMS to unwrap), so the agent, dashboard, and memory all work normally. This is standard, sellable, and compatible with everything in Layers 1–2. This is the one to build first.
- Actor-held key (true E2E): only the actor holds the key; Mesh genuinely cannot decrypt. Threat model: “even Mesh-at-runtime must not read my DMs.” The problem: if Mesh can’t decrypt, the agent can’t read the message either. An LLM cannot reason over ciphertext. So “the agent responds to your DM” and “Mesh can never see the plaintext” are in direct tension. Pure E2E kills the dashboard, server-side memory retrieval, and the agent’s ability to function.
3.3 The resolution: per-conversation data keys, released to the runtime only at turn time
The honest version of “actor-held” that doesn’t break the product:- Each private conversation has a data encryption key (DEK), random, per conversation. Bodies + confidential memories for that conversation are encrypted under the DEK (AES-256-GCM, the existing primitive).
- The DEK is wrapped (envelope) once per key-holder:
- wrapped under the org KMS key (CMEK path) — always, if tenant CMEK is on;
- wrapped under each participant actor’s key (actor path) — if actor-held is on.
- At turn time, the runtime unwraps the DEK transiently in memory using whichever key it can access, decrypts what it needs for that turn, and never persists the plaintext or the unwrapped DEK. The plaintext lives only for the duration of the turn, in process memory, the same way the master key already does.
You store the ciphertext once (one DEK per conversation). You store the wrapped key N times — once per actor, once per org. You do NOT store the message N times.
The Layer-2 audience set and the Layer-3 wrapped-DEK set are NOT globally
identical. memory_entry_audience (Layer 2, per-memory, frozen at write) and
the wrapped-DEK set (Layer 3, per-conversation, mutable as membership changes)
diverge the instant membership changes — a per-conversation DEK wrapped to a
late-joiner would let them decrypt bodies and memories written before they
joined, which Layer 2 denies. Crypto would grant what authz forbids. §3.3a
resolves this with per-audience-epoch DEKs. The two sets are aligned by
construction only within an epoch, not globally.
3.3a DEK epochs — the fix for late-joiners and revocation
A single long-lived per-conversation DEK cannot express membership changes and cannot revoke. Resolution: DEKs are per-audience-epoch, not per-conversation.- The conversation has a current DEK. On any membership change (join, leave, actor removal) the DEK is rotated: a new epoch begins with a fresh DEK, and content from the new epoch onward is encrypted under it.
- A DEK is wrapped only to the actors who were members during its epoch. A late-joiner is wrapped the current epoch’s DEK, never prior epochs’ — so they cannot decrypt history from before they joined. A departed member’s wrapped copies of prior epochs are destroyed, and they are never wrapped the new epoch — so “remove the wrapped DEK = revoke” becomes true going forward.
- Honest limit: rotation revokes access to future ciphertext, not data already unwrapped/exfiltrated before revocation. This is standard crypto reality (you cannot un-see plaintext someone already read); the datasheet says so rather than implying deletion is retroactive secrecy.
Prior art — the Double Ratchet’s two properties, without the Double Ratchet. The per-audience-epoch DEK is deliberately the coarse-grained analogue of the forward-secrecy and post-compromise-security guarantees Signal’s Double Ratchet (spec) provides per message. Signal derives a fresh key for every message so that a key compromise exposes neither earlier traffic (forward secrecy) nor later traffic once the ratchet steps past it (post-compromise “self-healing”). Mesh does not need per-message ratcheting — the runtime holds a conversation’s plaintext for the whole turn anyway, so per-message keys would buy nothing against our actual threat model — but it does want the lesson: rotate on the events that matter (here, membership epochs rather than every message), keep the payload count per key bounded, and make revocation a key-lifecycle operation rather than a delete flag. The epoch boundary is our ratchet step; the HKDF per-record subkeys below are the intra-epoch equivalent of deriving distinct keys so no single key covers unbounded traffic. Framed this way, §3.3a is the ratchet’s intent mapped onto a server-side, plaintext-reading system, not a naive attempt to run the messaging ratchet where it doesn’t fit.AAD binding (defense-in-depth): every ciphertext’s AES-GCM Additional Authenticated Data binds
(runtime_agent_id, conversation_id, epoch, record_id), so
a ciphertext cannot be replayed under a different conversation/epoch even by someone
who can write the DB.
Nonce discipline at scale: AES-256-GCM with random 96-bit nonces has a
birthday-bound (~2³² messages per key before reuse risk). At “vastly more numerous
payloads,” a single DEK could approach that. Mitigations, in order: (a) per-epoch
DEKs already cap the payload count per key; (b) derive a per-record subkey via HKDF
from the DEK keyed by record_id so each record effectively has its own key; or (c)
use a deterministic nonce = HKDF(record_id, epoch) counter rather than random. We
pick (a)+(b). Do not ship a single DEK per conversation with random nonces over
unbounded records.
3.4 The actor key: verification + custody
A magic-link enrollment flow, formalized:- Actor triggers key enrollment; Mesh sends a magic link over the already-trusted connector (e.g. a Slack DM to that actor’s verified connector identity). This proves control of the connector identity — the same identity the perspective graph already trusts.
-
On click, the actor’s browser generates the key client-side. Recommended: a
BIP39 24-word mnemonic (24 words = 256 bits, built-in checksum, mature
tooling). The mnemonic derives the actor key; Mesh stores only the public half
/ a wrapping pubkey, never the mnemonic or private key.
Supply-chain caveat: if the keygen JavaScript is served by Mesh, the actor-held-vs-org distinction is only as strong as our web supply chain — a malicious/XSS’d bundle could exfiltrate the mnemonic at generation time, defeating the “even Mesh can’t read it” claim. Actor-held is therefore not marketed as protection against a compromised Mesh; it protects against org-admin over-reach and DB/backup theft. Harden with strict CSP, SRI-pinned bundles,
crypto.getRandomValuesentropy (notMath.random), and — for the paranoid tenant — a reproducible-build / standalone keygen so the browser need not trust a freshly-served script. Entropy source and bundle integrity are a BLOCKER for the actor-held path, not a detail. - The actor is instructed to store the mnemonic safely (it is the only recovery path — see 3.6 for the consequence).
- Mesh wraps existing conversation DEKs to the new actor pubkey (asymmetric wrap, so the actor need not be online for a DEK created while they’re away — a sender can wrap to their published pubkey).
Post-quantum prior art — decide the wrap primitive with harvest-now-decrypt-later in mind. The actor-held wrap protects long-lived secrets: a DEK wrapped to an actor’s X25519 pubkey today is readable by anyone who records the ciphertext now and breaks the discrete log later — the classic harvest-now-decrypt-later exposure that matters precisely for at-rest data meant to stay confidential for years. Signal has already walked this path: PQXDH (spec) is X3DH re-derived to add post-quantum forward secrecy by mixing an ML-KEM shared secret into the agreement, and ML-KEM Braid (spec) extends continuous key agreement with PQ forward secrecy and post-compromise security. The design implication for Mesh is not to adopt either protocol — it is to make the actor-wrap PQ-hybrid from day one if the actor-held path ships: wrap the DEK under a hybrid X25519 + ML-KEM KEM (classical AND post-quantum, so it’s no weaker than today even if ML-KEM is later faulted) rather than retrofitting after a standards shift forces a re-wrap of every stored ciphertext. This is cheap to choose now and expensive to change once actor keys are enrolled. Not a v1 CMEK blocker — a prerequisite for the actor-held fast-follow.
Enrollment protocol + key-derivation must be specified before implementation.
Before the actor-held path ships, the design owes: magic-link single-use + short
TTL + replay rejection + connector-identity provenance binding; a versioned,
domain-separated KDF from the BIP39 seed to the wrapping keypair (BIP39 defines
seed derivation only, not the X25519/HPKE keypair — use the seed as HPKE
DeriveKeyPair ikm with HKDF label domain separation, RFC 9180) with published
test vectors; and rotation/recovery semantics (what access looks like
while a key is being replaced). These are prerequisites for the actor-held path,
not v1 blockers for the CMEK path.
Actor-only mode has no server-side runtime path. In actor-only mode Mesh holds only the actor’s public key and cannot unwrap the DEK at turn time — so the agent cannot read the message and server-side retrieval/dashboard cannot function. Actor-only is therefore not a server-side agent mode: it is only coherent with client-mediated decryption (the actor’s device unwraps and supplies plaintext for the turn) or a confidential-compute enclave the org cannot inspect. v1 does not build either. So actor-only is documented as a future, client-mediated posture, and the default CMEK path is the only one where the hosted agent works normally.
3.5 Threat model, stated honestly (the buyer will ask)
Rogue-operator scope. The “rogue Mesh operator” row is ✅ only for an operator whose reach is limited to the database/backups. An operator who also controls the runtime is the “compromised runtime process” row (❌ for both). These are two different adversaries, listed separately so CMEK protection is not overstated. Metadata is not encrypted. The conversation graph — who talks to whom, when, conversation shapes, participant sets — remains plaintext under both modes and is itself sensitive. That is a stated, accepted limitation, not a hidden one. The actor-only column is mode-dependent. The ✅ for “org admin over-reach” is only true in actor-only mode. Under the recommended default (§3.6: CMEK on, actor-held as an additional wrap), a CMEK wrap of every DEK always exists, so an org admin with KMS access can read employee DMs — the row is ❌ in the default posture and ✅ only when the org fallback wrap is absent. The honest datasheet: if you want employees’ DMs to be unreadable by their own org, you must run actor-only — and accept no legal hold / no recovery. You cannot have both org recoverability and org-blindness; that is a real either/or, not a slider. CMEK scope matches v1 encryption scope. “Mesh encrypts all tenant data under the org key” is scoped in v1 to confidential-shape bodies + confidential-band memories (§3.7). Channel/public content, the metadata graph, search indexes, and logs are not encrypted in v1. The threat table and any buyer claim apply to the confidential subset only.The uncomfortable truth for the datasheet: because the agent must read plaintext to function, no configuration makes Mesh a zero-knowledge system for data it actively processes. What we sell is: at-rest confidentiality, blast-radius reduction, org-vs-actor key custody choice, and cryptographic erasure. We must not market it as “we can never read your messages,” because for any message the agent answered, we could. Overclaiming here is how a security-conscious buyer walks — and how we end up in a breach post-mortem that says we lied.
3.6 Cryptographic erasure & the recovery cliff
- Erasure as a feature: destroy all wrapped copies of a conversation’s DEKs (every epoch, every recipient, including the CMEK wrap) and the ciphertext is permanently unreadable. This is a clean answer to GDPR/CCPA “delete my data.”
Per-actor erasure is a no-op under the CMEK-on default. “Delete one actor’s wrapped-DEK copy removes their access” is only true if that copy was ever their access path. Under the default (CMEK on), turn processing and the dashboard go through the org/CMEK wrap, not the actor’s key — so deleting the actor’s wrapped copy revokes nothing and the plaintext stays org-recoverable. Honest framing: per-actor cryptographic erasure is meaningful only in actor-only mode. In the default mode, “delete this actor’s data” is an application-level delete of the rows plus destroying the conversation’s DEKs outright (which affects all participants) — not a per-actor crypto revocation. Erasure ends at the ciphertext boundary. Destroying DEKs makes stored ciphertext unreadable. It cannot claw back plaintext already delivered to actor devices, sent in model-provider payloads, or landed in logs, exports, caches, or unencrypted backups. The GDPR/CCPA story must pair crypto-erasure with a plaintext-copy retention/deletion policy for each of those sinks; crypto-erasure is a component of deletion, not the whole of it.
- The cliff: if actor-held and the actor loses the mnemonic with no org-CMEK
fallback wrap, that data is gone forever. This must be a loud, opt-in choice.
Default posture: CMEK on, actor-held as an additional wrap, so the org retains
a recovery path unless the tenant explicitly chooses actor-only (and signs the
waiver). This mirrors
E4’s existing “the key is deliberately not in the database / losing it means every ciphertext is opaque” reasoning — we already understand this failure mode for the master key; actor keys make it per-actor.
3.7 What NOT to do in v1 (scope discipline)
- Do not encrypt everything. Encrypt bodies of confidential-shape conversations and confidential-band memories. Channel/public content stays plaintext (it’s public anyway, and searchable). This keeps FTS/pgvector retrieval working for the public corpus and confines the crypto-complexity to the data that needs it.
- Do not build actor-held E2E before CMEK. CMEK is sellable, simpler, and doesn’t fight the product. Actor-held is a fast-follow for the paranoid tenant.
- Searchability caveat to flag now: encrypted bodies cannot be server-side full-text or vector searched. Confidential memory retrieval within an authorized turn works (decrypt-then-rank in-process for that turn’s eligible set), but cross-conversation FTS over confidential content is off the table under E2E. This is a real product limitation, not a bug, and it must be decided deliberately.
Migration & rollout plan
Ordered by dependency and by “smallest reversible step first.” Each is ~1 PR. Phase A — Layer 1 (agent types), no crypto, low riskagent_type+principal_actor_id+principal_stateschema + the binding-aware consistency constraint (migration).- Principal-binding flow (expected-principal-use, compare-and-set);
personalagents engage no one until bound. - Tier-(-1) engagement-gate short-circuit + decision recording (extends the
existing
engagement_disposition/tier/signalcolumns — a type-gatedobserveis explainable like every other verdict). - Acceptance gate: (a) a
personalagentobserves a non-principal’s@mention; (b) distinct actor identities are preserved (no principal confusion across connectors); (c) binding is deterministic and idempotent (compare-and-set, second attempt is a no-op); (d) no cross-agent prompt/memory bleed; (e) deleting one agent does not mutate another. All must pass before rollout.
visibility enum with band; derive default band from conversation shape
at write time. Backfill must snapshot membership at message-observation time,
not the current roster: join against the message-graph membership events so a
churned group gets each historical memory the audience it actually had when
written. Where historical membership is unrecoverable, fail closed — mark the
memory confidential with an empty/uninferrable audience (not eligible) rather
than granting today’s roster access to old secrets. Do NOT blanket-map legacy
perspective→public: the existing revision contract permits operator-authored
entries with no source actor/message, so perspective does not prove channel
provenance. Only independently verified public evidence gets public; authenticated workspace-public channels get workspace scope;
source-less/ambiguous rows fail closed to confidential. conversation→
confidential (audience = observation-time participants).
6. memory_entry_audience table + populate for confidential entries.
7. Retrieval SQL: the bidirectional subset eligibility filter, in both prompt
assembly and recall. This is risks.md#r13, the highest-severity gap.
8. The risks.md#r22 regression net becomes mandatory here, and must cover both
privacy directions, not just leak-outward: (a) plant a DM memory, assert it is
filtered from a third actor’s turn (outward); (b) plant a memory in a hostile
public channel, assert it is not elevated to instruction in a principal DM and
is excluded when the down-scope knob is on (inward); (c) current-revision-only
retrieval; (d) subject_actor_id changes ranking but never eligibility; (e)
per-actor memory-write quota is enforced. The “visibility violations = exactly 0,
cannot be averaged” metric from f2 applies to both directions.
9. Operator API: metadata/body split + operator⇄actor linkage (couples with R16 /
Track D6 role work).
Phase C — Layer 3 (encryption), tenant-gated, opt-in
10. Per-audience-epoch DEK + envelope wrap under org KMS (CMEK). Reuse
internal/secrets. Persist an epoch id with each DEK and scope wraps to the
actors of that epoch (§3.3a) — not a single per-conversation DEK, which would
let late-joiners read prior history and break future-only revocation. Encrypt
confidential bodies + confidential memory bodies, with AAD binding + per-record
HKDF subkeys.
11. Dual-read/dual-write window on key_version (E4 already wants this for the master
key) so encryption rolls out online, not as an outage.
12. Rehearsed key-rotation + restore drill (E4 checkboxes) — extended to per-conv
DEKs.
13. Actor-held wrap: magic-link enrollment over the trusted connector, BIP39
client-side keygen, asymmetric DEK wrap to actor pubkeys, cryptographic-erasure
delete path.
Do not start Phase C before Phase B’s regression net is green. Encrypting data
whose authorization boundary is still leaky just makes a leak harder to detect.
Open decisions
Each has a recommendation; resolving these dials in the design before any code.- Naming:
openvssharedvsgeneralfor the third agent type? (Recommendation:open.) assistantprivilege semantics: beyond “sees its own confidential memory about the principal,” does the principal get any authority over the agent that a non-principal doesn’t (e.g. can only the principal edit persona / rotate keys)? (Recommendation: yes.)- Default encryption posture for hosted tenants: CMEK-on-by-default, or opt-in per tenant? (Recommendation: on-by-default for hosted, given the “trusting us with their data” framing.)
- Actor-only (no org fallback) — do we even offer it? It’s the strongest privacy story and the sharpest support/data-loss knife. (Recommendation: offer it, behind a signed waiver, not as default.)
- Confidential-content search: accept that confidential content is not cross-conversation searchable under E2E? (Recommendation: yes; flag it in the datasheet.)
workspaceband: define its re-emission semantics for v1, or ship only{public, confidential}and addworkspacelater? (Recommendation: drop it from v1 — an underspecified band in a shipped enum is a latent bug.)- RBAC / SSO / SCIM: is enterprise-grade role management + directory sync a GA gate, or a fast-follow? (Recommendation: GA gate for the hosted enterprise tier — two roles won’t survive a real procurement review.)
- Offboarding runbook: on employee exit, do we deactivate their
personalagent, revoke their epoch wraps, and retain org-recoverable data for legal hold? (Recommendation: yes to all three; needs sign-off on retention.) - Operator access audit: commit to an immutable audit event on every operator body-read now (couples with §2.5), or defer? (Recommendation: now — it’s cheap to add at the same time as the metadata/body split and expensive to retrofit.)
- Actor-wrap crypto agility: if the actor-held path ships, is the DEK wrap PQ-hybrid (X25519 + ML-KEM) from the first enrolled key, or classical-only with a later migration? (Recommendation: PQ-hybrid from day one — at-rest secrets are the canonical harvest-now-decrypt-later target, hybrid is no weaker than classical today, and retrofitting means re-wrapping every stored ciphertext. See §3.4.)