Skip to main content

Agent provisioning

Agents are created at runtime, through the web UI, into the database. There is no other way to create one, and that is deliberate — see the blank slate principle in philosophy.md. This document describes the provisioning contract and the setup flow that implements it.

Zero agents is a valid state

A fresh Mesh install has no agents. That state is fully supported:
  • the process starts and stays up
  • migrations run
  • /health reports healthy
  • the log says, once, that no agents are configured
  • inbound connector requests are rejected with an opaque 401
The rejection is opaque on purpose. An unconfigured Mesh should not tell an unauthenticated caller whether it has agents, which workspaces it serves, or whether a given app id is known to it. “No agent” and “wrong signature” look identical from the outside. Nothing about the empty state is degraded mode. It is the intended first boot, and it is what the setup flow is designed around.

What lives in the environment, and what does not

Environment variables configure the process. They never configure an agent. In the environment:
  • DATABASE_URL
  • HTTP_ADDR
  • PUBLIC_URL
  • MESH_ENV
  • OPENROUTER_API_KEY (legacy deployment-wide fallback only; newly provisioned agents store their own provider key)
  • OTEL_* (traces and logs; see telemetry.md)
  • the master encryption key
In the database:
  • agents — name, slug, status, model roles
  • connectors attached to each agent
  • every connector credential, encrypted
The master key is the one credential that cannot live in the database, and the reason is worth being explicit about: it is the key that decrypts the secrets table. Storing it alongside the ciphertext it protects would make the encryption decorative. It has to be supplied from outside — process environment, a mounted file, or a KMS the process can call — and it is the one secret an operator is responsible for holding. Everything else is runtime state. A credential in an env var is a credential that requires a redeploy to rotate, cannot be scoped to one connector, and cannot be audited per agent.

Credentials belong to a connector, not to an agent

Each agent has connectors; each connector owns its own credentials. For Slack that means a bot token and a signing secret per connector row, encrypted with AES-GCM in the secrets table and owned by that connector. The consequence is that one agent can be installed in many Slack workspaces at once. Each installation is its own connector with its own token and its own signing secret, and inbound events route on the pair (api_app_id, team_id). The agent is one identity in Mesh with several places it shows up, which is the same relationship a human employee has to several Slack workspaces. Routing looks up the connector by that pair before verifying the signature, because the signing secret is per-connector — there is nothing to verify with until the lookup has happened. Lookup then verify, and a failed lookup returns the same opaque 401 as a failed verification.

Setup flow

The first-boot wizard is the intended path to a working install. The backend for it lives in internal/admin (a JSON API under /api/); the web UI drives it. The numbered steps below are API steps — one per provisioning decision the backend records. They do not correspond 1:1 to the numbers the stepper shows a user, which are user-visible phases: the UI collapses the provider choice, the API key, and the model roles under one “Model” knot, and the hostname checkpoint, the Slack manifest, and the credentials under one “Slack” knot. So “step 9” here is a contract reference, not something to match against the sidebar.
  1. Create the administrator account. First boot with no users lands on setup. The administrator is a row in users — an email plus a bcrypt password hash, a one-way hash with its salt embedded in the output. Verification re-hashes the candidate with the stored parameters and compares the result; the plaintext is never stored and never compared as a string. This account is the human’s login for the web UI, and the wizard says so explicitly — the agent created next is a different thing with its own identity, and conflating the two is the first confusion the copy must prevent. Subsequent visits authenticate against this account through a login page; sessions are database-backed cookies (see 0006_users_sessions.sql), and this first account — an owner — invites the rest of the team from Settings → Operators: a one-time link per person that sets their password (0043_operators.sql), with a role (0050_operator_roles.sql): another owner, or an agent operator who is given specific agents and sees nothing else, so they survive restarts — and are already multi-replica-correct for the day the rest of the process is (see step 9).
  2. Name the agent. Supply its display name and slug. Mesh saves a draft immediately; the draft URL resumes the wizard after a reload.
  3. Connect a model provider. OpenRouter and Ramp Router are supported. Mesh validates the key and stores it encrypted as an agent-owned credential, so each agent has its own billing identity across its connector installations.
  4. Pick the models. Primary, fast and heartbeat roles use the chosen provider’s model catalog. OpenRouter’s primary defaults to openrouter/auto; unset fast and heartbeat roles use the primary model.
  5. Create the profile and picture. Write the persona or generate an editable draft from the agent’s purpose. Saving creates an append-only persona revision, bounded at 16 KiB. Upload a PNG/JPEG, keep the geometric default, or generate an avatar using an image model from OpenRouter’s live image catalog. Ramp Router supports profile generation; image generation is not enabled for that provider.
  6. Choose a connector. Slack and Telegram are available. Other displayed connector options remain disabled.
  7. Connect Slack automatically or manually. After an operator connects a workspace configuration refresh token in Settings, Create Slack app creates the app, persists its client/signing credentials, configures a signed per-app events URL and OAuth permissions, and attempts to upload the avatar. Install in Slack opens Slack’s approval flow. The callback exchanges the authorization code server-side and validates the app, workspace and scopes before saving the bot credential. Progress survives refreshes and interrupted setup; an uncertain creation is never repeated blindly. Connect an existing Slack app manually retains the hostname checkpoint, generated manifest and App ID/signing-secret/bot-token form. Both paths reuse the same connector persistence and runtime activation. Neither uses Socket Mode or requires an app-level socket token. Details, failure recovery, and the remaining live Slack validation are in Automatic Slack setup.
  8. Activate and test the conversation. Mesh saves encrypted credentials and reloads the running registry. Open the agent in Slack, send a test message, and invite it to the channels where it should participate. A saved installation can retry activation without obtaining another OAuth code.
Repeat from step 2 for each additional agent. Mesh still assumes one serving replica: registry reload updates the process that handled setup. Multi-replica registry propagation is separate work; database-backed sessions and serialized provisioning claims do not remove that constraint.

Set your hostname before you register the Slack app

For automatically provisioned apps, use Update Slack configuration on each managed connection after a hostname change. The following manual repair guidance applies to apps created through the manifest-and-credentials option. PUBLIC_URL is consumed in exactly one place: the manifest’s settings.event_subscriptions.request_url, which is PUBLIC_URL + /slack/events. Slack persists that string on the app. It does not hold a reference to this deployment that follows it around — it holds an address, and it posts events to that address until someone edits it. So the ordering is not a preference. If you intend to serve Mesh on a custom domain:
  1. point the DNS record at the deployment and terminate TLS in front of it
  2. restart Mesh with PUBLIC_URL set to the final https:// address
  3. then generate the manifest and create the Slack app
Do it the other way around and you own a repair job. The Slack app keeps posting to the hostname it was created with; Mesh, now answering on a different one, never sees the events. Nothing raises an error — not in Mesh, which is simply not being called, and not in Slack, which is delivering successfully to an address that still resolves. The agent goes quiet, and the only signal is that it stopped replying. When the old hostname stops resolving entirely, Slack’s delivery retries fail and the app’s event subscription can be disabled outright. The repair is small once you know it is needed: open the app’s Event Subscriptions page, replace the request URL with the current one, and let Slack re-verify. The installed bot token and signing secret are unaffected, so nothing needs re-installing and nothing needs re-pasting into Mesh. The agent detail page’s Slack app configuration section renders the request URL this instance currently expects, along with a link to each connected app’s Event Subscriptions page, precisely so this can be checked without deriving the URL by hand. One related constraint the wizard also surfaces at that checkpoint: Slack posts from its own network and requires a publicly reachable HTTPS URL. A PUBLIC_URL on localhost, a 127.0.0.0/8 address, an RFC1918 address, or plain http:// cannot receive events at all — url_verification fails and no message ever arrives. That is fine for local development, where nothing is expected to reach Slack; it is not a configuration any real install can keep.

Why the manifest matters

The worst part of self-hosting anything that talks to Slack is the app setup: scopes, event subscriptions, a Request URL that has to be reachable before Slack will accept it, and a dozen toggles across several screens. Every one is a place to get it subtly wrong and get a silent failure back. A manifest collapses all of it. Scopes and event subscriptions are declared in the document, so the user’s job is reduced to two actions — paste the manifest, paste back three values. For an open-source project, that path is the first impression, and it is worth more than most features. One sequencing wrinkle the wizard’s copy handles: Slack may probe the Request URL before this agent’s credentials are saved, and an unconfigured route answers the standard opaque 401. That is fine — verification succeeds on retry once step 9 completes, and the wizard tells the user so.

Adding more agents

Provisioning is the same flow every time. The second agent is created exactly like the first, because the first was never special:
  • one Mesh deployment
  • many agents, each with its own identity, credentials, connectors, and models
  • each one appearing in Slack as its own coworker
Observers in a channel see colleagues, not personas of one bot. That is the point of the multi-agent runtime, and it is only true because no agent is configured differently from the others.

Continuity with the MVP

This is not new scope. roadmap.md (repository) has listed a “simple agent provisioning flow” and a “basic web UI for setup and inspection” under must-have since it was written, and its demo script starts with “create an agent, connect Slack.” The env-var bootstrap that briefly existed was the deviation from that plan, not the plan. Removing it restores the original design.

Repairing permissions on a managed Slack app

Open the agent’s managed Slack setup. Installed Slack permissions checks the existing bot token using Slack’s auth.test identity and X-OAuth-Scopes response header. Missing groups:read or files:read, for example, appears as an explicit approval requirement. A failed lookup is not reported as a healthy installation. Use Check permissions to repeat the live check; checks also run when this view opens. This is an operator-driven check, not a background reconciliation job. Choose Update Slack permissions. Mesh exports the existing app manifest, adds the shared required bot scopes while preserving unrelated settings, and opens OAuth approval for that same app and workspace. A person authorized by Slack must approve; updating a manifest alone cannot change a token’s grants. If the management grant was revoked or expired, reconnect Slack provisioning in instance settings. Cancelling OAuth leaves current credentials in place. Success validates the workspace, app, bot identity and required scopes before saving credentials, then uses the existing activation/reload path. Activation failures retain the existing retry flow. After approval, verify a fresh message and upload in the affected channel. Scope verification does not prove channel membership or restore historical uploads whose original audience evidence was unknown. This flow does not modify historical access evidence or borrow another agent’s bot credentials. Manually connected apps still use their existing manual setup path.