Skip to main content

Tools

Strategy update (2026-09-18): The tool integration strategy defines the planned shared catalog, named connections, progressive discovery, and Add capability flow. The delivery plan tracks implementation. Descriptions of existing behavior below remain the current contract until those changes ship. A tool is a granted capability with an audit trail, not a function pointer. Mesh already has the two hard boundaries a tool system needs. The agent harness (agent-harness.md) makes the loop that calls tools durable, budgeted, and attributable. The execution layer (sandboxed-execution.md) makes the place exec-class tools run isolated, credential-scoped, and pluggable. This document is the layer between them: what a tool is, how tools are classified, registered, and gated, where each class actually executes, and which cloud services implement the classes that need infrastructure. One thing is deliberately settled before the taxonomy, because it shapes all of it:
Mesh is cloud-first. The default assumption of every popular agent harness — that tools run on the machine the harness runs on, in the user’s own filesystem, browser, and network — is exactly the assumption Mesh rejects. A Mesh agent’s shell is a rented sandbox, its browser is a rented browser session, and its web access goes through declared providers. Nothing about an agent’s reach is derived from the host the control plane happens to be deployed on.
That is not only a security posture. It is what makes fifty co-tenanted agents on one small control-plane box possible, and it is what makes every capability a row that can be audited, budgeted, and revoked.

The first tool already exists

The memory verbs shipped with the memory slice: a provider-neutral tool contract in internal/model (Tool, ToolCall, RoleTool), OpenRouter tool-call support, bounded tool rounds, and fail-closed validation of everything the model asks for. That slice proved the wire protocol. There are four of them — remember, recall, amend, forget — and the split is itself a contract rather than a convenience. remember creates and cannot modify; correction is amend; retirement is forget. A verb that cannot supersede cannot produce a result that honestly describes a supersession, which is half of the truthful-confirmation contract in memory-retrieval.md. The other half is that every reply following a tool round is handed an authoritative ledger of what actually committed, built from the store’s return values rather than the model’s requests. Note also what these tools deliberately do NOT accept: any parameter naming the source of a memory. Provenance is derived from the persisted inbound message’s graph record, so a model has no channel through which to attribute a claim to someone else. The first generalization is now built in internal/tool: an immutable registry binds provider-visible schemas to effect declarations, validation, and handlers; one executor validates and invokes calls against an already policy-scoped registry; and the four memory verbs use that path rather than a turn-loop switch. Tool calls and typed observations remain persisted steps of a durable Run. The round count remains capped at three and enforced structurally: on the last permitted round the tool list is removed from the request. That is enough for recall → amend → answer and short of an autonomous loop.

Effect classes

Tools are classified by effect surface — what a call can touch — because that is the property that determines where the call runs, what has to be provisioned for it, how cancellation treats it, and how suspicious the default policy should be. Four classes:
query tools read: memory recall, conversation-graph lookups (“what did we decide in this thread last week”), and web search or page-fetch through a declared hosted provider. A remote MCP tool the operator has classified as Read at review is a query tool too (docs/runtime/remote-mcp.md): the operator’s classification, not the server’s hint, is what places it here. Query tools run in the control plane. That is not an exception to the rule against executing tools in the harness process — that rule exists for exec-class calls, and a query tool has no shell, no filesystem, and no arbitrary network reach. The constraint that keeps the class honest: a query tool may only talk to Mesh’s own database (perspective-scoped, per context-resolution.md) or to a provider that was registered and credentialed at runtime. A control-plane tool that fetches an arbitrary URL is not a query tool; it is a server-side request forgery primitive pointed at whatever network the control plane deploys into. Arbitrary URLs are fetched by the browser class or from inside a workspace, where egress policy applies and the blast radius is a rented container, not the control plane. workspace tools are the exec class: run a command, read and write files, clone, build, test. They execute only inside a workspace materialized by the execution layer, one workspace per run, provisioned lazily on first use. sandboxed-execution.md is normative for all of it — the backend contract, credential minting, egress default-deny, compute budgets. Nothing in this document weakens it. browser tools drive a cloud browser session: navigate, read the page, click, type, screenshot. A browser session is a second provisioned resource type, sibling to the workspace, with its own backend contract and its own section below. It is deliberately not modelled as a workspace capability: the lifecycle (a session humans can watch and take over), the budget unit (session minutes), and the vendor landscape are all different. effect tools act on the outside world through an API Mesh holds a credential for: post a Slack message beyond the turn’s own reply, open a GitHub issue, push a commit, send an email. Two things distinguish the class. First, every effect call is an outbound event, not a side effect — same persistence, idempotency, and delivery-state machinery as replies (turn-pipeline.md). Second, the class interacts with the commit boundary: an effect call is what moves a Run from buffering to committed in the interrupt model, so it is the class whose calls must be recorded before they run and never silently re-run after a resume. The class is declared on the tool, not inferred from its behavior, and the declaration is load-bearing: a tool whose implementation reaches outside its declared class is a bug of the same severity as a backend that ignores egress policy.

Authority is the agent’s account, not the harness

Every popular harness puts a human approval prompt in front of effectful calls, and it is worth being exact about why: those harnesses act with the user’s own credentials. The agent is borrowing a person’s GitHub login, so of course a person has to be asked before it pushes. The approval gate is a patch over borrowed identity. Mesh rejects the premise. A Mesh agent is provisioned like a new coworker: it has its own GitHub account, its own Linear seat, its own Slack identity, and the operator grants those accounts exactly the permissions the role needs — write on these repositories and not those, member and not admin. Every tool call is made with the agent’s own credential, stored encrypted and scoped to that agent (../agent-provisioning.md, migration 0005). What the agent may do is therefore decided in the external system, by the permissions attached to the agent’s account, and Mesh’s job is to make every action attributable and reconstructable afterward — not to interpose a human on each one. Three consequences follow, and the rest of this document is written against them:
  • There is no per-call approval gate, and none is planned. An agent that can open a pull request opens it, the same way an intern with write access does, and the review happens where reviews happen: on the pull request. The policy column below is enabled | disabled; there is no requires_approval value. If a deployment ever wants a human decision in the loop for a specific workflow, awaiting_input and the WorkGraph human_gate node exist for decisions the agent needs (“which of these two fixes?”), and that is a different thing from a permission check.
  • Scoping happens at the account. The operator who wants an agent unable to delete repositories does not configure Mesh; they do not grant the agent’s GitHub account that permission. Mesh’s per-agent enablement and per-connector offer gating decide which capabilities are in the prompt, which is a question of relevance and blast radius, not of authorization.
  • The residual risk is prompt injection, and the answer is still scoping. A message in a channel can try to talk an agent into an action. An intern can be socially engineered too; the reason that is survivable is that the intern’s access was bounded before the conversation started. Structural separation of untrusted content (roadmap F3) narrows the attack; account permissions bound the damage; the ledger makes it explainable. A yes/no prompt to a human who cannot see the injected text does none of those.
The credential is the agent’s. Whether the transport that carries it is a first-party Go handler, a CLI inside the workspace, or a remote MCP server is the next section’s question and is orthogonal to this one.

Three tiers: which transport a capability uses

Mesh has three ways to give an agent a capability. Choose by runtime coupling and who maintains the implementation. Acting service identity is explicit in every tier; infrastructure payer credentials can have instance defaults. The planned shared integration layer unifies setup and discovery across these paths without collapsing their execution boundaries. Tier 1 is the smallest set and should stay that way: workspace_exec, workspace_read_file, workspace_write_file, and checkout_github_repo are first-party because each does something an external server cannot — mint a workspace, install a credential helper, bind a run to a checkout. A first-party package must not be written for a third-party API that has no Mesh-primitive coupling, unless Mesh deliberately owns a small provider-neutral capability contract with multiple adapters, as it does for web search and reading. Tiers 2 and 3 cover ordinary vendor APIs; dogfooding alone does not justify a wrapper. Tier 2 is the path the GitHub workflow already takes: the checkout installs a checksum-pinned gh, the workspace holds GH_TOKEN for the agent’s GitHub account, and the model is told to gh pr create. It is also the tier the local-first harnesses converged on, for a reason that holds here too: a CLI’s schema is not in the request, whereas offered MCP schemas consume context on model calls. Today Mesh offers the operator-approved subset, which can still be large. Planned progressive activation will load relevant native and MCP schemas on demand. The CLI is not free — the skill that teaches it (below) is rendered into the prompt — but a skill is a short procedure loaded when the work is code, not a per-request catalog of every callable the vendor exposes. Mesh’s version is stricter than theirs, because the CLI runs in a rented sandbox holding one scoped credential rather than on the operator’s machine with everything on it. The authority boundary of a tier-2 call is exactly “what the workspace’s credentials can do” — which is the account-scoping rule above, applied mechanically. What Mesh gives up is per-call structure: a gh pr create inside workspace_exec is one effectful exec to the ledger, not a typed create_pull_request step. That is why the commit boundary is crossed at the class level (any effectful workspace or effect call) rather than per tool, and why a workflow that needs a typed observation of a specific external effect is a tier-1 candidate. Tier 3 is the third-party surface, specified under Registration and gating. The credential is the agent’s own login to that service — an operator signs in to the vendor’s MCP server as the agent once — and it is stored per agent, as an agent-owned secret exactly like the agent’s GitHub or Linear key, never on the provider row. That placement is load-bearing: a tool_provider describes a server (endpoint), and several agents may enable tools from the same server, so a credential on the provider would make every one of those agents act as the same external account — which is the borrowed-identity failure this whole section exists to rule out. The binding that carries the credential is (agent, provider), as are reviewed tool fingerprints. Current connections are agent-wide across conversations; see Remote MCP tools. The target named-connection model retains single-agent ownership of every acting account, as required by the identity guardrail (repository), and keeps credentials off provider catalog definitions. Linear is the first consumer: its remote MCP server replaces the first-party Linear package once it reaches parity, and the package is retired rather than maintained alongside. Knowledge about how to use a tier-2 CLI belongs in a skill — a SKILL.md in the AgentSkills shape, shipped inside the owning package and rendered into the prompt as untrusted content, the way inbound bodies are — not in Go string literals. Skills are curated into the repository; they are not pulled from a public registry at runtime, because a registry of prompt text is a supply chain. Status: not built. Today the gh workflow is taught by internal/prompt.codeWorkSection, hand-written Go rendered from the offered tool set, and that stays the mechanism until a skill loader exists. The migration is one scheduled item: a package-level SKILL.md crawled with the manifest, rendered where codeWorkSection renders today, with the same “mention only offered tools” rule; codeWorkSection is deleted in the same change, not before.

The tool contract

The provider wire schema stays deliberately small in internal/model.Tool. An executable definition in internal/tool binds that schema to Mesh-only policy and behavior:
Replay is per-tool, not per-class. Most query and workspace reads are replayable; workspace_exec and workspace_write_file are effectful; every effect call is effectful. The harness already commits a replayability classification on each RunStep — this field is where it comes from. Replay is also what the commit boundary reads, and the rule is deliberately the conservative one: an effectful call of any non-query class crosses the run’s commit boundary before it runs, including a workspace_exec that only ran the tests. This is the rule #216 implements in turn.Engine.executeRecordedTool (run.ToolEffect is the decision, RecordCommit the write-ahead stamp); before it merges only the reply crosses, and this paragraph describes the contract, not main. Mesh cannot see inside an exec — the same command line that runs a test suite can git push — so “effectful with respect to its own workspace but replayable with respect to the world” is a distinction the harness cannot verify and therefore does not make. The cost is that a run interrupted after a test-only exec is resumed by reporting rather than re-executing; that is the right cost, because the workspace is torn down at run end anyway, so a resumed coding run is lossy regardless. A tool that wants to be re-executed on resume declares replayable, which is a promise about the world, not about the workspace. The one exception is Internal: a Mesh-owned effectful verb (the memory verbs) writes only to Mesh’s own Postgres and stays inside the boundary (see interrupt-model.md, “The enumerated effect surface”). A handler error fails the turn unless the handler tags it. Two tags turn it into an observation the model sees on its next completion: meshTool.ErrInvalidToolInput (the arguments were wrong in a way the model can fix) and, for query tools, meshTool.ErrToolUnavailable. The default stays fatal so an infrastructure failure (an unreachable store, a ledger write) is never mistaken for bad input. Argument decoding is the one input check every first-party handler shares, so it goes through one decoder: meshTool.DecodeArguments[T] requires a top-level JSON object (a bare null would otherwise run the tool on its defaults) and rejects unknown fields and anything after the object (the crawl’s shape validator canonicalizes but does not enforce additionalProperties: false, so the handler is where that is enforced), and every refusal it returns carries ErrInvalidToolInput with a message written for the model. An unknown field is named along with the accepted top-level fields; a wrong type names the field and the expected JSON kind. Messages name fields, never argument values. Before this, handlers returned the bare decoder error, and a model that invented label on file_linear_issue lost the whole turn to json: unknown field "label" instead of retrying without it. Needs is how provisioning stays lazy and legible. A run that never invokes a tool with Needs: ["workspace"] never provisions a workspace. A tool whose needs include a named secret is invocable only when the run’s minted grant set covers it. The declaration is also the answer to “what could this run have reached,” which must stay a query, not an investigation. An Observation is typed and persisted, never a bare string blob:
The contract exists today. Artifact storage and the durable ArtifactID link are still follow-on work; an executor refuses to label an observation truncated without an artifact id, and the run repository currently refuses artifact-bearing observations rather than silently losing the link. Once artifact persistence lands, large outputs — a build log, a page dump — are stored as artifacts and truncated into model context, not the other way around. The artifact row is what the UI shows and what a WorkGraph edge can carry; the context excerpt is a view of it.

The tool package format

This is the declared forward path for how first-party tools are structured on disk. Its schema is deliberately forward-compatible toward a future installable package — but that path is a kept-open option, not a commitment; remote MCP is Mesh’s third-party extension surface today (see “One third-party boundary” below). It is not a detour from the contract above — it is that contract’s on-disk shape. The Definition above is what a tool is to the runtime. This section is how tools are packaged so that a human can see them, an operator can enable them, and — in a future state — a third party can ship them, all without the control plane having to execute a line of the code to learn what they are.

A folder is a package, not a tool

The unit on disk is an integration package, and it is deliberately not the same unit as a Definition. A Definition models exactly one provider-visible callable: one name, one input schema, one Class, one Replay policy, one Needs set, one handler. But a real integration exposes many callables of differing shape — the Linear package reads issues (a query) and creates them (an effect); the GitHub package reads PRs, writes comments, and can grant workspace access. Collapsing those onto a single package-level class would either over-approve the reads or force every operation through one coarse compound tool. So the format splits two layers explicitly:
  • Package — folder-level metadata, shared config surface, and lifecycle hooks. Identity, versioning, the credential(s) the whole install needs, and the enable/disable/health lifecycle live here.
  • Tools — one or more [[tools]] entries, each of which becomes exactly one Definition: its own name, description, parameters schema, class, replay, needs, and handler binding.
The registry keeps registering callables, not packages — NewRegistry is unchanged, it just receives the Definitions the crawler built from every [[tools]] entry across every package. Every enabled package lives under tools/ with its own directory carrying a declarative manifest plus its compiled-in handlers:
The seed set is the tools we already run on ourselves — Linear, GitHub, Sentry, DeepSource, Datadog — chosen because dogfooding is the cheapest correctness test, and because they cover the shapes the format has to survive: an issue tracker, a code host, two observability providers, and a code-quality gate. The long tail is explicitly anticipated: someone else runs Jira, Trello, GitLab, Gitea; someone else has PagerDuty instead of Datadog. The format is designed for dozens to hundreds of these, so nothing about adding the eleventh package may require editing a central switch statement — a package is present because its folder is present.

The manifest

manifest.toml is the declarative half of the package. TOML is the chosen format — it is typed, it takes comments, humans hand-edit it, and it has none of YAML’s indentation or implicit-typing footguns for a file operators will read and write. The manifest carries package identity and lifecycle, the shared configuration surface an operator fills in, and one [[tools]] block per callable — each block carrying exactly the declarative fields a Definition requires, including the input schema:
Every [[tools]] entry carries its own input schema. This is load-bearing: normalizeDefinition rejects a Definition whose model.Tool.Parameters is absent or is not a JSON object, so a crawler could not produce a provider-visible tool at all without this field. The schema is a JSON Schema (draft 2020-12) document, given inline or via a bounded relative file reference under the package folder (no traversal, size-capped — see Versioning). Keeping the schema in Go instead would silently break the “identical manifest for future third-party packages” claim, because a third-party loader has no Go to read it from. An optional observation output schema is encouraged: it makes the future RPC/WASM boundary far less ambiguous about what a handler is allowed to return.

Config disposition is a closed vocabulary, not three loose booleans

The earlier three-boolean form (secret / encrypt / visible_to_user) admitted contradictory and unsafe states — secret=true, encrypt=false, or secret=true, visible_to_user=true — and never said which actor a value was visible to. It is replaced by a single closed disposition enum with a strict truth table the server defines and enforces: The invariant secret ⇒ encrypted && never returned is not expressible as an illegal state anymore — there is one field, and each value carries its whole disposition. Unknown disposition values fail the manifest closed. Where each disposition is stored, today: secret values go to secrets (encrypted, owner-scoped, captured write-only through the tools tab); config and public values go to agent_tool_config (0035), one row per (agent, package, key), written from the same tools tab and validated by internal/toolconfig against the manifest that declares them — unknown key, secret key, or wrong type is refused at the boundary. The turn engine reads an agent’s values once per turn and hands all of them to the package’s handlers (effects.Runtime.ConfigValue) and the public ones to the prompt, rendered as the agent’s own tool settings. That is the difference the two non-secret dispositions make: a public default repository is something the agent can name and act on; a config value reaches its handlers and nothing else. Need mapping is explicit and resolves exactly once. A secret field declares the stable secret-kind it satisfies via need, and a tool’s needs = ["secret:linear_api_key"] refers to that same kind. The crawler verifies that every secret:<kind> Need names a kind provided by exactly one secret config field in the package — no dangling Needs, no ambiguous double-provides. Secret ownership is derived by Mesh, not declared by the manifest. The manifest says a secret is needed and of what kind; it does not get to say who owns the stored credential. Ownership is resolved from the concrete agent-tool/provider installation at capture time. Migration 0005_connector_scoped_secrets is the precedent and the reason: install credentials are owned by owner_type='agent_connector' (owner_id = agent_connectors.id), not blanket owner_type='agent', precisely because one agent can hold many provider installs and a per-agent unique constraint made the second install unstorable — and worse, invited cross-install credential confusion. A manifest that could name its own owner would reintroduce exactly that bug. So the manifest supplies kind; the installation supplies owner.

The invariant: the manifest is crawlable without executing tool code

This is the load-bearing rule of the entire format, and it is stated here because everything else depends on it:
A manifest MUST be fully readable — parsed, validated, and rendered — without executing any of the package’s code. Discovering what a package is, what callables it exposes, what class each declares, what secrets it needs, and what its config surface looks like is a pure read of a data file. This invariant governs discovery and rendering only. It does not claim that code runs exclusively on Definition.Execute — lifecycle hooks (on_enable and friends) are real, effectful code that runs at enable/disable/capture time. Those are not discovery; they are the durable lifecycle defined in “Lifecycle” below, and they run only after an explicit operator action, never during a crawl.
Data you can read is safe; code you have to run is not. Keeping those two apart is what buys three properties at once, for free:
  1. The UI renders from manifests. Startup crawls tools/, parses each manifest.toml, and the operator surface (Track D) is a projection of that set — every offerable tool, its class, and its config fields, with zero tool-specific UI code and zero tool code executed to draw the screen.
  2. Enablement and secret capture happen before the tool ever runs. C1.5’s flow — “enabling a tool whose Needs declares a secret blocks until the secret is captured” — is only possible because the requirement is readable in advance. If learning a tool’s needs required running it, secret-gated enablement would be a chicken-and-egg problem.
  3. Third-party plugins become safe to catalog. The moment code runs just to describe itself, importing an untrusted plugin means executing untrusted code to render a settings page. The invariant is the security boundary that makes a future plugin ecosystem tractable: Mesh can read, display, and reason about a plugin it has not yet decided to trust with execution.

Startup crawl

A directory crawl on its own cannot discover a compiled Go handler just because tool.go sits next to the TOML — a Go package only enters the binary when something imports it, so “the folder is present” does not make its code linked and callable. Left implicit, that gap collapses back into a central handwritten list, init() side effects, or some unspecified binding magic — the very switch statement this format exists to delete. So binding is a two-phase contract, stated explicitly:
  1. Build time — generation. A code generator walks the in-tree tools/ packages and emits a deterministic compiled-handler catalog: a generated map from (package_id, handler_name) to the actual Go function, compiled into the binary. The generator also embeds and attests the first-party manifests, so what the binary ships with is fixed at build time, not re-read from an arbitrary disk at boot.
  2. Startup — parse + verify, never execute. The runtime parses every (embedded, first-party) manifest without executing a handler, then verifies a 1:1 match between each [[tools]] handler reference and an entry in the generated catalog, and between each declared package hook and its bound function. It then builds the Definitions and hands them to NewRegistry.
Mismatch behavior is defined, not left to chance:
  • Missing binding (a handler with no catalog entry), extra binding (a catalog entry no manifest references), or duplicate (name collision within or across packages, which NewRegistry already rejects) is a hard error with an operator-visible diagnostic naming the package, the tool, and the handler.
  • For an in-tree first-party built-in, an invalid or unmatched manifest fails build / start / readiness — a shipped capability silently vanishing is worse than a loud failure, so we fail loud.
  • Quarantine-with-process-survival (skip the one bad package, keep serving the rest) is reserved for the future third-party path, where one untrusted package must never be able to take down the process. First-party fails closed at the process; third-party fails closed at the package.

Lifecycle

The manifest declares which lifecycle junctures a package participates in; the runtime binds each declared hook to a compiled-in Go handler. The hook set: (The per-call Validate → Execute path is not a package hook — it is the Definition contract, one per [[tools]] entry, unchanged from the top of this doc.) These hooks are external effects and inherit the effect guarantees. on_enable/on_disable provision or tear down webhooks, revoke sessions, mint provider state — exactly the kind of external side effect the document already requires be persistent, idempotent, and delivery-guaranteed for tool calls. A naïve “just call the hook” would let a crash after the provider succeeds but before local state commits duplicate a webhook on retry, or leave one live after a disable. So lifecycle runs through a durable lifecycle executor, distinct from the per-call executor but built on the same discipline:
  • Attributed transition records. Every attempt writes a durable record: who requested it, which package/install, from-state → to-state, pinned to the exact config version and secret version in force at attempt time.
  • Stable idempotency keys. Each provider-facing action carries a key derived from (install, transition, config+secret version), so a retry after an ambiguous crash reconciles instead of duplicating.
  • Typed outcomes + retry classification. Success / retriable / terminal are explicit; deadlines bound each attempt; retriable failures back off, terminal failures surface to the operator.
  • State machine. Enable moves disabled → enabling → enabled | error, and the package is not offerable until it reaches enabled. Disable gates new calls immediately, then reconciles cleanup asynchronously; a failed teardown parks in error with a compensation task, it does not silently strand provider state.
  • Serialization. Concurrent enable / disable / rotation / health-check on the same install are serialized; in-flight calls are drained or cancelled per a declared policy before a disable completes. Rotation re-runs on_secret_capture against the new secret version before cutting over.

Out of the gates: compiled-in; future: the same manifest, a different loader

An honest constraint has to be stated plainly, because it shapes the phasing: Go has no safe way to hot-load a zip of native tool code. The plugin package is fragile, platform-bound, and version-locked to the host binary, and a downloaded shared object is arbitrary in-process code with the control plane’s full authority — exactly the ambient-authority pattern Mesh exists to reject. So the phasing is deliberate and the architecture is stable across it:
  • Out of the gates, every tool’s hooks are compiled-in Go, one in-tree package per tools/<name>/ folder, registered by the startup crawl. The manifest is the declarative sidecar; the code ships in the Mesh binary. This is first-party only — we are not accepting third-party plugins yet.
  • Future state, third parties ship a tool as a zipped folder containing the identical manifest.toml plus code, and Mesh runs that code behind a different loading mechanism — an out-of-process boundary (subprocess/RPC or WASM), never dynamically linked into the control plane.
The folder-and-manifest architecture is identical on day one and at hundreds of plugins. Only the code-loading boundary changes, and it changes outward — from compiled-in, to sandboxed out-of-process — which is the direction that adds safety, not removes it.

Versioning and trust boundaries

“Forward-compatible” is a claim that needs mechanism, not just intent — a future loader can only safely read today’s file if the compatibility and trust boundaries are pinned now. The schema therefore fixes:
  • manifest_version — the schema version of the manifest file itself. Present from v1 (shown above). A reader that does not understand a manifest_version refuses the package rather than guessing.
  • Immutable package identity + version — package.id is globally unique and stable; package.version is semver. Grants and audit records pin to (id, version, digest), never to a mutable folder name.
  • Mesh API/ABI range — package.mesh_api declares the runtime API range the package targets; a package outside the running Mesh’s supported range is refused, loud.
  • Hook protocol version — the out-of-process hook/RPC contract carries its own version, independent of manifest_version, so the wire protocol and the file format evolve separately.
  • Strict unknown-field policy — first-party manifests reject unknown fields at build time (typos fail loud). The third-party path uses the same strictness so a field a newer publisher added cannot silently change meaning under an older Mesh.
The third-party archive path additionally requires, before a manifest may influence any grant:
  • Provenance + integrity — a content digest over the package and a publisher signature; Mesh verifies both, and every grant/audit record is pinned to the verified digest.
  • Staged, atomic activation — an archive is validated in a staging area and only atomically activated on success; a failed validation never half-installs. Updates are atomic with rollback and explicit config migration between package versions.
  • Archive hardening — ingestion rejects path traversal, symlinks, oversized manifests, excessive file counts, and decompression bombs.
The installable-code path above and the later normative statement that remote MCP is Mesh’s third-party extension surface must not read as two competing answers to “how does a third party extend Mesh.” They are reconciled explicitly:
  • Remote MCP is the third-party execution surface. A third party who wants Mesh to run their capability attaches it as a remote MCP server — already out-of-process, already “describe-then-trust,” already class-and-policy assigned at registration. That is the boundary an implementer builds to today.
  • The package format is the first-party structure. It is how tools that ship with Mesh are organized on disk, and its schema is deliberately forward-compatible. Forward-compatible is a property, not a committed roadmap item: we are not promising an installable-plugin loader. If we ever ship one, the versioning + trust mechanisms above are its prerequisites — they exist so the option stays open, not because it is scheduled.
So: one unambiguous third-party boundary today — remote MCP. The package format is first-party structure whose schema does not foreclose a future installable path. This is the path we are moving forward on.

Relationship to remote MCP

Remote MCP (below) and the tool package format are not competitors, and the boundary between them is the one drawn just above: MCP is the third-party execution surface; the package format is first-party structure. A remote MCP server is already an out-of-process, declaratively-described capability — its tool list is metadata Mesh reads before it trusts, and its class and policy are operator-assigned at registration. The package format brings that same “describe-then-trust” discipline to the tools that ship with Mesh. Both obey the manifest invariant — Mesh learns what a thing is by reading, and runs it only after an explicit grant — which is exactly why the schema can stay forward-compatible toward a future installable path without that path being a commitment today.

The exec-class tools that exist

The first workspace-class tools are built, on the package format above, and they are the reason the execution layer exists: The shape follows the rule at the top of this document — a tool call and the process that made it are not the same object. The four workspace-class tools execute against a workspace the execution layer materialized for the run (sandboxed-execution.md §What ships today), reached through the per-agent runtime on the context exactly the way the Linear and GitHub handlers reach their API clients. create_github_pull_request is the one effect in the table: it runs in the control plane against GitHub’s API, as the fallback for opening the PR when gh is not available in the workspace. Under the three-tier rule it is a tier-2 capability wearing a tier-1 wrapper, kept only until gh is guaranteed in the workspace image; it is not the pattern to copy for the next GitHub operation. Two decisions in the workspace tools are worth naming:
  • A failing command is content, not an error. workspace_exec returns the exit code and both streams verbatim (clamped, head and tail kept), because the agent iterates on compiler output and test failures. So are the layer’s own degraded states — no backend configured, transport down — which come back as observations naming the fix, because a Go error from a handler fails the turn and “your operator has not configured a sandbox” is something the agent should be able to say.
  • checkout_github_repo is the GitHub package’s workspace grant. It proves the credential and reads the default branch from the control plane first, then clones into the workspace with git configured to authenticate from the GH_TOKEN the workspace was granted and to commit as the agent’s GitHub account. The workflow it sets up — edit with workspace_write_file, test with workspace_exec, commit and push with git, open the PR with gh pr create or create_github_pull_request — is what the tool descriptions teach the model.

Registration and gating

Tools follow the same rule as agents, connectors, and execution backends: configured at runtime, in the database, never through process environment. The shipped MCP connection policy lives in two tables from migration 0039:
Each approved fingerprint hashes the complete remote tool definition, including its input/output schemas and annotations. Discovery and execution compare it against the agent credential’s current catalog. New and changed supported tools are included automatically unless excluded. Preferences belong to the agent binding, never to the shared provider row. See Remote MCP for the normative discovery, approval, transport and provenance contract. The credential lives on the (agent, provider) binding and nowhere else. A provider row describes a server that any number of agents may use; each agent presents its own account. An agent without a binding has no credential and no offerable tools from that provider. Remote definitions enter the same tool registry and run ledger as built-ins, while their enablement uses the binding’s tool exclusions and current catalog. The existing agent_tools and agent_tool_config tables continue to govern built-in capabilities; MCP does not synthesize built-in policy rows or change their settings. Built-in tools are on by default, with explicit per-agent opt-outs. Tools become available when their requirements are configured: workspace tools need an execution backend, web tools need their effective provider settings, and tools that declare credentials in Needs need that agent’s own saved credential. No acting account is inherited from the instance or another agent. There is no requires_approval policy — see Authority is the agent’s account. Enabling a tool says this capability may appear in this agent’s prompt; what the agent may then do with it is bounded by the permissions on the agent’s own account in the external system. When the agent genuinely needs a human decision (“which of these two fixes?”), loop mode parks the run in awaiting_input and graph mode uses a human_gate node; either way any actor in the thread can answer, and the answer is an attributed message in the conversation graph — not a private modal only the requester can see. That is the multiplayer property the rest of Mesh is built around, applied to handoff. Per-connector gating exists because the same agent in a private DM and a shared channel should not necessarily have the same reach. The gate is evaluated at tool-offer time, not just call time: a tool the policy forbids in this conversation is not present in the model request at all. Models negotiate with error messages; they do not negotiate with absence. The in-memory enforcement seam exists today: Registry.Select returns a reduced registry, and that exact registry supplies both the model’s schemas and the executor’s name lookup. Built-in tools have database-backed per-agent enablement and secret-presence gates. Remote MCP connections add per-agent reviewed tool catalogs and exact conversation/audience scopes. Generalized per-connector policy for built-in tools remains a follow-up.

MCP is the extension surface, and it is remote

Implementation: Remote MCP tools provides the Streamable HTTP path (bearer token or OAuth), per-agent accounts, automatic activation of supported tools, persistent exclusions, and external provenance. New tools use durable effect semantics. Explicit unchanged legacy queries retain query behavior; changed definitions return to effect semantics. Legacy native tools remain available. Mesh does not grow a bespoke plugin API for third-party tools. The Model Context Protocol is the ecosystem’s answer, and the cloud-first constraint picks the transport: Mesh speaks remote MCP (Streamable HTTP with bearer-token authentication) to servers registered as tool_providers. Local stdio MCP servers — the default in local-first harnesses — are out of scope on purpose: a subprocess of the control plane is exactly the ambient-authority pattern the execution layer exists to kill. Connecting an MCP account makes its supported tools available immediately. The operator may exclude tools or turn the connection off. No classification checklist is required. The server’s readOnlyHint remains advisory; new operations use effect semantics so uncertain writes cannot be replayed. The credential is the agent’s own account, so the server enforces that account’s permissions. No workspace, model, or connector credential rides along on an MCP call.

Where each class runs: the cloud service map

The execution layer’s rule — Mesh integrates with a contract, vendors are backends — extends to every class that needs infrastructure. Named products below are the current best fits, recorded so the first implementation has a target; none of them may be named above their layer’s seam. Workspace backends. The managed-sandbox-API backend that sandboxed-execution.md ships first should be Modal Sandboxes: it has a supported Go SDK (the control plane is Go, and a sidecar in another language just to reach a vendor would be absurd), gVisor-isolated containers, exec with streamed output, file access, sandbox naming and tagging, idle timeouts, filesystem snapshots for the layered cold-start design, and — decisive for the credential model — egress control up to a domain allowlist, so default-deny-with-allowlist is a backend capability rather than a Mesh-side fence. The roadmap’s later backends are unchanged: an edge container platform (Cloudflare Sandboxes/Containers) is the sleep-and-wake Resumable case, Kubernetes is the own-perimeter case, local is last and dev-only. Browser backends. Cloudflare Browser Run (née Browser Rendering) is the first target: managed headless Chrome on Cloudflare’s network, a warm pool so sessions start without cold-start tax, a raw CDP endpoint the control plane can drive from Go, Live View (humans watch the session in real time), human in the loop (a person can click, type a password, solve the wall the agent hit, and hand control back), and session recordings as structured JSON. Those last three are not conveniences; they map one-to-one onto Mesh primitives — Live View links belong in progress emissions, human takeover is awaiting_input wearing a browser, and recordings are the observation artifact for an entire session. Cloudflare’s Kitesurf browser (lighter, no full Chromium) is a capability variant for read-mostly page work, not a separate integration. Browserbase or Steel slot in as alternate backends behind the same contract. A browser backend is not itself a model-facing tool. These are two halves of one capability at different seams: the model receives generic browser_navigate, browser_read, browser_act, and browser_screenshot tool definitions; the runtime provisions a session and executes those calls through whichever browser backend the agent was granted. Cloudflare belongs in the browser-backend descriptor and catalog, never in a tool name or in the harness. A vendor product may expose more than one integration shape without collapsing the seams. Browser Run’s stateless content and screenshot endpoints can back a provider-hosted query or one-shot browser tool, while its persistent CDP sessions back BrowserBackend. A shared vendor credential does not make those the same runtime contract: installation may coordinate the credential, but each call is still classified, gated, budgeted, and audited at the seam it uses. Query providers. Web search and page-to-markdown fetch are hosted APIs (Exa, Brave, Tavily, or similar) registered as tool_providers. The provider is a row; the tool is web_search. No scraping from the control plane. Effect providers. Slack actions ride the existing connector. GitHub, when it arrives, is an app installation whose short-lived installation tokens are exactly the minted-credential shape the execution layer already specifies — one integration serving both effect tools (open an issue) and workspace grants (clone this repo).

Browser sessions

A browser session is a provisioned resource with the same discipline as a workspace:
And a backend contract that is an intersection plus declared capabilities, exactly like execution backends:
Act is deliberately coarse in the contract; the tool layer above it exposes browser_navigate, browser_read, browser_act, browser_screenshot. Every action is a persisted tool call, and the released session’s recording is the run’s proof of what the agent actually did on the web — the browser equivalent of the prompt snapshot. The human-handoff flow is worth spelling out because no local-first harness can have it: the agent hits a login wall, the run parks (awaiting_human on the session, awaiting_input on the run), the thread gets a Live View link, any actor in the conversation — not just the requester — opens it, types the credential or clears the challenge, and hands back. The takeover is attributed to the actor who did it, and the credential never transits Mesh at all.

Budgets

Tool use adds no new budget philosophy, only new meters. The harness already owns max iterations, tool calls, tokens, and wall clock; the execution layer adds compute-seconds. Browser sessions add session-minutes, and query providers add per-provider call ceilings (search APIs bill per request, and a retry loop against a paid API is a money leak with no CPU signature). Everything lands in the same place: exhaustion is a terminal status with a reason, surfaced to the channel, and drawing more graph nodes still cannot manufacture more budget.

Progressive tool loading

An agent’s tool schemas are not free context. They are priced into the same input budget as its memory and its history (internal/turn/context_budget.go subtracts the whole request estimate, tool schemas included, and gives the system prompt what is left), so an agent connected to three MCP servers spends part of every turn carrying capabilities it is not using — and loses memory and history to pay for them. So a turn starts with the agent’s own verbs (memory, progress, commitments) plus one tool, tool_search, and loads an integration’s schemas only once the model has asked for them. A search returns the matching operations and offers their schemas on the next model request. Whether a turn does this at all is measured, not configured. It is not free: tool_search carries a bounded summary of the connected integrations in its description, and that travels on every call. So Mesh compares the two costs — carrying every schema, against carrying the index plus the handful a search loads — and an agent on the wrong side of that line is simply offered everything, as it always was. An agent with three tools carries all three. An agent with forty carries an index. Neither needs an operator to have known which it was. (Callable is the wrong word, and the distinction is the whole design: the executor resolves against the full authorized registry, so an operation the operator granted was always callable. What a search changes is whether the model has been shown how to call it.) What this changes is what the model is shown. What it may do is untouched:
  • The searchable index is built from the definitions the turn was already authorized to use, so anything reachable through search is something the operator already granted. There is no second authorization path.
  • Execution resolves against the full authorized registry, not the loaded subset, so a call already in flight cannot be invalidated by a later eviction. Offer-time policy, reviewed MCP fingerprints and call-time verification are identical either way.
  • An agent below the threshold behaves exactly as it did before. Nothing is stored on progressive loading’s behalf either way.
Loaded schemas are bounded by a share of the input budget. When a search would exceed it, the least recently used tools are unloaded and the model is told which, because a model that is not told will call one of them. Re-using a tool refreshes it, so the tool in active use is the last to go. A single tool larger than the whole budget is kept rather than evicted into an empty set. Activations are recorded in the continuation checkpoint, so a paused task does not have to rediscover what it was using. They are a request to re-activate, never a grant: a resumed turn looks each one up in the authority it has now and refuses anything revoked or redefined while the task waited, telling the model what it lost rather than letting it plan around a tool it no longer has. The summary in that description is why a model can search at all: it cannot look for a capability it does not know exists. It is also the whole cost of the mechanism, which is why the threshold above exists and why it is expressed in tokens rather than in a tool count — a handful of very large schemas is worth loading on demand where a handful of small ones is not. On a synthetic agent with 20 connected operations the first call carries the index instead of every schema, and internal/turn/tool_search_test.go logs both figures rather than asserting a single flattering one.

Ordering

The roadmap’s sequencing survives contact with this document, and the memory slice sharpens the first steps. What matters: the durable Run lands before the loop gets more rounds, and the graph substrate lands before exec-class tools make loops long and stateful (../roadmap.md (repository), Phases 3–3.5).
  1. Generalize the memory rounds into the harness loop. Runs and RunSteps from agent-harness.md, each tool round as a persisted step, tool_calls and observations as rows, the budget set enforced, and the memory verbs become the first registered tools instead of a hardcoded special case. This is also where the tool package format lands: the memory verbs move into tools/<name>/{manifest.toml, hooks} folders crawled at startup, so the registry is built from manifests rather than hand-registered in Go. No new infrastructure vendors.
  2. Query tools. tool_providers + agent_tools tables, the gating policy, and hosted web_search as the first provider-backed tool. This is the cheapest slice that makes agents visibly smarter, and it exercises registration, gating, and observation persistence with zero sandbox risk.
  3. The execution layer with Modal as the first backend, per sandboxed-execution.md: workspace tools (exec, read_file, write_file), lazy provisioning, egress allowlist, minted repo credentials, compute budgets. This is the slice where an agent first writes and runs code.
  4. Browser sessions on Browser Run: the backend contract, the four browser tools, Live View in progress emissions, human handoff through awaiting_input, recordings as artifacts.
  5. Effect tools and remote MCP: every effectful workspace or effect call crosses the commit boundary before it runs, and a resumed run past that boundary reports rather than re-executes; the remote MCP client so operators can attach third-party capability, authenticated as the agent’s own account, without Mesh shipping code.
Slices 1–2 can land against loop-mode runs while the WorkGraph substrate is still being built; they neither assume nor preclude graph execution. Slices 3–5 inherit graph-compatibility from the invariants they build on — workspace- per-run and session-per-run mean parallel nodes cannot share mutable state by construction.

Design guardrails

  • do not require executing a tool’s code to read its manifest — discovery, rendering, and enablement are pure reads of manifest.toml; execution happens only through Definition.Execute after an explicit grant
  • do not add a tool by editing a central switch — a tool is registered because its tools/<name>/ folder and manifest were crawled at startup
  • do not dynamically link third-party tool code into the control plane — the future plugin loader is out-of-process (subprocess/RPC or WASM), never Go plugin
  • do not execute exec-class tool calls in the harness process; do not fetch arbitrary URLs from it either
  • do not offer a tool the policy forbids — gate at offer time, not call time
  • do not infer a tool’s effect class from behavior; declare it, and treat a class violation as a security bug
  • do not configure a tool, provider, or backend through process environment
  • do not let a query tool read outside its RuntimeAgent perspective
  • do not model a browser session as a workspace, or bill it like one
  • do not import an MCP tool without an operator-assigned class and policy
  • do not run local stdio MCP servers in the control plane
  • do not put a per-call human approval in front of an effect — authority is the permissions on the agent’s own account in the external system, and a gate that substitutes for scoping is borrowed-credential thinking
  • do not let a human handoff become a private modal — awaiting_input and human_gate are attributed thread messages any actor can answer
  • do not write a first-party package for a third-party API with no Mesh-primitive coupling — that is what a CLI in the workspace or remote MCP is for, and both act as the agent’s own account
  • do not add CLI teaching to Go string literals beyond what codeWorkSection already carries — new CLI knowledge waits for the skill loader, and the existing section moves into a SKILL.md in the change that builds it
  • do not store an MCP provider’s credential on the provider row — it is the agent’s own login and lives on the (agent, provider) binding, or two agents sharing a server would act as one account
  • do not store a tool’s raw output in model context when it belongs in an artifact
  • do not let a provider name appear above its layer’s seam
  • do not let per-request provider billing escape the budget set