Tools
Strategy update (2026-09-18): The tool integration strategy defines the planned shared catalog, named connections, progressive discovery, and Add capability flow. The delivery plan tracks implementation. Descriptions of existing behavior below remain the current contract until those changes ship. A tool is a granted capability with an audit trail, not a function pointer. Mesh already has the two hard boundaries a tool system needs. The agent harness (agent-harness.md) makes the loop that calls tools durable,
budgeted, and attributable. The execution layer
(sandboxed-execution.md) makes the place exec-class
tools run isolated, credential-scoped, and pluggable. This document is the layer
between them: what a tool is, how tools are classified, registered, and gated,
where each class actually executes, and which cloud services implement the
classes that need infrastructure.
One thing is deliberately settled before the taxonomy, because it shapes all of
it:
Mesh is cloud-first. The default assumption of every popular agent harness — that tools run on the machine the harness runs on, in the user’s own filesystem, browser, and network — is exactly the assumption Mesh rejects. A Mesh agent’s shell is a rented sandbox, its browser is a rented browser session, and its web access goes through declared providers. Nothing about an agent’s reach is derived from the host the control plane happens to be deployed on.That is not only a security posture. It is what makes fifty co-tenanted agents on one small control-plane box possible, and it is what makes every capability a row that can be audited, budgeted, and revoked.
The first tool already exists
The memory verbs shipped with the memory slice: a provider-neutral tool contract ininternal/model (Tool, ToolCall, RoleTool), OpenRouter tool-call support,
bounded tool rounds, and fail-closed validation of everything the model asks for.
That slice proved the wire protocol.
There are four of them — remember, recall, amend, forget — and the split is
itself a contract rather than a convenience. remember creates and cannot modify;
correction is amend; retirement is forget. A verb that cannot supersede cannot
produce a result that honestly describes a supersession, which is half of the
truthful-confirmation contract in
memory-retrieval.md. The other half is that every reply
following a tool round is handed an authoritative ledger of what actually
committed, built from the store’s return values rather than the model’s requests.
Note also what these tools deliberately do NOT accept: any parameter naming the
source of a memory. Provenance is derived from the persisted inbound message’s
graph record, so a model has no channel through which to attribute a claim to
someone else.
The first generalization is now built in internal/tool: an immutable registry
binds provider-visible schemas to effect declarations, validation, and handlers;
one executor validates and invokes calls against an already policy-scoped
registry; and the four memory verbs use that path rather than a turn-loop switch.
Tool calls and typed observations remain persisted steps of a durable Run.
The round count remains capped at three and enforced structurally: on the
last permitted round the tool list is removed from the request. That is enough for
recall → amend → answer and short of an autonomous loop.
Effect classes
Tools are classified by effect surface — what a call can touch — because that is the property that determines where the call runs, what has to be provisioned for it, how cancellation treats it, and how suspicious the default policy should be. Four classes:query tools read: memory recall, conversation-graph lookups (“what did we
decide in this thread last week”), and web search or page-fetch through a
declared hosted provider. A remote MCP tool the operator has classified as
Read at review is a query tool too (docs/runtime/remote-mcp.md): the
operator’s classification, not the server’s hint, is what places it here. Query tools run in the control plane. That is not an
exception to the rule against executing tools in the harness process — that rule
exists for exec-class calls, and a query tool has no shell, no filesystem, and
no arbitrary network reach. The constraint that keeps the class honest: a query
tool may only talk to Mesh’s own database (perspective-scoped, per
context-resolution.md) or to a provider that was
registered and credentialed at runtime. A control-plane tool that fetches an
arbitrary URL is not a query tool; it is a server-side request forgery
primitive pointed at whatever network the control plane deploys into. Arbitrary
URLs are fetched by the browser class or from inside a workspace, where egress
policy applies and the blast radius is a rented container, not the control
plane.
workspace tools are the exec class: run a command, read and write files,
clone, build, test. They execute only inside a workspace materialized by the
execution layer, one workspace per run, provisioned lazily on first use.
sandboxed-execution.md is normative for all of it —
the backend contract, credential minting, egress default-deny, compute budgets.
Nothing in this document weakens it.
browser tools drive a cloud browser session: navigate, read the page,
click, type, screenshot. A browser session is a second provisioned resource
type, sibling to the workspace, with its own backend contract and its own
section below. It is deliberately not modelled as a workspace capability: the
lifecycle (a session humans can watch and take over), the budget unit (session
minutes), and the vendor landscape are all different.
effect tools act on the outside world through an API Mesh holds a
credential for: post a Slack message beyond the turn’s own reply, open a GitHub
issue, push a commit, send an email. Two things distinguish the class. First,
every effect call is an outbound event, not a side effect — same persistence,
idempotency, and delivery-state machinery as replies
(turn-pipeline.md). Second, the class interacts with the
commit boundary: an effect call is what moves a Run from buffering to
committed in the interrupt model, so it is the class whose calls must be
recorded before they run and never silently re-run after a resume.
The class is declared on the tool, not inferred from its behavior, and the
declaration is load-bearing: a tool whose implementation reaches outside its
declared class is a bug of the same severity as a backend that ignores egress
policy.
Authority is the agent’s account, not the harness
Every popular harness puts a human approval prompt in front of effectful calls, and it is worth being exact about why: those harnesses act with the user’s own credentials. The agent is borrowing a person’s GitHub login, so of course a person has to be asked before it pushes. The approval gate is a patch over borrowed identity. Mesh rejects the premise. A Mesh agent is provisioned like a new coworker: it has its own GitHub account, its own Linear seat, its own Slack identity, and the operator grants those accounts exactly the permissions the role needs — write on these repositories and not those, member and not admin. Every tool call is made with the agent’s own credential, stored encrypted and scoped to that agent (../agent-provisioning.md, migration 0005). What
the agent may do is therefore decided in the external system, by the
permissions attached to the agent’s account, and Mesh’s job is to make every
action attributable and reconstructable afterward — not to interpose a human on
each one.
Three consequences follow, and the rest of this document is written against
them:
- There is no per-call approval gate, and none is planned. An agent that
can open a pull request opens it, the same way an intern with write access
does, and the review happens where reviews happen: on the pull request. The
policycolumn below isenabled | disabled; there is norequires_approvalvalue. If a deployment ever wants a human decision in the loop for a specific workflow,awaiting_inputand the WorkGraphhuman_gatenode exist for decisions the agent needs (“which of these two fixes?”), and that is a different thing from a permission check. - Scoping happens at the account. The operator who wants an agent unable to delete repositories does not configure Mesh; they do not grant the agent’s GitHub account that permission. Mesh’s per-agent enablement and per-connector offer gating decide which capabilities are in the prompt, which is a question of relevance and blast radius, not of authorization.
- The residual risk is prompt injection, and the answer is still scoping. A message in a channel can try to talk an agent into an action. An intern can be socially engineered too; the reason that is survivable is that the intern’s access was bounded before the conversation started. Structural separation of untrusted content (roadmap F3) narrows the attack; account permissions bound the damage; the ledger makes it explainable. A yes/no prompt to a human who cannot see the injected text does none of those.
Three tiers: which transport a capability uses
Mesh has three ways to give an agent a capability. Choose by runtime coupling and who maintains the implementation. Acting service identity is explicit in every tier; infrastructure payer credentials can have instance defaults. The planned shared integration layer unifies setup and discovery across these paths without collapsing their execution boundaries.
Tier 1 is the smallest set and should stay that way:
workspace_exec,
workspace_read_file, workspace_write_file, and checkout_github_repo are
first-party because each does something an external server cannot — mint a
workspace, install a credential helper, bind a run to a checkout. A first-party
package must not be written for a third-party API that has no Mesh-primitive
coupling, unless Mesh deliberately owns a small provider-neutral capability
contract with multiple adapters, as it does for web search and reading. Tiers
2 and 3 cover ordinary vendor APIs; dogfooding alone does not justify a wrapper.
Tier 2 is the path the GitHub workflow already takes: the checkout installs a
checksum-pinned gh, the workspace holds GH_TOKEN for the agent’s GitHub
account, and the model is told to gh pr create. It is also the tier the
local-first harnesses converged on, for a reason that holds here too: a CLI’s
schema is not in the request, whereas offered MCP schemas consume context on
model calls. Today Mesh offers the operator-approved subset, which can still be
large. Planned progressive activation
will load relevant native and MCP schemas on demand. The CLI is not free — the skill that teaches it (below) is
rendered into the prompt — but a skill is a short procedure loaded when the work
is code, not a per-request catalog of every callable the vendor exposes. Mesh’s
version is stricter than theirs, because the CLI runs in a
rented sandbox holding one scoped credential rather than on the operator’s
machine with everything on it. The authority boundary of a tier-2 call is
exactly “what the workspace’s credentials can do” — which is the account-scoping
rule above, applied mechanically. What Mesh gives up is per-call structure: a
gh pr create inside workspace_exec is one effectful exec to the ledger, not
a typed create_pull_request step. That is why the commit boundary is crossed
at the class level (any effectful workspace or effect call) rather than per
tool, and why a workflow that needs a typed observation of a specific external
effect is a tier-1 candidate.
Tier 3 is the third-party surface, specified under
Registration and gating. The credential is the
agent’s own login to that service — an operator signs in to the vendor’s MCP
server as the agent once — and it is stored per agent, as an agent-owned
secret exactly like the agent’s GitHub or Linear key, never on the provider
row. That placement is load-bearing: a tool_provider describes a server
(endpoint), and several agents may enable tools from the same
server, so a credential on the provider would make every one of those agents
act as the same external account — which is the borrowed-identity failure this
whole section exists to rule out. The binding that carries the credential is
(agent, provider), as are reviewed tool fingerprints. Current connections are
agent-wide across conversations; see Remote MCP tools. The target
named-connection model retains single-agent ownership of every acting account,
as required by the identity guardrail (repository),
and keeps credentials off provider catalog definitions. Linear is the
first consumer: its remote MCP server replaces the first-party Linear package
once it reaches parity, and the package is retired rather than maintained
alongside.
Knowledge about how to use a tier-2 CLI belongs in a skill — a SKILL.md in the
AgentSkills shape, shipped inside the owning package and rendered into the prompt
as untrusted content, the way inbound bodies are — not in Go string literals.
Skills are curated into the repository; they are not pulled from a public
registry at runtime, because a registry of prompt text is a supply chain.
Status: not built. Today the gh workflow is taught by
internal/prompt.codeWorkSection, hand-written Go rendered from the offered
tool set, and that stays the mechanism until a skill loader exists. The
migration is one scheduled item: a package-level SKILL.md crawled with the
manifest, rendered where codeWorkSection renders today, with the same
“mention only offered tools” rule; codeWorkSection is deleted in the same
change, not before.
The tool contract
The provider wire schema stays deliberately small ininternal/model.Tool. An
executable definition in internal/tool binds that schema to Mesh-only policy and
behavior:
Replay is per-tool, not per-class. Most query and workspace reads are
replayable; workspace_exec and workspace_write_file are effectful; every
effect call is effectful. The harness already commits a replayability
classification on each RunStep — this field is where it comes from.
Replay is also what the commit boundary reads, and the rule is deliberately
the conservative one: an effectful call of any non-query class crosses the
run’s commit boundary before it runs, including a workspace_exec that only
ran the tests. This is the rule #216
implements in turn.Engine.executeRecordedTool (run.ToolEffect is the
decision, RecordCommit the write-ahead stamp); before it merges only the reply
crosses, and this paragraph describes the contract, not main. Mesh cannot see inside an exec — the same command line that runs
a test suite can git push — so “effectful with respect to its own workspace
but replayable with respect to the world” is a distinction the harness cannot
verify and therefore does not make. The cost is that a run interrupted after a
test-only exec is resumed by reporting rather than re-executing; that is the
right cost, because the workspace is torn down at run end anyway, so a resumed
coding run is lossy regardless. A tool that wants to be re-executed on resume
declares replayable, which is a promise about the world, not about the
workspace. The one exception is Internal: a Mesh-owned effectful verb (the
memory verbs) writes only to Mesh’s own Postgres and stays inside the boundary
(see interrupt-model.md, “The enumerated effect
surface”).
A handler error fails the turn unless the handler tags it. Two tags turn it
into an observation the model sees on its next completion:
meshTool.ErrInvalidToolInput (the arguments were wrong in a way the model can
fix) and, for query tools, meshTool.ErrToolUnavailable. The default stays
fatal so an infrastructure failure (an unreachable store, a ledger write) is
never mistaken for bad input. Argument decoding is the one input check every
first-party handler shares, so it goes through one decoder:
meshTool.DecodeArguments[T] requires a top-level JSON object (a bare null
would otherwise run the tool on its defaults) and rejects unknown fields and
anything after the object (the
crawl’s shape validator canonicalizes but does not enforce
additionalProperties: false, so the handler is where that is enforced), and
every refusal it returns carries ErrInvalidToolInput with a message written
for the model. An unknown field is named along with the accepted top-level
fields; a wrong type names the field and the expected JSON kind. Messages name
fields, never argument values. Before this, handlers returned the bare decoder
error, and a model that invented label on file_linear_issue lost the whole
turn to json: unknown field "label" instead of retrying without it.
Needs is how provisioning stays lazy and legible. A run that never invokes a
tool with Needs: ["workspace"] never provisions a workspace. A tool whose
needs include a named secret is invocable only when the run’s minted grant set
covers it. The declaration is also the answer to “what could this run have
reached,” which must stay a query, not an investigation.
An Observation is typed and persisted, never a bare string blob:
ArtifactID link are
still follow-on work; an executor refuses to label an observation truncated
without an artifact id, and the run repository currently refuses artifact-bearing
observations rather than silently losing the link.
Once artifact persistence lands, large outputs — a build log, a page dump — are stored as artifacts and
truncated into model context, not the other way around. The artifact row is
what the UI shows and what a WorkGraph edge can carry; the context excerpt is a
view of it.
The tool package format
This is the declared forward path for how first-party tools are structured on disk. Its schema is deliberately forward-compatible toward a future installable package — but that path is a kept-open option, not a commitment; remote MCP is Mesh’s third-party extension surface today (see “One third-party boundary” below). It is not a detour from the contract above — it is that contract’s on-disk shape. TheDefinition above is what a tool is to the runtime. This section is how
tools are packaged so that a human can see them, an operator can enable them,
and — in a future state — a third party can ship them, all without the control
plane having to execute a line of the code to learn what they are.
A folder is a package, not a tool
The unit on disk is an integration package, and it is deliberately not the same unit as aDefinition. A Definition models exactly one provider-visible
callable: one name, one input schema, one Class, one Replay policy, one
Needs set, one handler. But a real integration exposes many callables of
differing shape — the Linear package reads issues (a query) and creates them
(an effect); the GitHub package reads PRs, writes comments, and can grant
workspace access. Collapsing those onto a single package-level class would
either over-approve the reads or force every operation through one coarse compound
tool. So the format splits two layers explicitly:
- Package — folder-level metadata, shared config surface, and lifecycle hooks. Identity, versioning, the credential(s) the whole install needs, and the enable/disable/health lifecycle live here.
- Tools — one or more
[[tools]]entries, each of which becomes exactly oneDefinition: its ownname,description,parametersschema,class,replay,needs, and handler binding.
NewRegistry is
unchanged, it just receives the Definitions the crawler built from every
[[tools]] entry across every package.
Every enabled package lives under tools/ with its own directory carrying a
declarative manifest plus its compiled-in handlers:
The manifest
manifest.toml is the declarative half of the package. TOML is the chosen
format — it is typed, it takes comments, humans hand-edit it, and it has none of
YAML’s indentation or implicit-typing footguns for a file operators will read and
write. The manifest carries package identity and lifecycle, the shared
configuration surface an operator fills in, and one [[tools]] block per
callable — each block carrying exactly the declarative fields a Definition
requires, including the input schema:
[[tools]] entry carries its own input schema. This is load-bearing:
normalizeDefinition rejects a Definition whose model.Tool.Parameters is
absent or is not a JSON object, so a crawler could not produce a provider-visible
tool at all without this field. The schema is a JSON Schema (draft 2020-12)
document, given inline or via a bounded relative file reference under the
package folder (no traversal, size-capped — see Versioning). Keeping the schema in
Go instead would silently break the “identical manifest for future third-party
packages” claim, because a third-party loader has no Go to read it from. An
optional observation output schema is encouraged: it makes the future RPC/WASM
boundary far less ambiguous about what a handler is allowed to return.
Config disposition is a closed vocabulary, not three loose booleans
The earlier three-boolean form (secret / encrypt / visible_to_user) admitted
contradictory and unsafe states — secret=true, encrypt=false, or
secret=true, visible_to_user=true — and never said which actor a value was
visible to. It is replaced by a single closed disposition enum with a strict
truth table the server defines and enforces:
The invariant
secret ⇒ encrypted && never returned is not expressible as an
illegal state anymore — there is one field, and each value carries its whole
disposition. Unknown disposition values fail the manifest closed.
Where each disposition is stored, today: secret values go to secrets
(encrypted, owner-scoped, captured write-only through the tools tab); config
and public values go to agent_tool_config (0035), one row per (agent,
package, key), written from the same tools tab and validated by
internal/toolconfig against the manifest that declares them — unknown key,
secret key, or wrong type is refused at the boundary. The turn engine reads an
agent’s values once per turn and hands all of them to the package’s handlers
(effects.Runtime.ConfigValue) and the public ones to the prompt, rendered as
the agent’s own tool settings. That is the difference the two non-secret
dispositions make: a public default repository is something the agent can
name and act on; a config value reaches its handlers and nothing else.
Need mapping is explicit and resolves exactly once. A secret field declares
the stable secret-kind it satisfies via need, and a tool’s
needs = ["secret:linear_api_key"] refers to that same kind. The crawler verifies
that every secret:<kind> Need names a kind provided by exactly one secret
config field in the package — no dangling Needs, no ambiguous double-provides.
Secret ownership is derived by Mesh, not declared by the manifest. The
manifest says a secret is needed and of what kind; it does not get to say
who owns the stored credential. Ownership is resolved from the concrete
agent-tool/provider installation at capture time. Migration
0005_connector_scoped_secrets is the precedent and the reason: install
credentials are owned by owner_type='agent_connector' (owner_id =
agent_connectors.id), not blanket owner_type='agent', precisely because one
agent can hold many provider installs and a per-agent unique constraint made the
second install unstorable — and worse, invited cross-install credential
confusion. A manifest that could name its own owner would reintroduce exactly that
bug. So the manifest supplies kind; the installation supplies owner.
The invariant: the manifest is crawlable without executing tool code
This is the load-bearing rule of the entire format, and it is stated here because everything else depends on it:A manifest MUST be fully readable — parsed, validated, and rendered — without executing any of the package’s code. Discovering what a package is, what callables it exposes, what class each declares, what secrets it needs, and what its config surface looks like is a pure read of a data file. This invariant governs discovery and rendering only. It does not claim that code runs exclusively onData you can read is safe; code you have to run is not. Keeping those two apart is what buys three properties at once, for free:Definition.Execute— lifecycle hooks (on_enableand friends) are real, effectful code that runs at enable/disable/capture time. Those are not discovery; they are the durable lifecycle defined in “Lifecycle” below, and they run only after an explicit operator action, never during a crawl.
- The UI renders from manifests. Startup crawls
tools/, parses eachmanifest.toml, and the operator surface (Track D) is a projection of that set — every offerable tool, its class, and its config fields, with zero tool-specific UI code and zero tool code executed to draw the screen. - Enablement and secret capture happen before the tool ever runs. C1.5’s
flow — “enabling a tool whose
Needsdeclares a secret blocks until the secret is captured” — is only possible because the requirement is readable in advance. If learning a tool’s needs required running it, secret-gated enablement would be a chicken-and-egg problem. - Third-party plugins become safe to catalog. The moment code runs just to describe itself, importing an untrusted plugin means executing untrusted code to render a settings page. The invariant is the security boundary that makes a future plugin ecosystem tractable: Mesh can read, display, and reason about a plugin it has not yet decided to trust with execution.
Startup crawl
A directory crawl on its own cannot discover a compiled Go handler just becausetool.go sits next to the TOML — a Go package only enters the binary when
something imports it, so “the folder is present” does not make its code linked and
callable. Left implicit, that gap collapses back into a central handwritten list,
init() side effects, or some unspecified binding magic — the very switch
statement this format exists to delete. So binding is a two-phase contract,
stated explicitly:
- Build time — generation. A code generator walks the in-tree
tools/packages and emits a deterministic compiled-handler catalog: a generated map from(package_id, handler_name)to the actual Go function, compiled into the binary. The generator also embeds and attests the first-party manifests, so what the binary ships with is fixed at build time, not re-read from an arbitrary disk at boot. - Startup — parse + verify, never execute. The runtime parses every
(embedded, first-party) manifest without executing a handler, then verifies
a 1:1 match between each
[[tools]]handlerreference and an entry in the generated catalog, and between each declared package hook and its bound function. It then builds theDefinitions and hands them toNewRegistry.
- Missing binding (a
handlerwith no catalog entry), extra binding (a catalog entry no manifest references), or duplicate (namecollision within or across packages, whichNewRegistryalready rejects) is a hard error with an operator-visible diagnostic naming the package, the tool, and the handler. - For an in-tree first-party built-in, an invalid or unmatched manifest fails build / start / readiness — a shipped capability silently vanishing is worse than a loud failure, so we fail loud.
- Quarantine-with-process-survival (skip the one bad package, keep serving the rest) is reserved for the future third-party path, where one untrusted package must never be able to take down the process. First-party fails closed at the process; third-party fails closed at the package.
Lifecycle
The manifest declares which lifecycle junctures a package participates in; the runtime binds each declared hook to a compiled-in Go handler. The hook set:
(The per-call
Validate → Execute path is not a package hook — it is the
Definition contract, one per [[tools]] entry, unchanged from the top of this
doc.)
These hooks are external effects and inherit the effect guarantees.
on_enable/on_disable provision or tear down webhooks, revoke sessions, mint
provider state — exactly the kind of external side effect the document already
requires be persistent, idempotent, and delivery-guaranteed for tool calls. A
naïve “just call the hook” would let a crash after the provider succeeds but
before local state commits duplicate a webhook on retry, or leave one live after
a disable. So lifecycle runs through a durable lifecycle executor, distinct
from the per-call executor but built on the same discipline:
- Attributed transition records. Every attempt writes a durable record: who requested it, which package/install, from-state → to-state, pinned to the exact config version and secret version in force at attempt time.
- Stable idempotency keys. Each provider-facing action carries a key derived
from
(install, transition, config+secret version), so a retry after an ambiguous crash reconciles instead of duplicating. - Typed outcomes + retry classification. Success / retriable / terminal are explicit; deadlines bound each attempt; retriable failures back off, terminal failures surface to the operator.
- State machine. Enable moves
disabled → enabling → enabled | error, and the package is not offerable until it reachesenabled. Disable gates new calls immediately, then reconciles cleanup asynchronously; a failed teardown parks inerrorwith a compensation task, it does not silently strand provider state. - Serialization. Concurrent enable / disable / rotation / health-check on the
same install are serialized; in-flight calls are drained or cancelled per a
declared policy before a disable completes. Rotation re-runs
on_secret_captureagainst the new secret version before cutting over.
Out of the gates: compiled-in; future: the same manifest, a different loader
An honest constraint has to be stated plainly, because it shapes the phasing: Go has no safe way to hot-load a zip of native tool code. Theplugin
package is fragile, platform-bound, and version-locked to the host binary, and a
downloaded shared object is arbitrary in-process code with the control plane’s
full authority — exactly the ambient-authority pattern Mesh exists to reject.
So the phasing is deliberate and the architecture is stable across it:
- Out of the gates, every tool’s hooks are compiled-in Go, one in-tree
package per
tools/<name>/folder, registered by the startup crawl. The manifest is the declarative sidecar; the code ships in the Mesh binary. This is first-party only — we are not accepting third-party plugins yet. - Future state, third parties ship a tool as a zipped folder containing the
identical
manifest.tomlplus code, and Mesh runs that code behind a different loading mechanism — an out-of-process boundary (subprocess/RPC or WASM), never dynamically linked into the control plane.
Versioning and trust boundaries
“Forward-compatible” is a claim that needs mechanism, not just intent — a future loader can only safely read today’s file if the compatibility and trust boundaries are pinned now. The schema therefore fixes:manifest_version— the schema version of the manifest file itself. Present from v1 (shown above). A reader that does not understand amanifest_versionrefuses the package rather than guessing.- Immutable package identity + version —
package.idis globally unique and stable;package.versionis semver. Grants and audit records pin to(id, version, digest), never to a mutable folder name. - Mesh API/ABI range —
package.mesh_apideclares the runtime API range the package targets; a package outside the running Mesh’s supported range is refused, loud. - Hook protocol version — the out-of-process hook/RPC contract carries its own
version, independent of
manifest_version, so the wire protocol and the file format evolve separately. - Strict unknown-field policy — first-party manifests reject unknown fields at build time (typos fail loud). The third-party path uses the same strictness so a field a newer publisher added cannot silently change meaning under an older Mesh.
- Provenance + integrity — a content digest over the package and a publisher signature; Mesh verifies both, and every grant/audit record is pinned to the verified digest.
- Staged, atomic activation — an archive is validated in a staging area and only atomically activated on success; a failed validation never half-installs. Updates are atomic with rollback and explicit config migration between package versions.
- Archive hardening — ingestion rejects path traversal, symlinks, oversized manifests, excessive file counts, and decompression bombs.
One third-party boundary: MCP is the surface, packages are the structure
The installable-code path above and the later normative statement that remote MCP is Mesh’s third-party extension surface must not read as two competing answers to “how does a third party extend Mesh.” They are reconciled explicitly:- Remote MCP is the third-party execution surface. A third party who wants Mesh to run their capability attaches it as a remote MCP server — already out-of-process, already “describe-then-trust,” already class-and-policy assigned at registration. That is the boundary an implementer builds to today.
- The package format is the first-party structure. It is how tools that ship with Mesh are organized on disk, and its schema is deliberately forward-compatible. Forward-compatible is a property, not a committed roadmap item: we are not promising an installable-plugin loader. If we ever ship one, the versioning + trust mechanisms above are its prerequisites — they exist so the option stays open, not because it is scheduled.
Relationship to remote MCP
Remote MCP (below) and the tool package format are not competitors, and the boundary between them is the one drawn just above: MCP is the third-party execution surface; the package format is first-party structure. A remote MCP server is already an out-of-process, declaratively-described capability — its tool list is metadata Mesh reads before it trusts, and its class and policy are operator-assigned at registration. The package format brings that same “describe-then-trust” discipline to the tools that ship with Mesh. Both obey the manifest invariant — Mesh learns what a thing is by reading, and runs it only after an explicit grant — which is exactly why the schema can stay forward-compatible toward a future installable path without that path being a commitment today.The exec-class tools that exist
The firstworkspace-class tools are built, on the package format above, and
they are the reason the execution layer exists:
The shape follows the rule at the top of this document — a tool call and the
process that made it are not the same object. The four
workspace-class tools
execute against a workspace the execution layer materialized for the run
(sandboxed-execution.md §What ships today),
reached through the per-agent runtime on the context exactly the way the Linear
and GitHub handlers reach their API clients. create_github_pull_request is
the one effect in the table: it runs in the control plane against GitHub’s
API, as the fallback for opening the PR when gh is not available in the
workspace. Under the three-tier rule
it is a tier-2 capability wearing a tier-1 wrapper, kept only until gh is
guaranteed in the workspace image; it is not the pattern to copy for the next
GitHub operation. Two decisions in the workspace tools are worth naming:
- A failing command is content, not an error.
workspace_execreturns the exit code and both streams verbatim (clamped, head and tail kept), because the agent iterates on compiler output and test failures. So are the layer’s own degraded states — no backend configured, transport down — which come back as observations naming the fix, because a Go error from a handler fails the turn and “your operator has not configured a sandbox” is something the agent should be able to say. checkout_github_repois the GitHub package’s workspace grant. It proves the credential and reads the default branch from the control plane first, then clones into the workspace with git configured to authenticate from theGH_TOKENthe workspace was granted and to commit as the agent’s GitHub account. The workflow it sets up — edit withworkspace_write_file, test withworkspace_exec, commit and push with git, open the PR withgh pr createorcreate_github_pull_request— is what the tool descriptions teach the model.
Registration and gating
Tools follow the same rule as agents, connectors, and execution backends: configured at runtime, in the database, never through process environment. The shipped MCP connection policy lives in two tables from migration0039:
(agent, provider) binding and nowhere else. A
provider row describes a server that any number of agents may use; each agent
presents its own account. An agent without a binding has no credential and no
offerable tools from that provider. Remote definitions enter the same tool
registry and run ledger as built-ins, while their enablement uses the binding’s
tool exclusions and current catalog. The existing agent_tools and
agent_tool_config tables continue to govern built-in capabilities; MCP does
not synthesize built-in policy rows or change their settings.
Built-in tools are on by default, with explicit per-agent opt-outs. Tools become
available when their requirements are configured: workspace tools need an
execution backend, web tools need their effective provider settings, and tools
that declare credentials in Needs need that agent’s own saved credential.
No acting account is inherited from the instance or another agent. There is no requires_approval policy — see
Authority is the agent’s account.
Enabling a tool says this capability may appear in this agent’s prompt; what
the agent may then do with it is bounded by the permissions on the agent’s own
account in the external system. When the agent genuinely needs a human decision
(“which of these two fixes?”), loop mode parks the run in awaiting_input and
graph mode uses a human_gate node; either way any actor in the thread can
answer, and the answer is an attributed message in the conversation graph — not
a private modal only the requester can see. That is the multiplayer property the
rest of Mesh is built around, applied to handoff.
Per-connector gating exists because the same agent in a private DM and a shared
channel should not necessarily have the same reach. The gate is evaluated at
tool-offer time, not just call time: a tool the policy forbids in this
conversation is not present in the model request at all. Models negotiate with
error messages; they do not negotiate with absence.
The in-memory enforcement seam exists today: Registry.Select returns a reduced
registry, and that exact registry supplies both the model’s schemas and the
executor’s name lookup. Built-in tools have database-backed per-agent enablement
and secret-presence gates. Remote MCP connections add per-agent reviewed tool
catalogs and exact conversation/audience scopes. Generalized per-connector policy
for built-in tools remains a follow-up.
MCP is the extension surface, and it is remote
Implementation: Remote MCP tools provides the Streamable HTTP path (bearer token or OAuth), per-agent accounts, automatic activation of supported tools, persistent exclusions, and external provenance. New tools use durable effect semantics. Explicit unchanged legacy queries retain query behavior; changed definitions return to effect semantics. Legacy native tools remain available. Mesh does not grow a bespoke plugin API for third-party tools. The Model Context Protocol is the ecosystem’s answer, and the cloud-first constraint picks the transport: Mesh speaks remote MCP (Streamable HTTP with bearer-token authentication) to servers registered astool_providers. Local stdio MCP servers — the default in
local-first harnesses — are out of scope on purpose: a subprocess of the control
plane is exactly the ambient-authority pattern the execution layer exists to
kill.
Connecting an MCP account makes its supported tools available immediately. The
operator may exclude tools or turn the connection off. No classification checklist
is required. The server’s readOnlyHint remains advisory; new operations use effect
semantics so uncertain writes cannot be replayed. The credential is the agent’s own
account, so the server enforces that account’s permissions. No workspace, model, or
connector credential rides along on an MCP call.
Where each class runs: the cloud service map
The execution layer’s rule — Mesh integrates with a contract, vendors are backends — extends to every class that needs infrastructure. Named products below are the current best fits, recorded so the first implementation has a target; none of them may be named above their layer’s seam. Workspace backends. The managed-sandbox-API backend thatsandboxed-execution.md ships first should be
Modal Sandboxes: it has a supported Go SDK (the control plane is Go, and a
sidecar in another language just to reach a vendor would be absurd),
gVisor-isolated containers, exec with streamed output, file access, sandbox
naming and tagging, idle timeouts, filesystem snapshots for the layered
cold-start design, and — decisive for the credential model — egress control up
to a domain allowlist, so default-deny-with-allowlist is a backend capability
rather than a Mesh-side fence. The roadmap’s later backends are unchanged:
an edge container platform (Cloudflare Sandboxes/Containers) is the
sleep-and-wake Resumable case, Kubernetes is the own-perimeter case, local is
last and dev-only.
Browser backends. Cloudflare Browser Run (née Browser Rendering) is the
first target: managed headless Chrome on Cloudflare’s network, a warm pool so
sessions start without cold-start tax, a raw CDP endpoint the control plane can
drive from Go, Live View (humans watch the session in real time), human
in the loop (a person can click, type a password, solve the wall the agent
hit, and hand control back), and session recordings as structured JSON.
Those last three are not conveniences; they map one-to-one onto Mesh
primitives — Live View links belong in progress emissions, human takeover is
awaiting_input wearing a browser, and recordings are the observation artifact
for an entire session. Cloudflare’s Kitesurf browser (lighter, no full Chromium)
is a capability variant for read-mostly page work, not a separate integration.
Browserbase or Steel slot in as alternate backends behind the same contract.
A browser backend is not itself a model-facing tool. These are two halves of
one capability at different seams: the model receives generic
browser_navigate, browser_read, browser_act, and browser_screenshot tool
definitions; the runtime provisions a session and executes those calls through
whichever browser backend the agent was granted. Cloudflare belongs in the
browser-backend descriptor and catalog, never in a tool name or in the harness.
A vendor product may expose more than one integration shape without collapsing
the seams. Browser Run’s stateless content and screenshot endpoints can back a
provider-hosted query or one-shot browser tool, while its persistent CDP
sessions back BrowserBackend. A shared vendor credential does not make those
the same runtime contract: installation may coordinate the credential, but each
call is still classified, gated, budgeted, and audited at the seam it uses.
Query providers. Web search and page-to-markdown fetch are hosted APIs
(Exa, Brave, Tavily, or similar) registered as tool_providers. The provider
is a row; the tool is web_search. No scraping from the control plane.
Effect providers. Slack actions ride the existing connector. GitHub, when
it arrives, is an app installation whose short-lived installation tokens are
exactly the minted-credential shape the execution layer already specifies —
one integration serving both effect tools (open an issue) and workspace
grants (clone this repo).
Browser sessions
A browser session is a provisioned resource with the same discipline as a workspace:Act is deliberately coarse in the contract; the tool layer above it exposes
browser_navigate, browser_read, browser_act, browser_screenshot. Every
action is a persisted tool call, and the released session’s recording is the
run’s proof of what the agent actually did on the web — the browser equivalent
of the prompt snapshot.
The human-handoff flow is worth spelling out because no local-first harness can
have it: the agent hits a login wall, the run parks (awaiting_human on the
session, awaiting_input on the run), the thread gets a Live View link, any
actor in the conversation — not just the requester — opens it, types the
credential or clears the challenge, and hands back. The takeover is attributed
to the actor who did it, and the credential never transits Mesh at all.
Budgets
Tool use adds no new budget philosophy, only new meters. The harness already owns max iterations, tool calls, tokens, and wall clock; the execution layer adds compute-seconds. Browser sessions add session-minutes, and query providers add per-provider call ceilings (search APIs bill per request, and a retry loop against a paid API is a money leak with no CPU signature). Everything lands in the same place: exhaustion is a terminal status with a reason, surfaced to the channel, and drawing more graph nodes still cannot manufacture more budget.Progressive tool loading
An agent’s tool schemas are not free context. They are priced into the same input budget as its memory and its history (internal/turn/context_budget.go
subtracts the whole request estimate, tool schemas included, and gives the
system prompt what is left), so an agent connected to three MCP servers spends
part of every turn carrying capabilities it is not using — and loses memory and
history to pay for them.
So a turn starts with the agent’s own verbs (memory, progress, commitments)
plus one tool, tool_search, and loads an integration’s schemas only once the
model has asked for them. A search returns the matching operations and offers
their schemas on the next model request.
Whether a turn does this at all is measured, not configured. It is not free:
tool_search carries a bounded summary of the connected integrations in its
description, and that travels on every call. So Mesh compares the two costs —
carrying every schema, against carrying the index plus the handful a search
loads — and an agent on the wrong side of that line is simply offered
everything, as it always was. An agent with three tools carries all three. An
agent with forty carries an index. Neither needs an operator to have known
which it was.
(Callable is the wrong word, and the distinction is the whole design: the
executor resolves against the full authorized registry, so an operation the
operator granted was always callable. What a search changes is whether the
model has been shown how to call it.)
What this changes is what the model is shown. What it may do is
untouched:
- The searchable index is built from the definitions the turn was already authorized to use, so anything reachable through search is something the operator already granted. There is no second authorization path.
- Execution resolves against the full authorized registry, not the loaded subset, so a call already in flight cannot be invalidated by a later eviction. Offer-time policy, reviewed MCP fingerprints and call-time verification are identical either way.
- An agent below the threshold behaves exactly as it did before. Nothing is stored on progressive loading’s behalf either way.
internal/turn/tool_search_test.go logs
both figures rather than asserting a single flattering one.
Ordering
The roadmap’s sequencing survives contact with this document, and the memory slice sharpens the first steps. What matters: the durable Run lands before the loop gets more rounds, and the graph substrate lands before exec-class tools make loops long and stateful (../roadmap.md (repository), Phases 3–3.5).
- Generalize the memory rounds into the harness loop. Runs and RunSteps
from
agent-harness.md, each tool round as a persisted step,tool_callsand observations as rows, the budget set enforced, and the memory verbs become the first registered tools instead of a hardcoded special case. This is also where the tool package format lands: the memory verbs move intotools/<name>/{manifest.toml, hooks}folders crawled at startup, so the registry is built from manifests rather than hand-registered in Go. No new infrastructure vendors. - Query tools.
tool_providers+agent_toolstables, the gating policy, and hostedweb_searchas the first provider-backed tool. This is the cheapest slice that makes agents visibly smarter, and it exercises registration, gating, and observation persistence with zero sandbox risk. - The execution layer with Modal as the first backend, per
sandboxed-execution.md: workspace tools (exec,read_file,write_file), lazy provisioning, egress allowlist, minted repo credentials, compute budgets. This is the slice where an agent first writes and runs code. - Browser sessions on Browser Run: the backend contract, the four browser
tools, Live View in progress emissions, human handoff through
awaiting_input, recordings as artifacts. - Effect tools and remote MCP: every effectful
workspaceoreffectcall crosses the commit boundary before it runs, and a resumed run past that boundary reports rather than re-executes; the remote MCP client so operators can attach third-party capability, authenticated as the agent’s own account, without Mesh shipping code.
Design guardrails
- do not require executing a tool’s code to read its manifest — discovery,
rendering, and enablement are pure reads of
manifest.toml; execution happens only throughDefinition.Executeafter an explicit grant - do not add a tool by editing a central switch — a tool is registered because its
tools/<name>/folder and manifest were crawled at startup - do not dynamically link third-party tool code into the control plane — the
future plugin loader is out-of-process (subprocess/RPC or WASM), never Go
plugin - do not execute exec-class tool calls in the harness process; do not fetch arbitrary URLs from it either
- do not offer a tool the policy forbids — gate at offer time, not call time
- do not infer a tool’s effect class from behavior; declare it, and treat a class violation as a security bug
- do not configure a tool, provider, or backend through process environment
- do not let a query tool read outside its RuntimeAgent perspective
- do not model a browser session as a workspace, or bill it like one
- do not import an MCP tool without an operator-assigned class and policy
- do not run local stdio MCP servers in the control plane
- do not put a per-call human approval in front of an effect — authority is the permissions on the agent’s own account in the external system, and a gate that substitutes for scoping is borrowed-credential thinking
- do not let a human handoff become a private modal —
awaiting_inputandhuman_gateare attributed thread messages any actor can answer - do not write a first-party package for a third-party API with no Mesh-primitive coupling — that is what a CLI in the workspace or remote MCP is for, and both act as the agent’s own account
- do not add CLI teaching to Go string literals beyond what
codeWorkSectionalready carries — new CLI knowledge waits for the skill loader, and the existing section moves into aSKILL.mdin the change that builds it - do not store an MCP provider’s credential on the provider row — it is the
agent’s own login and lives on the
(agent, provider)binding, or two agents sharing a server would act as one account - do not store a tool’s raw output in model context when it belongs in an artifact
- do not let a provider name appear above its layer’s seam
- do not let per-request provider billing escape the budget set