> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Tools

# Tools

**Strategy update (2026-09-18):** The
[tool integration strategy](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/tool-integration-strategy.md) defines the planned shared
catalog, named connections, progressive discovery, and Add capability flow. The
[delivery plan](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap/tool-integration-delivery.md) tracks implementation.
Descriptions of existing behavior below remain the current contract until those
changes ship.

A tool is a granted capability with an audit trail, not a function pointer.

Mesh already has the two hard boundaries a tool system needs. The agent harness
([`agent-harness.md`](/runtime/agent-harness)) makes the loop that calls tools durable,
budgeted, and attributable. The execution layer
([`sandboxed-execution.md`](/runtime/sandboxed-execution)) makes the place exec-class
tools run isolated, credential-scoped, and pluggable. This document is the layer
between them: what a tool *is*, how tools are classified, registered, and gated,
where each class actually executes, and which cloud services implement the
classes that need infrastructure.

One thing is deliberately settled before the taxonomy, because it shapes all of
it:

> **Mesh is cloud-first.** The default assumption of every popular agent harness
> — that tools run on the machine the harness runs on, in the user's own
> filesystem, browser, and network — is exactly the assumption Mesh rejects.
> A Mesh agent's shell is a rented sandbox, its browser is a rented browser
> session, and its web access goes through declared providers. Nothing about an
> agent's reach is derived from the host the control plane happens to be
> deployed on.

That is not only a security posture. It is what makes fifty co-tenanted agents
on one small control-plane box possible, and it is what makes every capability a
row that can be audited, budgeted, and revoked.

## The first tool already exists

The memory verbs shipped with the memory slice: a provider-neutral tool contract
in `internal/model` (`Tool`, `ToolCall`, `RoleTool`), OpenRouter tool-call support,
bounded tool rounds, and fail-closed validation of everything the model asks for.
That slice proved the wire protocol.

There are four of them — `remember`, `recall`, `amend`, `forget` — and the split is
itself a contract rather than a convenience. `remember` creates and cannot modify;
correction is `amend`; retirement is `forget`. A verb that cannot supersede cannot
produce a result that honestly describes a supersession, which is half of the
truthful-confirmation contract in
[`memory-retrieval.md`](/runtime/memory-retrieval). The other half is that every reply
following a tool round is handed an authoritative ledger of what actually
committed, built from the store's return values rather than the model's requests.

Note also what these tools deliberately do NOT accept: any parameter naming the
source of a memory. Provenance is derived from the persisted inbound message's
graph record, so a model has no channel through which to attribute a claim to
someone else.

The first generalization is now built in `internal/tool`: an immutable registry
binds provider-visible schemas to effect declarations, validation, and handlers;
one executor validates and invokes calls against an already policy-scoped
registry; and the four memory verbs use that path rather than a turn-loop switch.
Tool calls and typed observations remain persisted steps of a durable Run.

The round count remains capped at three and enforced structurally: on the
last permitted round the tool list is removed from the request. That is enough for
`recall` → `amend` → answer and short of an autonomous loop.

## Effect classes

Tools are classified by **effect surface** — what a call can touch — because
that is the property that determines where the call runs, what has to be
provisioned for it, how cancellation treats it, and how suspicious the default
policy should be. Four classes:

```
query      — reads against Mesh state or a declared hosted provider
workspace  — shell, files, git, builds: anything inside a sandboxed workspace
browser    — actions inside a cloud browser session
effect     — externally visible actions through a connector or provider API
```

**`query`** tools read: memory recall, conversation-graph lookups ("what did we
decide in this thread last week"), and web search or page-fetch through a
*declared hosted provider*. A remote MCP tool the operator has classified as
Read at review is a query tool too (`docs/runtime/remote-mcp.md`): the
operator's classification, not the server's hint, is what places it here. Query tools run in the control plane. That is not an
exception to the rule against executing tools in the harness process — that rule
exists for exec-class calls, and a query tool has no shell, no filesystem, and
no arbitrary network reach. The constraint that keeps the class honest: a query
tool may only talk to Mesh's own database (perspective-scoped, per
[`context-resolution.md`](/runtime/context-resolution)) or to a provider that was
registered and credentialed at runtime. A control-plane tool that fetches an
**arbitrary URL** is not a query tool; it is a server-side request forgery
primitive pointed at whatever network the control plane deploys into. Arbitrary
URLs are fetched by the `browser` class or from inside a workspace, where egress
policy applies and the blast radius is a rented container, not the control
plane.

**`workspace`** tools are the exec class: run a command, read and write files,
clone, build, test. They execute only inside a workspace materialized by the
execution layer, one workspace per run, provisioned lazily on first use.
[`sandboxed-execution.md`](/runtime/sandboxed-execution) is normative for all of it —
the backend contract, credential minting, egress default-deny, compute budgets.
Nothing in this document weakens it.

**`browser`** tools drive a cloud browser session: navigate, read the page,
click, type, screenshot. A browser session is a second provisioned resource
type, sibling to the workspace, with its own backend contract and its own
section below. It is deliberately *not* modelled as a workspace capability: the
lifecycle (a session humans can watch and take over), the budget unit (session
minutes), and the vendor landscape are all different.

**`effect`** tools act on the outside world through an API Mesh holds a
credential for: post a Slack message beyond the turn's own reply, open a GitHub
issue, push a commit, send an email. Two things distinguish the class. First,
every effect call is an outbound event, not a side effect — same persistence,
idempotency, and delivery-state machinery as replies
([`turn-pipeline.md`](/runtime/turn-pipeline)). Second, the class interacts with the
commit boundary: an effect call is what moves a Run from `buffering` to
`committed` in the interrupt model, so it is the class whose calls must be
recorded before they run and never silently re-run after a resume.

The class is declared on the tool, not inferred from its behavior, and the
declaration is load-bearing: a tool whose implementation reaches outside its
declared class is a bug of the same severity as a backend that ignores egress
policy.

## Authority is the agent's account, not the harness

Every popular harness puts a human approval prompt in front of effectful calls,
and it is worth being exact about why: those harnesses act with **the user's own
credentials**. The agent is borrowing a person's GitHub login, so of course a
person has to be asked before it pushes. The approval gate is a patch over
borrowed identity.

Mesh rejects the premise. A Mesh agent is provisioned like a new coworker: it
has its own GitHub account, its own Linear seat, its own Slack identity, and the
operator grants those accounts exactly the permissions the role needs — write
on these repositories and not those, member and not admin. Every tool call is
made with the agent's own credential, stored encrypted and scoped to that agent
([`../agent-provisioning.md`](/agent-provisioning), migration `0005`). What
the agent may do is therefore decided **in the external system, by the
permissions attached to the agent's account**, and Mesh's job is to make every
action attributable and reconstructable afterward — not to interpose a human on
each one.

Three consequences follow, and the rest of this document is written against
them:

* **There is no per-call approval gate, and none is planned.** An agent that
  can open a pull request opens it, the same way an intern with write access
  does, and the review happens where reviews happen: on the pull request. The
  `policy` column below is `enabled | disabled`; there is no `requires_approval`
  value. If a deployment ever wants a human decision in the loop for a specific
  workflow, `awaiting_input` and the WorkGraph `human_gate` node exist for
  *decisions the agent needs* ("which of these two fixes?"), and that is a
  different thing from a permission check.
* **Scoping happens at the account.** The operator who wants an agent unable to
  delete repositories does not configure Mesh; they do not grant the agent's
  GitHub account that permission. Mesh's per-agent enablement and per-connector
  offer gating decide *which capabilities are in the prompt*, which is a
  question of relevance and blast radius, not of authorization.
* **The residual risk is prompt injection, and the answer is still scoping.** A
  message in a channel can try to talk an agent into an action. An intern can be
  socially engineered too; the reason that is survivable is that the intern's
  access was bounded before the conversation started. Structural separation of
  untrusted content (roadmap F3) narrows the attack; account permissions bound
  the damage; the ledger makes it explainable. A yes/no prompt to a human who
  cannot see the injected text does none of those.

The credential is the agent's. Whether the transport that carries it is a
first-party Go handler, a CLI inside the workspace, or a remote MCP server is
the next section's question and is orthogonal to this one.

### Three tiers: which transport a capability uses

Mesh has three ways to give an agent a capability. Choose by runtime coupling
and who maintains the implementation. Acting service identity is explicit in
every tier; infrastructure payer credentials can have instance defaults. The
[planned shared integration layer](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/tool-integration-strategy.md) unifies setup
and discovery across these paths without collapsing their execution boundaries.

| Tier | What | When | Who maintains it |
| - | - | - | - |
| **1 · First-party package** | `tools/<name>/` with compiled Go handlers | The capability is coupled to a Mesh primitive: materializing a workspace, granting it a credential, crossing the commit boundary, reading the perspective-scoped graph | Mesh |
| **2 · CLI in the workspace** | The vendor's own CLI (`gh`, `git`, …) installed in the sandbox, driven through `workspace_exec`, authenticated by a credential the workspace was granted | The service has a CLI the model already knows and the call needs nothing from Mesh beyond a shell and a scoped credential | The vendor |
| **3 · Remote MCP** | A remote MCP server registered as a `tool_provider`, authenticated as the agent's account | The service has no good CLI, or a third party owns the capability, or the operator wants to attach a capability without Mesh shipping code for it | The vendor or the third party |

Tier 1 is the smallest set and should stay that way: `workspace_exec`,
`workspace_read_file`, `workspace_write_file`, and `checkout_github_repo` are
first-party because each does something an external server cannot — mint a
workspace, install a credential helper, bind a run to a checkout. **A first-party
package must not be written for a third-party API that has no Mesh-primitive
coupling, unless Mesh deliberately owns a small provider-neutral capability
contract with multiple adapters, as it does for web search and reading.** Tiers
2 and 3 cover ordinary vendor APIs; dogfooding alone does not justify a wrapper.

Tier 2 is the path the GitHub workflow already takes: the checkout installs a
checksum-pinned `gh`, the workspace holds `GH_TOKEN` for the agent's GitHub
account, and the model is told to `gh pr create`. It is also the tier the
local-first harnesses converged on, for a reason that holds here too: a CLI's
schema is not in the request, whereas offered MCP schemas consume context on
model calls. Today Mesh offers the operator-approved subset, which can still be
large. Planned [progressive activation](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/tool-integration-strategy.md#what-the-agent-sees)
will load relevant native and MCP schemas on demand. The CLI is not free — the skill that teaches it (below) is
rendered into the prompt — but a skill is a short procedure loaded when the work
is code, not a per-request catalog of every callable the vendor exposes. Mesh's
version is stricter than theirs, because the CLI runs in a
rented sandbox holding one scoped credential rather than on the operator's
machine with everything on it. The authority boundary of a tier-2 call is
exactly "what the workspace's credentials can do" — which is the account-scoping
rule above, applied mechanically. What Mesh gives up is per-call structure: a
`gh pr create` inside `workspace_exec` is one effectful exec to the ledger, not
a typed `create_pull_request` step. That is why the commit boundary is crossed
at the class level (any effectful `workspace` or `effect` call) rather than per
tool, and why a workflow that needs a typed observation of a specific external
effect is a tier-1 candidate.

Tier 3 is the third-party surface, specified under
[Registration and gating](#registration-and-gating). The credential is the
agent's own login to that service — an operator signs in to the vendor's MCP
server *as the agent* once — and it is stored **per agent**, as an agent-owned
secret exactly like the agent's GitHub or Linear key, never on the provider
row. That placement is load-bearing: a `tool_provider` describes a server
(endpoint), and several agents may enable tools from the same
server, so a credential on the provider would make every one of those agents
act as the same external account — which is the borrowed-identity failure this
whole section exists to rule out. The binding that carries the credential is
`(agent, provider)`, as are reviewed tool fingerprints. Current connections are
agent-wide across conversations; see [Remote MCP tools](/runtime/remote-mcp). The target
named-connection model retains single-agent ownership of every acting account,
as required by the [identity guardrail (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap.md#agent-owned-acting-identities),
and keeps credentials off provider catalog definitions. Linear is the
first consumer: its remote MCP server replaces the first-party Linear package
once it reaches parity, and the package is retired rather than maintained
alongside.

Knowledge about how to use a tier-2 CLI belongs in a skill — a `SKILL.md` in the
AgentSkills shape, shipped inside the owning package and rendered into the prompt
as untrusted content, the way inbound bodies are — not in Go string literals.
Skills are curated into the repository; they are not pulled from a public
registry at runtime, because a registry of prompt text is a supply chain.
**Status: not built.** Today the `gh` workflow is taught by
`internal/prompt.codeWorkSection`, hand-written Go rendered from the offered
tool set, and that stays the mechanism until a skill loader exists. The
migration is one scheduled item: a package-level `SKILL.md` crawled with the
manifest, rendered where `codeWorkSection` renders today, with the same
"mention only offered tools" rule; `codeWorkSection` is deleted in the same
change, not before.

## The tool contract

The provider wire schema stays deliberately small in `internal/model.Tool`. An
executable definition in `internal/tool` binds that schema to Mesh-only policy and
behavior:

```go theme={null}
type Definition struct {
    Tool     model.Tool // Name, Description, Parameters: the provider-visible part
    Class    Class      // query | workspace | browser | effect
    Replay   Replay     // replayable | effectful
    Needs    []string   // declared capabilities/grants
    Validate Validator  // raw request -> canonical authorized arguments
    Execute  Handler    // validated Call -> typed Observation
}
```

`Replay` is per-tool, not per-class. Most `query` and `workspace` reads are
replayable; `workspace_exec` and `workspace_write_file` are effectful; every
`effect` call is effectful. The harness already commits a replayability
classification on each RunStep — this field is where it comes from.

`Replay` is also what the commit boundary reads, and the rule is deliberately
the conservative one: **an effectful call of any non-query class crosses the
run's commit boundary before it runs**, including a `workspace_exec` that only
ran the tests. This is the rule [#216](https://github.com/TextureHQ/mesh/pull/216)
implements in `turn.Engine.executeRecordedTool` (`run.ToolEffect` is the
decision, `RecordCommit` the write-ahead stamp); before it merges only the reply
crosses, and this paragraph describes the contract, not `main`. Mesh cannot see inside an exec — the same command line that runs
a test suite can `git push` — so "effectful with respect to its own workspace
but replayable with respect to the world" is a distinction the harness cannot
verify and therefore does not make. The cost is that a run interrupted after a
test-only exec is resumed by *reporting* rather than re-executing; that is the
right cost, because the workspace is torn down at run end anyway, so a resumed
coding run is lossy regardless. A tool that wants to be re-executed on resume
declares `replayable`, which is a promise about the world, not about the
workspace. The one exception is `Internal`: a Mesh-owned effectful verb (the
memory verbs) writes only to Mesh's own Postgres and stays inside the boundary
(see [`interrupt-model.md`](/runtime/interrupt-model), "The enumerated effect
surface").

A handler error fails the turn unless the handler tags it. Two tags turn it
into an observation the model sees on its next completion:
`meshTool.ErrInvalidToolInput` (the arguments were wrong in a way the model can
fix) and, for query tools, `meshTool.ErrToolUnavailable`. The default stays
fatal so an infrastructure failure (an unreachable store, a ledger write) is
never mistaken for bad input. Argument decoding is the one input check every
first-party handler shares, so it goes through one decoder:
`meshTool.DecodeArguments[T]` requires a top-level JSON object (a bare `null`
would otherwise run the tool on its defaults) and rejects unknown fields and
anything after the object (the
crawl's shape validator canonicalizes but does not enforce
`additionalProperties: false`, so the handler is where that is enforced), and
every refusal it returns carries `ErrInvalidToolInput` with a message written
for the model. An unknown field is named along with the accepted top-level
fields; a wrong type names the field and the expected JSON kind. Messages name
fields, never argument values. Before this, handlers returned the bare decoder
error, and a model that invented `label` on `file_linear_issue` lost the whole
turn to `json: unknown field "label"` instead of retrying without it.

`Needs` is how provisioning stays lazy and legible. A run that never invokes a
tool with `Needs: ["workspace"]` never provisions a workspace. A tool whose
needs include a named secret is invocable only when the run's minted grant set
covers it. The declaration is also the answer to "what could this run have
reached," which must stay a query, not an investigation.

An `Observation` is typed and persisted, never a bare string blob:

```go theme={null}
type Observation struct {
    Content    string          // what the model sees, bounded
    Truncated  bool
    ArtifactID *uuid.UUID      // full output, stored out of context
    Metadata   json.RawMessage // exit codes, URLs visited, bytes read, …
}
```

The contract exists today. Artifact storage and the durable `ArtifactID` link are
still follow-on work; an executor refuses to label an observation truncated
without an artifact id, and the run repository currently refuses artifact-bearing
observations rather than silently losing the link.

Once artifact persistence lands, large outputs — a build log, a page dump — are stored as artifacts and
truncated *into* model context, not the other way around. The artifact row is
what the UI shows and what a WorkGraph edge can carry; the context excerpt is a
view of it.

## The tool package format

*This is the declared forward path for how first-party tools are structured on
disk. Its schema is deliberately forward-compatible toward a future installable
package — but that path is a kept-open option, not a commitment; **remote MCP is
Mesh's third-party extension surface today** (see "One third-party boundary"
below). It is not a detour from the contract above — it is that contract's on-disk
shape.*

The `Definition` above is what a tool *is* to the runtime. This section is how
tools are *packaged* so that a human can see them, an operator can enable them,
and — in a future state — a third party can ship them, all without the control
plane having to execute a line of the code to learn what they are.

#### A folder is a *package*, not a tool

The unit on disk is an **integration package**, and it is deliberately *not* the
same unit as a `Definition`. A `Definition` models exactly one provider-visible
callable: one name, one input schema, one `Class`, one `Replay` policy, one
`Needs` set, one handler. But a real integration exposes *many* callables of
differing shape — the Linear package reads issues (a `query`) **and** creates them
(an `effect`); the GitHub package reads PRs, writes comments, and can grant
workspace access. Collapsing those onto a single package-level `class` would
either over-approve the reads or force every operation through one coarse compound
tool. So the format splits two layers explicitly:

* **Package** — folder-level metadata, *shared* config surface, and lifecycle
  hooks. Identity, versioning, the credential(s) the whole install needs, and the
  enable/disable/health lifecycle live here.
* **Tools** — one or more `[[tools]]` entries, each of which becomes exactly one
  `Definition`: its own `name`, `description`, `parameters` schema, `class`,
  `replay`, `needs`, and handler binding.

The registry keeps registering **callables, not packages** — `NewRegistry` is
unchanged, it just receives the `Definition`s the crawler built from every
`[[tools]]` entry across every package.

Every enabled package lives under `tools/` with its own directory carrying a
declarative manifest plus its compiled-in handlers:

```
tools/
  linear/
    manifest.toml
    tool.go          # handlers for each [[tools]] entry + lifecycle hooks
  github/
    manifest.toml
    tool.go
  sentry/     …
  deepsource/ …
  datadog/    …
```

The seed set is the tools we already run on ourselves — **Linear, GitHub, Sentry,
DeepSource, Datadog** — chosen because dogfooding is the cheapest correctness
test, and because they cover the shapes the format has to survive: an issue
tracker, a code host, two observability providers, and a code-quality gate. The
long tail is explicitly anticipated: someone else runs Jira, Trello, GitLab,
Gitea; someone else has PagerDuty instead of Datadog. The format is designed for
**dozens to hundreds** of these, so nothing about adding the eleventh package may
require editing a central switch statement — a package is present because its
folder is present.

### The manifest

`manifest.toml` is the declarative half of the package. **TOML is the chosen
format** — it is typed, it takes comments, humans hand-edit it, and it has none of
YAML's indentation or implicit-typing footguns for a file operators will read and
write. The manifest carries package identity and lifecycle, the shared
configuration surface an operator fills in, and one `[[tools]]` block per
callable — each block carrying *exactly* the declarative fields a `Definition`
requires, including the input schema:

```toml theme={null}
# tools/linear/manifest.toml
manifest_version = 1              # schema version of this file (see Versioning)

[package]
id          = "mesh.linear"       # stable, globally-unique package identity
version     = "1.0.0"             # package version, semver
description = "Read and write Linear issues."
mesh_api    = ">=1.0.0 <2.0.0"    # Mesh runtime API range this package targets

# Lifecycle hooks the PACKAGE participates in. Discovery/rendering is not a hook;
# these are the durable, effectful transitions defined in "Lifecycle" below.
[package.hooks]
on_enable         = true
on_disable        = true
on_secret_capture = true
health_check      = true

# Shared configuration surface. Each field declares one closed disposition, so the
# enablement UI and the secret-capture flow render straight from here. See the
# disposition table below — the three-boolean form is gone.
[[config]]
key         = "api_key"
type        = "string"
disposition = "secret"         # secret => encrypted at rest && never returned
need        = "linear_api_key" # stable secret-kind; the Need this field resolves
required    = true
description = "Linear personal API key (team-scoped)."

[[config]]
key         = "team_key"
type        = "string"
disposition = "public"         # rendered in cleartext, editable, safe in audit view
required    = true
default     = "MESH"
description = "Default Linear team key for created issues."

# One [[tools]] entry per provider-visible callable. Each becomes one Definition.
[[tools]]
name        = "linear_search_issues"
description = "Search Linear issues by query."
class       = "query"          # query | workspace | browser | effect
replay      = "replayable"     # replayable | effectful
needs       = ["secret:linear_api_key"]
handler     = "SearchIssues"   # binds to the compiled-in handler catalog
parameters  = "schemas/search_issues.input.json"   # bounded relative file ref
observation = "schemas/search_issues.output.json"  # optional, sharpens RPC/WASM boundary

[[tools]]
name        = "linear_create_issue"
description = "Create a Linear issue."
class       = "effect"         # a write is an effect, even in the same package
replay      = "effectful"
needs       = ["secret:linear_api_key"]
handler     = "CreateIssue"
parameters  = "schemas/create_issue.input.json"
```

**Every `[[tools]]` entry carries its own input schema.** This is load-bearing:
`normalizeDefinition` *rejects* a `Definition` whose `model.Tool.Parameters` is
absent or is not a JSON object, so a crawler could not produce a provider-visible
tool at all without this field. The schema is a **JSON Schema (draft 2020-12)**
document, given inline or via a **bounded relative file reference** under the
package folder (no traversal, size-capped — see Versioning). Keeping the schema in
Go instead would silently break the "identical manifest for future third-party
packages" claim, because a third-party loader has no Go to read it from. An
optional `observation` output schema is encouraged: it makes the future RPC/WASM
boundary far less ambiguous about what a handler is allowed to return.

#### Config disposition is a closed vocabulary, not three loose booleans

The earlier three-boolean form (`secret` / `encrypt` / `visible_to_user`) admitted
contradictory and unsafe states — `secret=true, encrypt=false`, or
`secret=true, visible_to_user=true` — and never said *which* actor a value was
visible to. It is replaced by a single closed `disposition` enum with a strict
truth table the server defines and enforces:

| `disposition` | Encrypted at rest | Returned to a reader | Rendered in operator UI |
| - | - | - | - |
| `secret` | **always** | **never** (write-only; redacted hint only) | capture form only, never the value |
| `config` | no | to operators, not to the agent/model | cleartext, editable |
| `public` | no | yes | cleartext, editable, safe in audit view |

The invariant `secret ⇒ encrypted && never returned` is not expressible as an
illegal state anymore — there is one field, and each value carries its whole
disposition. Unknown disposition values fail the manifest closed.

Where each disposition is stored, today: `secret` values go to `secrets`
(encrypted, owner-scoped, captured write-only through the tools tab); `config`
and `public` values go to `agent_tool_config` (`0035`), one row per (agent,
package, key), written from the same tools tab and validated by
`internal/toolconfig` against the manifest that declares them — unknown key,
secret key, or wrong type is refused at the boundary. The turn engine reads an
agent's values once per turn and hands all of them to the package's handlers
(`effects.Runtime.ConfigValue`) and the `public` ones to the prompt, rendered as
the agent's own tool settings. That is the difference the two non-secret
dispositions make: a `public` default repository is something the agent can
name and act on; a `config` value reaches its handlers and nothing else.

**Need mapping is explicit and resolves exactly once.** A `secret` field declares
the stable secret-**kind** it satisfies via `need`, and a tool's
`needs = ["secret:linear_api_key"]` refers to that same kind. The crawler verifies
that every `secret:<kind>` Need names a `kind` provided by exactly one `secret`
config field in the package — no dangling Needs, no ambiguous double-provides.

**Secret *ownership* is derived by Mesh, not declared by the manifest.** The
manifest says a secret is *needed* and of what *kind*; it does **not** get to say
who owns the stored credential. Ownership is resolved from the concrete
agent-tool/provider *installation* at capture time. Migration
`0005_connector_scoped_secrets` is the precedent and the reason: install
credentials are owned by `owner_type='agent_connector'` (owner\_id =
`agent_connectors.id`), *not* blanket `owner_type='agent'`, precisely because one
agent can hold many provider installs and a per-agent unique constraint made the
second install unstorable — and worse, invited cross-install credential
confusion. A manifest that could name its own owner would reintroduce exactly that
bug. So the manifest supplies *kind*; the installation supplies *owner*.

### The invariant: the manifest is crawlable without executing tool code

This is the load-bearing rule of the entire format, and it is stated here because
everything else depends on it:

> **A manifest MUST be fully readable — parsed, validated, and rendered — without
> executing any of the package's code.** Discovering what a package is, what
> callables it exposes, what class each declares, what secrets it needs, and what
> its config surface looks like is a pure read of a data file. This invariant
> governs **discovery and rendering only.** It does *not* claim that code runs
> exclusively on `Definition.Execute` — lifecycle hooks (`on_enable` and friends)
> are real, effectful code that runs at enable/disable/capture time. Those are
> not discovery; they are the durable lifecycle defined in "Lifecycle" below, and
> they run only after an explicit operator action, never during a crawl.

Data you can read is safe; code you have to run is not. Keeping those two apart is
what buys three properties at once, for free:

1. **The UI renders from manifests.** Startup crawls `tools/`, parses each
   `manifest.toml`, and the operator surface (Track D) is a projection of that
   set — every offerable tool, its class, and its config fields, with zero
   tool-specific UI code and zero tool code executed to draw the screen.
2. **Enablement and secret capture happen before the tool ever runs.** C1.5's
   flow — "enabling a tool whose `Needs` declares a secret blocks until the secret
   is captured" — is only possible because the requirement is *readable in
   advance*. If learning a tool's needs required running it, secret-gated
   enablement would be a chicken-and-egg problem.
3. **Third-party plugins become safe to catalog.** The moment code runs just to
   describe itself, importing an untrusted plugin means executing untrusted code
   to render a settings page. The invariant is the security boundary that makes a
   future plugin ecosystem tractable: Mesh can read, display, and reason about a
   plugin it has not yet decided to trust with execution.

### Startup crawl

A directory crawl on its own **cannot** discover a compiled Go handler just
because `tool.go` sits next to the TOML — a Go package only enters the binary when
something imports it, so "the folder is present" does not make its code linked and
callable. Left implicit, that gap collapses back into a central handwritten list,
`init()` side effects, or some unspecified binding magic — the very switch
statement this format exists to delete. So binding is a **two-phase contract**,
stated explicitly:

1. **Build time — generation.** A code generator walks the in-tree `tools/`
   packages and emits a **deterministic compiled-handler catalog**: a generated
   map from `(package_id, handler_name)` to the actual Go function, compiled into
   the binary. The generator also embeds and attests the first-party manifests,
   so what the binary ships with is fixed at build time, not re-read from an
   arbitrary disk at boot.
2. **Startup — parse + verify, never execute.** The runtime parses every
   (embedded, first-party) manifest *without executing a handler*, then verifies
   a **1:1 match** between each `[[tools]]` `handler` reference and an entry in
   the generated catalog, and between each declared package hook and its bound
   function. It then builds the `Definition`s and hands them to `NewRegistry`.

Mismatch behavior is defined, not left to chance:

* **Missing binding** (a `handler` with no catalog entry), **extra binding** (a
  catalog entry no manifest references), or **duplicate** (`name` collision within
  or across packages, which `NewRegistry` already rejects) is a hard error with an
  operator-visible diagnostic naming the package, the tool, and the handler.
* For an **in-tree first-party built-in**, an invalid or unmatched manifest fails
  **build / start / readiness** — a shipped capability silently vanishing is worse
  than a loud failure, so we fail loud.
* **Quarantine-with-process-survival** (skip the one bad package, keep serving the
  rest) is reserved for the *future* third-party path, where one untrusted
  package must never be able to take down the process. First-party fails closed at
  the process; third-party fails closed at the package.

### Lifecycle

The manifest declares which lifecycle junctures a package participates in; the
runtime binds each declared hook to a compiled-in Go handler. The hook set:

| Hook | When it fires | Typical use |
| - | - | - |
| `on_enable` | operator turns the package on for an agent | provision webhooks, verify the credential works |
| `on_disable` | operator turns it off | tear down webhooks, revoke sessions |
| `on_secret_capture` | a `secret` config field is captured | validate the credential before the package settles to enabled |
| `health_check` | periodic / on demand | surface reachability in the operator UI |

(The per-call `Validate` → `Execute` path is *not* a package hook — it is the
`Definition` contract, one per `[[tools]]` entry, unchanged from the top of this
doc.)

**These hooks are external effects and inherit the effect guarantees.**
`on_enable`/`on_disable` provision or tear down webhooks, revoke sessions, mint
provider state — exactly the kind of external side effect the document already
requires be persistent, idempotent, and delivery-guaranteed for tool calls. A
naïve "just call the hook" would let a crash *after* the provider succeeds but
*before* local state commits duplicate a webhook on retry, or leave one live after
a disable. So lifecycle runs through a **durable lifecycle executor**, distinct
from the per-call executor but built on the same discipline:

* **Attributed transition records.** Every attempt writes a durable record: who
  requested it, which package/install, from-state → to-state, pinned to the exact
  **config version and secret version** in force at attempt time.
* **Stable idempotency keys.** Each provider-facing action carries a key derived
  from `(install, transition, config+secret version)`, so a retry after an
  ambiguous crash reconciles instead of duplicating.
* **Typed outcomes + retry classification.** Success / retriable / terminal are
  explicit; deadlines bound each attempt; retriable failures back off, terminal
  failures surface to the operator.
* **State machine.** Enable moves `disabled → enabling → enabled | error`, and the
  package is **not offerable** until it reaches `enabled`. Disable **gates new
  calls immediately**, then reconciles cleanup asynchronously; a failed teardown
  parks in `error` with a compensation task, it does not silently strand
  provider state.
* **Serialization.** Concurrent enable / disable / rotation / health-check on the
  same install are serialized; in-flight calls are drained or cancelled per a
  declared policy before a disable completes. Rotation re-runs `on_secret_capture`
  against the new secret version before cutting over.

### Out of the gates: compiled-in; future: the same manifest, a different loader

An honest constraint has to be stated plainly, because it shapes the phasing:
**Go has no safe way to hot-load a zip of native tool code.** The `plugin`
package is fragile, platform-bound, and version-locked to the host binary, and a
downloaded shared object is arbitrary in-process code with the control plane's
full authority — exactly the ambient-authority pattern Mesh exists to reject.

So the phasing is deliberate and the architecture is stable across it:

* **Out of the gates**, every tool's hooks are **compiled-in Go**, one in-tree
  package per `tools/<name>/` folder, registered by the startup crawl. The
  manifest is the declarative sidecar; the code ships in the Mesh binary. This is
  first-party only — we are *not* accepting third-party plugins yet.
* **Future state**, third parties ship a tool as a **zipped folder containing the
  identical `manifest.toml`** plus code, and Mesh runs that code behind a
  *different loading mechanism* — an out-of-process boundary (subprocess/RPC or
  WASM), never dynamically linked into the control plane.

The folder-and-manifest architecture is identical on day one and at hundreds of
plugins. Only the code-loading boundary changes, and it changes *outward* — from
compiled-in, to sandboxed out-of-process — which is the direction that adds
safety, not removes it.

### Versioning and trust boundaries

"Forward-compatible" is a claim that needs *mechanism*, not just intent — a
future loader can only safely read today's file if the compatibility and trust
boundaries are pinned now. The schema therefore fixes:

* **`manifest_version`** — the schema version of the manifest file itself.
  Present from v1 (shown above). A reader that does not understand a
  `manifest_version` refuses the package rather than guessing.
* **Immutable package identity + version** — `package.id` is globally unique and
  stable; `package.version` is semver. Grants and audit records pin to
  `(id, version, digest)`, never to a mutable folder name.
* **Mesh API/ABI range** — `package.mesh_api` declares the runtime API range the
  package targets; a package outside the running Mesh's supported range is
  refused, loud.
* **Hook protocol version** — the out-of-process hook/RPC contract carries its own
  version, independent of `manifest_version`, so the wire protocol and the file
  format evolve separately.
* **Strict unknown-field policy** — first-party manifests reject unknown fields at
  build time (typos fail loud). The third-party path uses the same strictness so
  a field a newer publisher added cannot silently change meaning under an older
  Mesh.

The third-party archive path additionally requires, *before* a manifest may
influence any grant:

* **Provenance + integrity** — a content **digest** over the package and a
  **publisher signature**; Mesh verifies both, and every grant/audit record is
  pinned to the verified digest.
* **Staged, atomic activation** — an archive is validated in a staging area and
  only atomically activated on success; a failed validation never half-installs.
  Updates are atomic with **rollback** and explicit config migration between
  package versions.
* **Archive hardening** — ingestion rejects path traversal, symlinks,
  oversized manifests, excessive file counts, and decompression bombs.

### One third-party boundary: MCP is the surface, packages are the structure

The installable-code path above and the later normative statement that **remote
MCP is Mesh's third-party extension surface** must not read as two competing
answers to "how does a third party extend Mesh." They are reconciled explicitly:

* **Remote MCP is the third-party *execution* surface.** A third party who wants
  Mesh to run *their* capability attaches it as a remote MCP server — already
  out-of-process, already "describe-then-trust," already class-and-policy assigned
  at registration. That is the boundary an implementer builds to today.
* **The package format is the first-party *structure*.** It is how tools that
  ship *with* Mesh are organized on disk, and its schema is deliberately
  forward-compatible. Forward-compatible is a *property*, **not** a committed
  roadmap item: we are not promising an installable-plugin loader. If we ever
  ship one, the versioning + trust mechanisms above are its prerequisites — they
  exist so the option stays open, not because it is scheduled.

So: **one unambiguous third-party boundary today — remote MCP.** The package
format is first-party structure whose schema does not foreclose a future
installable path. **This is the path we are moving forward on.**

### Relationship to remote MCP

Remote MCP (below) and the tool package format are not competitors, and the
boundary between them is the one drawn just above: **MCP is the third-party
execution surface; the package format is first-party structure.** A remote MCP
server is *already* an out-of-process, declaratively-described capability — its
tool list is metadata Mesh reads before it trusts, and its class and policy are
operator-assigned at registration. The package format brings that same
"describe-then-trust" discipline to the tools that ship *with* Mesh. Both obey the
manifest invariant — Mesh learns what a thing is by reading, and runs it only
after an explicit grant — which is exactly why the schema *can* stay
forward-compatible toward a future installable path without that path being a
commitment today.

### The exec-class tools that exist

The first `workspace`-class tools are built, on the package format above, and
they are the reason the execution layer exists:

| Tool | Package | Class / replay | Needs |
| - | - | - | - |
| `workspace_exec` | `tools/workspace` | workspace / effectful | `workspace` |
| `workspace_read_file` | `tools/workspace` | workspace / replayable | `workspace` |
| `workspace_write_file` | `tools/workspace` | workspace / effectful | `workspace` |
| `checkout_github_repo` | `tools/github` | workspace / effectful | `workspace`, `secret:github_pat` |
| `create_github_pull_request` | `tools/github` | effect / effectful | `secret:github_pat` |

The shape follows the rule at the top of this document — a tool call and the
process that made it are not the same object. The four `workspace`-class tools
execute against a workspace the execution layer materialized for the run
([`sandboxed-execution.md` §What ships today](/runtime/sandboxed-execution#what-ships-today)),
reached through the per-agent runtime on the context exactly the way the Linear
and GitHub handlers reach their API clients. `create_github_pull_request` is
the one `effect` in the table: it runs in the control plane against GitHub's
API, as the fallback for opening the PR when `gh` is not available in the
workspace. Under the [three-tier rule](#three-tiers-which-transport-a-capability-uses)
it is a tier-2 capability wearing a tier-1 wrapper, kept only until `gh` is
guaranteed in the workspace image; it is not the pattern to copy for the next
GitHub operation. Two decisions in the workspace tools are worth naming:

* **A failing command is content, not an error.** `workspace_exec` returns the
  exit code and both streams verbatim (clamped, head and tail kept), because the
  agent iterates on compiler output and test failures. So are the layer's own
  degraded states — no backend configured, transport down — which come back as
  observations naming the fix, because a Go error from a handler fails the turn
  and "your operator has not configured a sandbox" is something the agent should
  be able to *say*.
* **`checkout_github_repo` is the GitHub package's workspace grant.** It proves
  the credential and reads the default branch from the control plane first, then
  clones into the workspace with git configured to authenticate from the
  `GH_TOKEN` the workspace was granted and to commit as the agent's GitHub
  account. The workflow it sets up — edit with `workspace_write_file`, test with
  `workspace_exec`, commit and push with git, open the PR with `gh pr create` or
  `create_github_pull_request` — is what the tool descriptions teach the model.

## Registration and gating

Tools follow the same rule as agents, connectors, and execution backends:
**configured at runtime, in the database, never through process environment.**

The shipped MCP connection policy lives in two tables from migration `0039`:

```text theme={null}
tool_providers (
  id,
  kind,                 -- mcp
  name,                 -- initial server label; agent-facing names live below
  endpoint,             -- unique remote server URL
  created_at
)

agent_tool_providers (
  agent_id,
  provider_id,
  name,                 -- this agent's connection label
  credential_ref,       -- agent-owned encrypted secret
  enabled,              -- on after connection unless explicitly turned off
  tool_policy,          -- all (legacy selected lists migrate to exclusions)
  disabled_tools,       -- names explicitly excluded, including removed names
  approved_tools,       -- [{name, fingerprint, class: "query"}]
  conversation_scopes,  -- [{conversation_id, audience}]: reviewed snapshots
  version,              -- fences stale approvals and in-flight registries
  updated_by_user_id,
  updated_at
  -- PRIMARY KEY (agent_id, provider_id)
)
```

Each approved fingerprint hashes the complete remote tool definition, including
its input/output schemas and annotations. Discovery and execution compare it
against the agent credential's current catalog. New and changed supported tools
are included automatically unless excluded. Preferences belong to the agent
binding, never to the shared provider row. See [Remote MCP](/runtime/remote-mcp) for the
normative discovery, approval, transport and provenance contract.

The credential lives on the `(agent, provider)` binding and nowhere else. A
provider row describes a server that any number of agents may use; each agent
presents its own account. An agent without a binding has no credential and no
offerable tools from that provider. Remote definitions enter the same tool
registry and run ledger as built-ins, while their enablement uses the binding's
tool exclusions and current catalog. The existing `agent_tools` and
`agent_tool_config` tables continue to govern built-in capabilities; MCP does
not synthesize built-in policy rows or change their settings.

Built-in tools are on by default, with explicit per-agent opt-outs. Tools become
available when their requirements are configured: workspace tools need an
execution backend, web tools need their effective provider settings, and tools
that declare credentials in `Needs` need that agent's own saved credential.
No acting account is inherited from the instance or another agent. There is no `requires_approval` policy — see
[Authority is the agent's account](#authority-is-the-agents-account-not-the-harness).
Enabling a tool says *this capability may appear in this agent's prompt*; what
the agent may then do with it is bounded by the permissions on the agent's own
account in the external system. When the agent genuinely needs a human decision
("which of these two fixes?"), loop mode parks the run in `awaiting_input` and
graph mode uses a `human_gate` node; either way *any actor in the thread* can
answer, and the answer is an attributed message in the conversation graph — not
a private modal only the requester can see. That is the multiplayer property the
rest of Mesh is built around, applied to handoff.

Per-connector gating exists because the same agent in a private DM and a shared
channel should not necessarily have the same reach. The gate is evaluated at
tool-*offer* time, not just call time: a tool the policy forbids in this
conversation is not present in the model request at all. Models negotiate with
error messages; they do not negotiate with absence.

The in-memory enforcement seam exists today: `Registry.Select` returns a reduced
registry, and that exact registry supplies both the model's schemas and the
executor's name lookup. Built-in tools have database-backed per-agent enablement
and secret-presence gates. Remote MCP connections add per-agent reviewed tool
catalogs and exact conversation/audience scopes. Generalized per-connector policy
for built-in tools remains a follow-up.

### MCP is the extension surface, and it is remote

**Implementation:** [Remote MCP tools](/runtime/remote-mcp) provides the Streamable
HTTP path (bearer token or OAuth), per-agent accounts, automatic activation of
supported tools, persistent exclusions, and external provenance. New tools use
durable effect semantics. Explicit unchanged legacy queries retain query behavior;
changed definitions return to effect semantics. Legacy native tools remain available.

Mesh does not grow a bespoke plugin API for third-party tools. The Model Context
Protocol is the ecosystem's answer, and the cloud-first constraint picks the
transport: Mesh speaks **remote MCP** (Streamable HTTP with bearer-token authentication) to
servers registered as `tool_providers`. Local stdio MCP servers — the default in
local-first harnesses — are out of scope on purpose: a subprocess of the control
plane is exactly the ambient-authority pattern the execution layer exists to
kill.

Connecting an MCP account makes its supported tools available immediately. The
operator may exclude tools or turn the connection off. No classification checklist
is required. The server's `readOnlyHint` remains advisory; new operations use effect
semantics so uncertain writes cannot be replayed. The credential is the agent's own
account, so the server enforces that account's permissions. No workspace, model, or
connector credential rides along on an MCP call.

## Where each class runs: the cloud service map

The execution layer's rule — Mesh integrates with a *contract*, vendors are
*backends* — extends to every class that needs infrastructure. Named products
below are the current best fits, recorded so the first implementation has a
target; none of them may be named above their layer's seam.

**Workspace backends.** The managed-sandbox-API backend that
[`sandboxed-execution.md`](/runtime/sandboxed-execution) ships first should be
**Modal Sandboxes**: it has a supported Go SDK (the control plane is Go, and a
sidecar in another language just to reach a vendor would be absurd),
gVisor-isolated containers, exec with streamed output, file access, sandbox
naming and tagging, idle timeouts, filesystem snapshots for the layered
cold-start design, and — decisive for the credential model — egress control up
to a domain allowlist, so default-deny-with-allowlist is a backend capability
rather than a Mesh-side fence. The roadmap's later backends are unchanged:
an edge container platform (Cloudflare Sandboxes/Containers) is the
sleep-and-wake `Resumable` case, Kubernetes is the own-perimeter case, local is
last and dev-only.

**Browser backends.** **Cloudflare Browser Run** (née Browser Rendering) is the
first target: managed headless Chrome on Cloudflare's network, a warm pool so
sessions start without cold-start tax, a raw CDP endpoint the control plane can
drive from Go, **Live View** (humans watch the session in real time), **human
in the loop** (a person can click, type a password, solve the wall the agent
hit, and hand control back), and **session recordings** as structured JSON.
Those last three are not conveniences; they map one-to-one onto Mesh
primitives — Live View links belong in progress emissions, human takeover is
`awaiting_input` wearing a browser, and recordings are the observation artifact
for an entire session. Cloudflare's Kitesurf browser (lighter, no full Chromium)
is a capability variant for read-mostly page work, not a separate integration.
Browserbase or Steel slot in as alternate backends behind the same contract.

**A browser backend is not itself a model-facing tool.** These are two halves of
one capability at different seams: the model receives generic
`browser_navigate`, `browser_read`, `browser_act`, and `browser_screenshot` tool
definitions; the runtime provisions a session and executes those calls through
whichever browser backend the agent was granted. Cloudflare belongs in the
browser-backend descriptor and catalog, never in a tool name or in the harness.

A vendor product may expose more than one integration shape without collapsing
the seams. Browser Run's stateless content and screenshot endpoints can back a
provider-hosted query or one-shot browser tool, while its persistent CDP
sessions back `BrowserBackend`. A shared vendor credential does not make those
the same runtime contract: installation may coordinate the credential, but each
call is still classified, gated, budgeted, and audited at the seam it uses.

**Query providers.** Web search and page-to-markdown fetch are hosted APIs
(Exa, Brave, Tavily, or similar) registered as `tool_providers`. The provider
is a row; the tool is `web_search`. No scraping from the control plane.

**Effect providers.** Slack actions ride the existing connector. GitHub, when
it arrives, is an app installation whose short-lived installation tokens are
exactly the minted-credential shape the execution layer already specifies —
one integration serving both `effect` tools (open an issue) and workspace
grants (clone this repo).

## Browser sessions

A browser session is a provisioned resource with the same discipline as a
workspace:

```
browser_sessions (
  id,
  run_id,               -- at most one live session per run
  agent_id,
  backend_id,
  external_ref,         -- opaque, backend-owned
  status,               -- pending | live | awaiting_human | released | failed
  live_view_url,        -- if the backend declares the capability
  recording_ref,        -- artifact pointer, written on release
  created_by_actor_id,
  session_seconds_used,
  created_at,
  released_at
)
```

And a backend contract that is an intersection plus declared capabilities,
exactly like execution backends:

```go theme={null}
type BrowserBackend interface {
    Capabilities() BrowserCapabilities   // live_view, human_handoff, recording, …
    Acquire(ctx context.Context, spec SessionSpec) (Session, error)
    Act(ctx context.Context, s Session, action Action) (Observation, error)
    Release(ctx context.Context, s Session) error
}
```

`Act` is deliberately coarse in the contract; the tool layer above it exposes
`browser_navigate`, `browser_read`, `browser_act`, `browser_screenshot`. Every
action is a persisted tool call, and the released session's recording is the
run's proof of what the agent actually did on the web — the browser equivalent
of the prompt snapshot.

The human-handoff flow is worth spelling out because no local-first harness can
have it: the agent hits a login wall, the run parks (`awaiting_human` on the
session, `awaiting_input` on the run), the thread gets a Live View link, *any
actor in the conversation* — not just the requester — opens it, types the
credential or clears the challenge, and hands back. The takeover is attributed
to the actor who did it, and the credential never transits Mesh at all.

## Budgets

Tool use adds no new budget philosophy, only new meters. The harness already
owns max iterations, tool calls, tokens, and wall clock; the execution layer
adds compute-seconds. Browser sessions add **session-minutes**, and query
providers add **per-provider call ceilings** (search APIs bill per request, and
a retry loop against a paid API is a money leak with no CPU signature).
Everything lands in the same place: exhaustion is a terminal status with a
reason, surfaced to the channel, and drawing more graph nodes still cannot
manufacture more budget.

## Progressive tool loading

An agent's tool schemas are not free context. They are priced into the same
input budget as its memory and its history (`internal/turn/context_budget.go`
subtracts the whole request estimate, tool schemas included, and gives the
system prompt what is left), so an agent connected to three MCP servers spends
part of every turn carrying capabilities it is not using — and loses memory and
history to pay for them.

So a turn starts with the agent's own verbs (memory, progress, commitments)
plus one tool, `tool_search`, and loads an integration's schemas only once the
model has asked for them. A search returns the matching operations and offers
their schemas on the next model request.

**Whether a turn does this at all is measured, not configured.** It is not free:
`tool_search` carries a bounded summary of the connected integrations in its
description, and that travels on every call. So Mesh compares the two costs —
carrying every schema, against carrying the index plus the handful a search
loads — and an agent on the wrong side of that line is simply offered
everything, as it always was. An agent with three tools carries all three. An
agent with forty carries an index. Neither needs an operator to have known
which it was.
(Callable is the wrong word, and the distinction is the whole design: the
executor resolves against the full authorized registry, so an operation the
operator granted was always callable. What a search changes is whether the
model has been shown how to call it.)

What this changes is what the model is **shown**. What it may **do** is
untouched:

* The searchable index is built from the definitions the turn was already
  authorized to use, so anything reachable through search is something the
  operator already granted. There is no second authorization path.
* Execution resolves against the full authorized registry, not the loaded
  subset, so a call already in flight cannot be invalidated by a later
  eviction. Offer-time policy, reviewed MCP fingerprints and call-time
  verification are identical either way.
* An agent below the threshold behaves exactly as it did before. Nothing is
  stored on progressive loading's behalf either way.

Loaded schemas are bounded by a share of the input budget. When a search would
exceed it, the least recently used tools are unloaded and the model is told
which, because a model that is not told will call one of them. Re-using a tool
refreshes it, so the tool in active use is the last to go. A single tool larger
than the whole budget is kept rather than evicted into an empty set.

Activations are recorded in the continuation checkpoint, so a paused task does
not have to rediscover what it was using. They are a request to re-activate,
never a grant: a resumed turn looks each one up in the authority it has *now*
and refuses anything revoked or redefined while the task waited, telling the
model what it lost rather than letting it plan around a tool it no longer has.

The summary in that description is why a model can search at all: it cannot
look for a capability it does not know exists. It is also the whole cost of the
mechanism, which is why the threshold above exists and why it is expressed in
tokens rather than in a tool count — a handful of very large schemas is worth
loading on demand where a handful of small ones is not.

On a synthetic agent with 20 connected operations the first call carries the
index instead of every schema, and `internal/turn/tool_search_test.go` logs
both figures rather than asserting a single flattering one.

## Ordering

The roadmap's sequencing survives contact with this document, and the memory
slice sharpens the first steps. What matters: the durable Run lands before the
loop gets more rounds, and the graph substrate lands before exec-class tools
make loops long and stateful ([`../roadmap.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap.md), Phases 3–3.5).

1. **Generalize the memory rounds into the harness loop.** Runs and RunSteps
   from [`agent-harness.md`](/runtime/agent-harness), each tool round as a persisted
   step, `tool_calls` and observations as rows, the budget set enforced, and the
   memory verbs become the first registered tools instead of a hardcoded special
   case. This is also where the tool package format lands: the memory verbs move
   into `tools/<name>/{manifest.toml, hooks}` folders crawled at startup, so the
   registry is built from manifests rather than hand-registered in Go. No new
   infrastructure vendors.
2. **Query tools.** `tool_providers` + `agent_tools` tables, the gating
   policy, and hosted `web_search` as the first provider-backed tool. This is
   the cheapest slice that makes agents visibly smarter, and it exercises
   registration, gating, and observation persistence with zero sandbox risk.
3. **The execution layer with Modal as the first backend**, per
   [`sandboxed-execution.md`](/runtime/sandboxed-execution): workspace tools
   (`exec`, `read_file`, `write_file`), lazy provisioning, egress allowlist,
   minted repo credentials, compute budgets. This is the slice where an agent
   first writes and runs code.
4. **Browser sessions on Browser Run**: the backend contract, the four browser
   tools, Live View in progress emissions, human handoff through
   `awaiting_input`, recordings as artifacts.
5. **Effect tools and remote MCP**: every effectful `workspace` or `effect`
   call crosses the commit boundary before it runs, and a resumed run past that
   boundary reports rather than re-executes; the remote MCP client so operators
   can attach third-party capability, authenticated as the agent's own account,
   without Mesh shipping code.

Slices 1–2 can land against loop-mode runs while the WorkGraph substrate is
still being built; they neither assume nor preclude graph execution. Slices 3–5
inherit graph-compatibility from the invariants they build on — workspace-
per-run and session-per-run mean parallel nodes cannot share mutable state by
construction.

## Design guardrails

* do not require executing a tool's code to read its manifest — discovery,
  rendering, and enablement are pure reads of `manifest.toml`; execution happens
  only through `Definition.Execute` after an explicit grant
* do not add a tool by editing a central switch — a tool is registered because its
  `tools/<name>/` folder and manifest were crawled at startup
* do not dynamically link third-party tool code into the control plane — the
  future plugin loader is out-of-process (subprocess/RPC or WASM), never Go
  `plugin`
* do not execute exec-class tool calls in the harness process; do not fetch
  arbitrary URLs from it either
* do not offer a tool the policy forbids — gate at offer time, not call time
* do not infer a tool's effect class from behavior; declare it, and treat a
  class violation as a security bug
* do not configure a tool, provider, or backend through process environment
* do not let a query tool read outside its RuntimeAgent perspective
* do not model a browser session as a workspace, or bill it like one
* do not import an MCP tool without an operator-assigned class and policy
* do not run local stdio MCP servers in the control plane
* do not put a per-call human approval in front of an effect — authority is the
  permissions on the agent's own account in the external system, and a gate
  that substitutes for scoping is borrowed-credential thinking
* do not let a human handoff become a private modal — `awaiting_input` and
  `human_gate` are attributed thread messages any actor can answer
* do not write a first-party package for a third-party API with no
  Mesh-primitive coupling — that is what a CLI in the workspace or remote MCP
  is for, and both act as the agent's own account
* do not add CLI teaching to Go string literals beyond what `codeWorkSection`
  already carries — new CLI knowledge waits for the skill loader, and the
  existing section moves into a `SKILL.md` in the change that builds it
* do not store an MCP provider's credential on the provider row — it is the
  agent's own login and lives on the `(agent, provider)` binding, or two agents
  sharing a server would act as one account
* do not store a tool's raw output in model context when it belongs in an
  artifact
* do not let a provider name appear above its layer's seam
* do not let per-request provider billing escape the budget set


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.