> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Sandboxed execution

# Sandboxed execution

**Strategy update (2026-09-18):** The
[tool integration strategy](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/tool-integration-strategy.md) defines the planned shared
catalog, named connections, progressive discovery, and Add capability flow. The
[delivery plan](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap/tool-integration-delivery.md) tracks implementation.
Descriptions of existing behavior below remain the current contract until those
changes ship.

A tool call and the process that made it must not be the same object.

That is the whole document. Everything below is the consequence of taking it
seriously.

The obvious design executes tools in the harness process, on the harness host,
in whatever working directory the harness happens to be sitting in. It is
obvious because it requires no design at all — the filesystem is already there,
the credentials are already in the environment, and `exec` is one line. It is
also the design that produces every one of the following, and they are not
independent bugs:

* two concurrent runs corrupt each other's working directory
* a destructive command, or an instruction smuggled in through content the agent
  read, executes with the harness's full ambient authority
* concurrency is bounded by one machine, and one heavy run degrades every agent
  on the box
* nothing about the execution is a row anywhere, so nothing about it can be
  audited, resumed, or explained

Mesh gives execution an explicit boundary: a **workspace**, owned by exactly one
run, holding no credential it was not specifically granted, materialized by a
pluggable **execution layer** rather than by whatever host the harness happens
to be running on.

## Two wins, and they are not the same win

It is tempting to say "run it in the cloud" and consider all four problems
solved. They are not, and conflating them produces a system that has paid for
isolation without receiving it.

**Collisions are not caused by locality.** They are caused by a shared mutable
working directory. Two runs in the same directory corrupt each other whether
that directory is on a laptop or a rented worker. Moving execution to a remote
sandbox while still sharing a directory relocates the collision somewhere
harder to observe.

So the two properties are separate and both are required:

* **Remote execution** buys the security boundary, horizontal scale, and host
  integrity.
* **Workspace-per-run** buys the absence of collisions.

Remote sandboxes make the second property cheap, because a fresh sandbox is a
fresh filesystem. But it is only guaranteed if the workspace is a persisted
entity with an owner, rather than an implicit side effect of where a process
happened to run.

## The execution layer is a separate concern

Mesh does not integrate with a sandbox vendor. Mesh has an **execution layer**,
and a vendor is a backend that implements it. The first shipped backend is a
managed sandbox API; it is a backend, not the design.

That distinction is the architecture, not a hedge, so it is stated before the
rest of the design rather than as a footnote at the end. Nothing above the
execution layer names a backend. Tools do not know where they run. The harness
does not know where they run. Which system executes a command is one row of
configuration, and every other part of Mesh reads a declared capability rather
than a product name.

The backends this is expected to serve are genuinely different from each other,
which is why the boundary has to be designed now rather than extracted later:

* a managed sandbox API — the default, and what ships first
* a Kubernetes cluster, for an operator who already runs one and wants agent
  workloads inside their own perimeter
* an edge container platform, where sandboxes are addressed by a caller-supplied
  identity and can sleep and wake rather than being created and destroyed
* an isolate runtime, for the narrow case of running a snippet with no shell,
  no package installs, and no filesystem worth the name
* a local backend, for people working on Mesh itself

These do not share a cost model, a startup latency profile, a snapshot story, or
a network model. Kubernetes has no equivalent of a filesystem snapshot handle; a
managed sandbox API has no equivalent of a namespace; an isolate runtime has
neither and cannot run a build at all. So the contract cannot be the union of
everything each backend can do — that is how an interface ends up shaped like
whichever backend was written first.

Two of those entries are worth naming as specific pressure on the design, because
they are the cases most likely to break a contract that was written against one
backend:

**Lifecycle is not universally create-and-destroy.** Some platforms address a
sandbox by an identifier the caller chooses, and let it sleep and wake with its
filesystem intact. That is a genuinely different lifecycle from "create,
do work, terminate," and it is *better* for a conversational agent: a thread that
comes back to life an hour later can wake the same workspace rather than rebuild
it. So the contract must not assume that terminate means destroy. It means
*release* — the backend decides whether that is destruction or hibernation, and
reports which via capability. [`interrupt-model.md`](/runtime/interrupt-model) already establishes that
a conversation can resume long after the work paused; a lifecycle that only knows
how to destroy would throw away exactly the state that makes resumption cheap.

**Some backends cannot host the workload at all.** An isolate runtime that runs a
JavaScript or Python snippet is a legitimate backend for "evaluate this
expression" and a non-starter for "run the test suite." This is the case that
proves capabilities have to be declared rather than assumed: a backend that
cannot offer a shell is not a broken sandbox backend, it is a different capability
set, and the layer above must be able to see that before routing a tool to it
rather than discovering it from a failed exec.

Instead:

* The **contract is the small intersection**: acquire a workspace, exec in it,
  read and write files, release it. Every backend must implement all of it.
* Everything else is a **declared capability** the backend advertises, and the
  layer above degrades legibly when it is absent. Snapshotting is the clear
  example: a backend that cannot snapshot project state is slower to start and
  otherwise correct, and Mesh must say so rather than fail.
* **Semantics travel with the contract, not the backend.** Workspace-per-run,
  credential scoping, configurable egress, and budget enforcement are Mesh's
  invariants. A backend that cannot enforce one of them declares that it cannot,
  and it is a policy decision whether that backend is allowed to run a given
  agent. "Isolation" must never quietly mean something weaker because of which
  backend is configured.

```
execution_backends (
  id,
  kind,               -- managed sandbox api | kubernetes | local | …
  name,
  capabilities,       -- snapshots, egress control, gpu, persistent volumes, …
  credential_ref,     -- encrypted, owned like a connector credential
  status,
  created_at
)
```

A backend is configured the way a connector is configured: at runtime, in the
database, through the UI — never through process environment. That is the same
rule [`agent-provisioning.md`](/agent-provisioning) already establishes for agents and their
credentials, and for the same reason.

## The workspace

A workspace is a row, created for a run, materialized by the execution layer:

```
workspaces (
  id,
  run_id,                -- at most one workspace per run
  agent_id,
  backend_id,            -- which execution backend materialized this
  environment_id,        -- the declared environment this materializes
  image_ref,             -- base image: the toolchain layer
  snapshot_ref,          -- project-state layer, if the backend supports it
  external_ref,          -- the backend's own handle, opaque to everything above
  status,                -- pending | provisioning | ready | running | released | failed
  created_by_actor_id,   -- attribution
  compute_seconds_used,
  created_at,
  terminated_at
)
```

`external_ref` is deliberately opaque. Nothing outside the backend's own
implementation is allowed to parse it, which is the mechanical guarantee that no
caller has quietly grown a dependency on a particular vendor's identifier
format. It is also not assumed to be *assigned* by the backend: a platform that
addresses sandboxes by a caller-chosen identity is served by letting the backend
decide what goes in this column, rather than by modelling every backend as
returning a handle it invented.

`released` rather than `reaped` is deliberate. Mesh's invariant is that a
workspace stops being reachable by its run, not that a container is destroyed.
A backend that hibernates instead of destroying satisfies the invariant, and
saying `reaped` in the schema would have quietly encoded one backend's lifecycle
as the universal one.

### Provisioning is lazy

A run does not get a workspace. A run gets a workspace **when a tool first
reaches for one**.

This matters more in Mesh than it would in a coding-agent harness. Mesh lives in
Slack threads, and most turns in a Slack thread are conversational — a question,
a status check, a clarification. Those turns touch no filesystem at all.
Provisioning a sandbox for them buys nothing and costs startup latency on every
single reply, which is exactly the wrong tax on the interaction Mesh is
supposed to be good at.

So the status column starts at `pending` and most runs never leave it. The
useful consequence is that execution cost scales with *work* rather than with
conversation volume: a busy channel full of chatter is nearly free, and a run
that compiles something pays for what it used.

### One per run, and no more

At most one workspace per run, enforced by the schema rather than by convention.
Sub-agent work in Mesh is a child run, so a sub-agent gets its own workspace
with no additional machinery — which is precisely the property whose absence
causes sub-agents to trample each other in a single-directory harness.

`created_by_actor_id` is not decoration. [`philosophy.md`](/philosophy) commits to every
agent turn being explainable, and "which human's request caused this container
to exist" is part of that. Execution is the most consequential thing an agent
does; it does not get to be the one part of the system with no attribution.

## The loop stays in the control plane

The naive reading of "sandbox the agent" is to put the agent in the sandbox. Do
not do that.

Mesh splits it:

* **Control plane** owns the harness loop, run and step persistence, model
  calls, outbound connector emissions, and all database writes.
* **Sandbox** is the *effect surface*: shell commands, file reads and writes,
  git, dependency installs, tests, builds, linters.

The reasoning, in order of importance.

**It preserves the state machine.** [`agent-harness.md`](/runtime/agent-harness) makes a run durable
by committing each iteration as a transaction, and [`interrupt-model.md`](/runtime/interrupt-model)
builds coalescing, supersession, and `awaiting_input` on top of that durable
state. All of it depends on the loop's state living in Postgres. Move the loop
into an ephemeral container and every one of those properties becomes a
distributed-systems problem that was previously a transaction.

**It keeps the interesting credentials out of the blast radius.** A workspace
that only runs commands needs no model provider key, no Slack bot token, and no
database connection string. The credential surface inside the boundary is the
smallest it can be, by construction, rather than by discipline.

**It keeps the execution layer thin enough to have more than one
implementation.** This is the part that matters for pluggability. A backend that
only has to run commands and move files can be written against Kubernetes, a
managed sandbox API, or a local process in comparable effort. A backend that has
to host the agent loop must also reproduce durable run state, model access, and
outbound delivery — at which point there will only ever be one backend, whatever
the interface claims.

**It costs approximately nothing.** An exec round trip to a warm sandbox is
milliseconds. The model call it sits between is seconds. The split is invisible
in the latency budget.

**Cancellation safety is unaffected.** `RunStep` already carries a replayability
classification, and an effectful step is still effectful when it happens in a
sandbox. A step that pushed a commit cannot be silently redone regardless of
where the push ran from.

## Credentials are declared and minted, never dumped

Credential delivery follows the [agent-owned identity guardrail (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap.md#agent-owned-acting-identities):
an acting credential belongs to the agent running the workspace. A delivery grant
cannot authorize another agent’s account or a pooled bot identity. Modal/Fly
backend credentials configure shared infrastructure; they are not guest identities.

The rejected design is to pass the agent's full credential set into the sandbox
as environment variables. It is the natural first instinct, and it gives away
the entire benefit.

The threat is not the agent running a destructive command. A sandbox contains
that. The threat is **an instruction arriving through content** — a file in a
cloned repository, an issue comment, a fetched page, a dependency's postinstall
script — and the resulting action being *exfiltration*. A sandbox does not
contain exfiltration of a keyring that was handed over at boot, and
exfiltration is the failure that cannot be rolled back.

So:

* An **environment declares** what it needs: which repositories, which
  registries, which named secrets.
* Mesh **mints** short-lived, narrowly scoped credentials per run.
  * Repository access is a scoped installation token for the specific
    repository, with a short expiry. Not a long-lived personal token.
  * Everything else is reached through a per-run token that the sandbox
    exchanges against a broker endpoint in the control plane. Every secret read
    is a persisted event tied to the run and the actor who triggered it.
* The declaration is persisted, so *what could this run see* is a query rather
  than an investigation.

The ergonomics for the agent are unchanged: it still finds what it needs where
it expects. The difference is that the grant is bounded, expiring, and
attributable.

## Workspace network access

Agent workspaces have **unrestricted outbound network access by default**. Mesh
cannot predict which HTTP APIs or hosts an open-source installation needs.
Filesystem/process isolation and credential permissions remain separate from
network policy: open egress does not grant additional credentials, but code in
that workspace can send any data or credentials it already holds to arbitrary
reachable destinations. Provider-level network restrictions still apply.

Operators can opt into `mode: "allowlist"` under **Settings → Agent runtimes →
Workspace network access** (`GET/PUT /api/settings/workspace-egress`). Restricted
mode allows Mesh's built-in source hosts and package registries
(`workspace.DefaultEgressDomains`), operator-specified domains, and the Mesh
broker host. It is default-deny outside that effective list. The UI shows the
built-in hosts and defaults to refusing incapable backends when restriction is
enabled; an operator may explicitly accept open egress on those backends instead.

In open mode, the workspace spec carries no `execution.EgressPolicy`, including
on a backend with `Capabilities.EgressControl`. In restricted mode, Modal enforces
the domain allowlist and an empty CIDR allowlist at sandbox creation; raw IPs
cannot bypass it. On an incapable backend such as Fly Sprites, `uncontrolled:
"refuse"` rejects acquisition before a sandbox is created; `"allow"` accepts open
egress. That fallback setting is dormant in open mode. The policy follows both
warm restores and cold retries.

An install with **no saved policy** defaults to `mode: "open"`. A saved legacy
policy without `mode` retains its previous allowlist behavior on upgrade. Its
owner can explicitly switch to open mode. New API writes must specify the mode;
malformed or unreadable policies still fail acquisition rather than silently
widening access.

Policy changes apply to **newly created workspaces**. Existing running or held
(paused) sandboxes retain their creation-time policy until replaced. Updating
Mesh or saving this setting does not retrofit a live sandbox's network rules.
The credential-free document parser remains separately network-isolated; this
setting does not relax document-processing isolation or connector URL checks.

The allowlist bounds destinations, not what can be sent there. Some built-in
hosts accept uploads, so even restricted mode is not an exfiltration barrier by
itself. Scoped credentials and human review still matter.

Denied attempts are not yet recorded as structured events: a backend denial may
surface only as a connection error. A connection reset alone does not prove a
policy denial; inspect the configured policy rather than inferring one from the
hostname or asking for another token.

## Cold start is layered, not atomic

The objection to a fresh workspace per run is that a large repository would be
cloned every time. That objection is correct about a single atomic snapshot and
wrong about the design.

Workspace state has natural layers with different lifetimes:

* **Base image** — operating system and toolchain. Changes when the toolchain
  changes.
* **Project state** — the checked-out repository, installed dependencies, warm
  build caches. Changes on push.
* **Run artifacts** — build output, scratch files. Ephemeral by definition.

Snapshotting these together forces every toolchain update to discard warm
project state, and every code change to rebuild an image. So Mesh snapshots the
project-state directory independently of the base image and mounts it into a
fresh workspace at run start. A run begins by mounting a warm tree and fetching
a delta, not by cloning a repository.

This is a **declared capability**, not an assumption. Backends differ sharply
here: a managed sandbox API may expose directory-level snapshots, while a
Kubernetes backend would express the same idea as a volume and may not be able
to express it at all. A backend without the capability is slower on first touch
and otherwise correct, and Mesh reports that plainly instead of failing. This is
the clearest case of why the layer is an intersection plus capabilities rather
than one vendor's feature list.

This is the mechanism behind the question asked at environment creation: *what
does this agent need access to?* The answer is what gets warmed. A repository
registered once is fast on every subsequent run, and a toolchain bump does not
throw that away.

Maintaining a pool of pre-warmed sandboxes is a further optimization. It is
deliberately out of scope until measured start latency justifies it.

## Budgets extend to compute

[`agent-harness.md`](/runtime/agent-harness) commits to runs terminating legibly rather than
hanging. Compute is now a resource a run consumes, so it joins the budget set:

* max compute-seconds per run
* max concurrent workspaces per agent
* an idle reaper, because a crashed worker must not leak paid compute

Budget accounting lives above the backend, in Mesh, because it is a property of
the run rather than of the vendor. Backends report consumption; they do not own
the limit. Otherwise the same agent means two different things on two backends.

One workspace per run, not one per tool call, and none at all for a run that
never touches a file. Exhaustion is a terminal status with a reason, surfaced to
the channel, exactly like every other budget. "I stopped after twenty minutes of
compute" is a useful thing to say.

## The backend contract

The intersection every backend must implement:

```go theme={null}
type Backend interface {
    // Capabilities is read before use. Callers branch on declared capability,
    // never on backend identity.
    Capabilities() Capabilities

    // Acquire, not Create. A backend that addresses workspaces by a
    // caller-supplied identity may be waking an existing one; a backend that
    // only creates may ignore that and always build a new one. The caller does
    // not need to know which, and must not care.
    Acquire(ctx context.Context, spec WorkspaceSpec) (Workspace, error)

    Exec(ctx context.Context, w Workspace, cmd Command) (Result, error)
    ReadFile(ctx context.Context, w Workspace, path string) ([]byte, error)
    WriteFile(ctx context.Context, w Workspace, path string, b []byte) error

    // Release, not Terminate. The invariant is that the workspace stops being
    // reachable by its run. Whether that means destroyed or hibernated is the
    // backend's decision, reported via capability.
    Release(ctx context.Context, w Workspace) error
}
```

Optional behavior is a separate interface a backend may or may not satisfy, so
"can this backend snapshot?" is answered by the type system rather than by a
runtime error halfway through a run:

```go theme={null}
type Snapshotter interface {
    Snapshot(ctx context.Context, w Workspace, dir string) (SnapshotRef, error)
}

type EgressController interface {
    ApplyEgressPolicy(ctx context.Context, w Workspace, p EgressPolicy) error
}

// Resumable is declared by a backend that can release a workspace without
// destroying it and later reacquire the same filesystem. A conversational agent
// benefits directly: a thread that goes quiet for an hour can wake its workspace
// instead of rebuilding it.
type Resumable interface {
    Reacquire(ctx context.Context, ref ExternalRef) (Workspace, error)
}

// Shell is declared by a backend that can run arbitrary commands. An isolate
// runtime that only evaluates a snippet does not declare it, and the layer above
// routes exec-class tools elsewhere rather than discovering the gap from a
// failed exec.
type Shell interface {
    SupportsShell() bool
}
```

Three rules keep the seam honest, and all three are things that are easy to
violate accidentally and hard to undo:

**No caller branches on backend identity.** Not tools, not the harness, not the
UI. Anything that would need `if backend == "..."` is a missing capability
declaration, and the fix is to add the capability rather than the branch. This is
the single rule that keeps a pluggable layer pluggable, because it is always
cheaper in the moment to special-case one vendor.

**The contract is the intersection, not the union.** A backend that needs a
concept no other backend has expresses it inside its own configuration, not by
widening the interface everyone implements. An interface that grows to fit every
backend's best feature is how a second implementation becomes impossible to
write.

**Mesh's invariants are Mesh's, not the backend's.** Workspace-per-run,
credential scoping, configurable egress, and budget accounting live above the
backend. Where a backend cannot enforce one, it declares the gap and Mesh
decides by policy whether that backend may run a given agent. The word
"isolated" must mean the same thing regardless of what is configured underneath.

### Ordering: the local backend ships last

A local backend exists — requiring a third-party account before anything
executes at all is a real cost for someone running Mesh for the first time — but
it ships *after* a remote one, implemented against the same contract, and is
documented as unsafe for anything but working on Mesh itself. It still allocates
a distinct working directory per run, because the collision fix is not allowed
to regress on the development path.

The ordering is the point. Building local first and abstracting later produces an
interface shaped like a filesystem, and every remote backend then spends its
life pretending to be one. The first implementation of an interface defines it no
matter what the documentation claims, so the first implementation should be the
one with the most constraints.

## What ships today

The contract above is the design. This section is the honest inventory of what
is built against it, so a reader can tell a guarantee from an intention without
reading the code. Keep it current in the same PR that changes the code.

**Built.**

* The `Backend` contract (`internal/execution`) and its first, remote backend:
  Fly Sprites (`internal/execution/fly`). `Command.Dir` carries a working
  directory. The workspace layer delivers authorized credentials through a
  protected file before commands, without putting them in the create-time
  environment (see the credential bullet below).
* Backends as runtime configuration: the `agent_runtimes` settings section,
  configured from the dashboard or claimed by `MESH_FLY_API_TOKEN`, projected
  through the `execution.Provider` descriptor catalog
  (`internal/execution/backends`). Nothing above that catalog names a vendor;
  the workspace manager constructs whichever enabled backend it finds through
  the descriptor's `New`.
* **Which enabled backend a run uses (per-agent selection).** An
  install may enable more than one backend at once. The resolver
  (`internal/app/workspace_runtime.go`) chooses among the enabled set without
  ever branching on which vendor a name identifies:
  * **One enabled** → it is used, no configuration required.
  * **More than one enabled** → each agent may pin one via
    `runtime_agents.runtime_backend` (a catalog backend `Name`, set through the
    agent-detail page, which shows the selector *only* when more than one is
    enabled). An agent with no pin takes the **documented instance default**:
    the first enabled backend in catalog order, logged once at acquisition so an
    unpinned agent's implicit runtime is observable rather than looking random.
  * A pin naming a backend that is **not currently enabled** (turned off, or a
    name that predates a catalog change) is **refused** with an
    operator-actionable error — the resolver never silently falls through to a
    different vendor's runtime. The set of enabled backends is validated at the
    admin boundary too, so a pin can only be stored for a backend that is
    enabled at the time it is set.
  * **Disabling a backend un-pins its agents (the disable-time cascade).** When
    an operator turns a backend off, the same transaction that flips its
    `enabled` switch also NULLs `runtime_agents.runtime_backend` for every agent
    pinned to it (`ClearAgentRuntimeBackendPins`), and the save response reports
    the reset slugs so the settings page can say *"disabling `<backend>` reset N
    agent(s) to the instance default."* A reset agent falls back to the instance
    default on its next acquisition — it is never left pointing at a runtime the
    install no longer offers. This is the primary path; the *not-currently-
    enabled* refusal above is the defense-in-depth backstop for a pin that
    survives a clean disable (a race, a direct DB edit, or a backend that goes
    unhealthy without being disabled).
    The pin is a *selection among enabled backends*, never a credential: the
    backend's token remains an install-owned setting.
* **Release resolves by the row, not by the selection.** Selection answers
  where *new* work runs; ending an *existing* workspace uses the backend its row
  recorded (`workspaces.backend_kind`), looked up by name in the catalog
  (`workspace.KindResolver`, wired from `internal/app/workspace_runtime.go`).
  Both the per-run release and the process-wide reaper — which has no agent and
  sweeps every agent's rows — go through it, so a workspace an agent pinned to a
  non-default backend materialized is released by that backend rather than
  failing against the default and running on. A backend whose enable switch is
  off still releases what it minted while its credential is retained (turning
  off acquisition never turns off cleanup, the same rule credential revocation
  follows). Only a recorded backend with no catalog entry or no credential left
  is unreachable: the row stays live as evidence and the release reports that
  it must be released by hand. **Snapshot reclaim follows the same rule**: the
  reaper's snapshot sweep (`Reaper.SweepSnapshots`) lists every backend that
  has reclaimable snapshots (`ListReclaimableSnapshotKinds`, by
  `project_snapshots.backend_kind`), resolves each through the same
  `KindResolver`, and deletes that backend's artifacts through it — so a
  non-default enabled backend's retired snapshots are reclaimed instead of
  costing storage forever. The per-sweep row limit applies *per backend*
  (filtered before the `LIMIT`), so one backend's backlog cannot starve
  another's, and so is the sweep's time: kinds run in name order under the
  sweep's one deadline, each bounded to its fair share of what remains
  (unused time rolls forward), so a hung backend early in the order cannot
  spend the deadline and starve every later kind on every tick. An
  unreachable kind is skipped — warned about once per process per kind, then
  logged at debug on later ticks — and its rows stay unreclaimed as evidence; a backend without the snapshot capability has
  nothing to reclaim; neither blocks the other kinds. Resume verification and
  snapshot capture stay on the selection, because both are about continuing
  work there (a capture's own orphan cleanup deletes through the backend that
  just minted it). Nothing that maintains an existing workspace or snapshot row
  resolves the default backend alone.
* **Sized workspaces**: a vendor-neutral workspace tier (`execution.Size` —
  `small`, `standard`, `large`, `xlarge`, ordered smallest-first) that each
  backend maps to its own guest units at its leaf (`internal/execution/fly`,
  `internal/execution/modal`); nothing above the leaf names RAM or CPU. The
  ONLY human-facing knob is a per-install **ceiling** (the `max_size` config
  field on every backend descriptor, a closed enum validated at the form like
  the region enum) — size is a cost guardrail, not an agent-level functional
  dial. The tier a workspace actually asks for is the workspace manager's
  policy, clamped to the ceiling before the backend sees it.
  🚨 **The default tier today is a REVERSIBLE BRIDGE, not the end state.** It
  carries working headroom (`large`, above the `standard` build floor) so
  heavier builds — commongrid's `next build` is the one that forced it —
  succeed before OOM-aware auto-climb lands. Once auto-climb ships,
  the default drops to `small` (the smallest tier that boots) and headroom is
  reached by *climbing* tiers on OOM, not by a fat default; the enum is built
  for that reversal (a real `small` floor below the default, room above it up
  to the ceiling). `NODE_OPTIONS`/`--max-old-space-size` is explicitly out of
  scope — it is operator repo config, language-specific.
* The `workspaces` table (`0030`) and `internal/workspace`: one workspace per
  run enforced by `UNIQUE (run_id)`, the row written before the backend call,
  acquisition **lazy** on the first exec-class tool call, reuse within the run,
  release when the run ends however it ends, and a background sweep (at boot
  and every ten minutes, each sweep time-bounded) that releases workspaces
  whose run is already terminal or that outlived the longest possible run.
* The exec-class tools, on the crawlable package format: `workspace_exec`,
  `workspace_read_file`, `workspace_write_file` (`tools/workspace`, `Needs: ["workspace"]`), and the GitHub package's `checkout_github_repo` (`Needs: ["workspace", "secret:github_pat"]`), which clones a repository into the
  workspace, configures git to commit as the agent's GitHub account, and checks
  out a working branch. `create_github_pull_request` is the control-plane
  fallback for opening the PR when `gh` is not in the workspace image.
* A **progress notice** on the run's first workspace acquisition: one threaded
  message telling the requester a sandbox is being set up and a reply will
  follow, so a multi-minute coding run is not an hourglass and silence
  ([`agent-harness.md` §Progress emissions](/runtime/agent-harness#progress-emissions)).
* Explicit, agent-owned [credential delivery (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/workspace-credentials.md): a saved
  PAT alone does not grant sandbox access. A visible permission authorizes the agent's
  GitHub credential as `GH_TOKEN` and `GITHUB_TOKEN`, plus non-interactive
  guards for git and gh. Credentials are resolved from the agent's own secrets per
  acquisition and made visible to **every command** the workspace runs. The
  first live run proved a create-time environment does not reach a Sprite's
  exec sessions, and the backend's per-exec environment parameter rides the
  request URL (which vendor logs retain) and replaces the session environment,
  so neither is the credential path. Instead the workspace layer writes the
  grants once per workspace to a mode-0600 file under the sandbox user's home
  through the filesystem surface (a request body, not a URL) and prefixes each
  command with a fixed `sh -c '. "$1"; shift; exec "$@"'` wrapper that sources
  it; argv is never re-parsed. The token is therefore on the sandbox's disk for
  the run's lifetime — a deliberate trade: an agent that can run `env` can read
  either, while a credential in a third party's access logs is exposure nothing
  in the run can bound. Nothing else from the keyring goes in, git
  authenticates through a credential helper that reads `GH_TOKEN` at use time,
  and the workspace is destroyed at release. Permission changes and credential
  rotation invalidate active and held sandboxes; handles revalidate every operation
  and a background worker retries provider teardown. Previously copied long-lived
  tokens require revocation at GitHub.

**Not built, and what that means in practice.**

* **Minted credentials and the broker.** The grant is the agent's own personal
  access token for the run's lifetime, not a short-lived scoped installation
  token. Use a fine-grained PAT scoped to the repositories the agent works on.
* **Egress control on every backend.** Open access is the default. Opt-in
  allowlist enforcement is available on Modal; Fly Sprites declare
  `EgressControl: false`, so restricted installs must refuse that backend or
  explicitly accept its open egress. Denied attempts are not yet recorded as events.
* **Declared environments and the layered cold start.** Every run clones. There
  is no per-repository project-state snapshot yet; Sprites checkpoints are the
  obvious first implementation, and a warm clone per repository is the next
  meaningful latency win.
* **`Resumable`.** Release destroys the sprite. A conversation's next turn
  re-clones rather than waking the previous turn's workspace, so an agent must
  push its branch before the run ends — the tool descriptions say so.
* **Compute budgets.** A workspace's life is bounded by its run's wall clock
  and by per-command timeouts, not by a compute-seconds budget of its own.
* **Artifacts.** Large command output is clamped into the observation (head and
  tail kept, a visible marker for the middle) rather than stored whole.

## Design guardrails

* do not execute tool calls in the harness process
* do not share a working directory between runs, locally or remotely
* do not put the harness loop inside the sandbox
* do not hard-code an execution backend anywhere above the execution layer
* do not branch on backend identity; branch on declared capability
* do not widen the backend contract to fit one backend's best feature
* do not let "isolated" mean something weaker because of which backend is
  configured
* do not pass an agent's credential set into a sandbox wholesale
* distinguish filesystem/process isolation from optional network restrictions
* do not snapshot toolchain and project state as one unit
* do not provision a workspace for a run that never asked for one
* do not let a workspace stay reachable by a run that has ended
* do not assume release means destroy, or that acquire means create

### GitLab coding

[GitLab coding (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/gitlab-coding.md) extends this contract to GitLab.com HTTPS/PAT
workflows. GitHub and GitLab use the same checkout engine and independent
agent-owned delivery permissions. Both providers can coexist: Git credentials
are host-scoped and commit identities are repository-local. See
[sandbox credential delivery (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/workspace-credentials.md) for per-provider issuance,
revocation, and continuation behavior.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.