> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent harness

# Agent harness

The turn pipeline proves the loop: one message in, one model call, one reply out.
That is call-and-response. It is not an agent.

The agent harness is the layer that turns a single model call into a bounded,
resumable, multi-iteration run that can use tools, report progress, ask for help,
and, when the work genuinely decomposes, compile the objective into a durable
WorkGraph without losing the thread.

The relationship is simple:

> **The harness makes one loop reliable. The WorkGraph composes reliable loops.**

See [`work-graph.md`](/runtime/work-graph) for the graph contract.

## Why it exists

A useful agent turn is not one request. It is a loop:

* read the request
* decide whether a tool is needed
* call the tool
* read the observation
* decide again
* eventually produce a reply

Two things break when that loop lives entirely inside one in-memory function
call.

First, the run becomes invisible. There is no persisted state, so nothing can
report where the run is, nothing can resume it, and a worker crash silently
loses the whole turn. The user sees nothing and never learns why.

Second, the run becomes unattributable. Tool calls and observations are the most
consequential things an agent does, and if they only exist in a stack frame they
cannot be audited. [`philosophy.md`](/philosophy) says every agent turn is explainable.
An unpersisted tool-calling loop is the fastest way to make that untrue.

The harness exists to make long agent work durable, legible, and bounded.

### A precondition, stated once

Everything below makes a turn *more* capable and therefore *more* expensive:
tools, iterations, subprocesses, sandboxes, and eventually WorkGraph fan-out.
That is fine as long as one thing is true upstream of it:

> **A turn exists because someone addressed the agent, and there is at most one
> in flight per (agent, conversation).**

That is not this document's contract — it is [`addressing.md`](/runtime/addressing)'s —
but the harness is where its absence gets expensive. A harness this powerful, driven by a
router that treats every message in a busy channel as a request, is a fan-out of
tool-calling loops that ends with the host out of CPU and every co-tenanted agent
dead. Budgets below bound one run. They do not bound the *number* of runs.

WorkGraphs do not weaken this invariant. One conversation turn still owns one
root Run. Graph nodes are child executions *inside* that Run, not additional
conversation turns.

## Runs and run steps

A turn becomes a `Run`. Each iteration of a loop becomes a `RunStep`.

Run statuses:

* `queued` — accepted, not yet picked up
* `running` — a worker holds it
* `awaiting_tool` — a tool call is outstanding
* `awaiting_input` — the agent asked a question and is waiting on an actor
* `succeeded` — terminal, produced a final reply
* `failed` — terminal, produced an error
* `cancelled` — terminal, stopped by an actor
* `budget_exhausted` — terminal, hit a declared limit

Each iteration commits as one transaction: the step, its tool calls, its
observations, and the new run status land together. A worker that dies mid-run
resumes from the last committed step instead of starting over or disappearing.

This is the root fix, and it is worth being blunt about why. The failure mode we
are designing against is an agent that accepts a hard request and then goes
quiet for four minutes. That is not a messaging bug. It is a state bug: a long
turn held as one in-memory goroutine has no state to report progress *from* and
no state to resume *to*. Persist the run and progress reporting becomes a
read, not a guess.

This fix has a hard scope, and its boundary is the subject of another contract. A
durable Run makes *one turn* survive a crash. It does nothing for a promise that
outlives the turn — an agent that says *"on it, spinning up the implementation"*
and then reaches `succeeded` having created nothing has not crashed; the turn
succeeded and the promise evaporated. That is a Commitment, not a Run, and it is
owned by [`commitments.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/commitments.md). Do not expect the durable Run to
close it.

The same rows are the audit substrate. What one run step must capture — the
model call with its request, response, and token usage; the tool call with its
arguments and observation — and how an operator reads it back is specified in
[`observability.md`](/runtime/observability). Capture ships with the ledger, not as a
follow-up.

## Execution shape: loop or graph

A root Run has an execution shape:

```text theme={null}
loop   — one bounded agent loop directly owns the objective

graph  — one WorkGraph schedules bounded jobs for the objective
```

This is not a claim that every request deserves decomposition.

A short, coherent, sequential request should stay a loop. Conceptually that is a
one-node graph; the implementation does not need to materialize graph state for
it. Graph mode earns its complexity when work splits into independent branches,
needs explicit verification or fan-in, requires differently scoped execution
contexts, or needs a human gate before an effect.

A conservative classifier may use `model_fast` to decide whether graph planning
is worth attempting. The classifier decides **shape**, not topology. If it is
uncertain or fails, the safe fallback is `loop`.

When shape is `graph`, `model_primary` (or a future dedicated planning role) may
propose topology. The proposal is typed data and is handed to the deterministic
admission layer described in [`work-graph.md`](/runtime/work-graph). The model never
turns its own proposal directly into scheduled work.

A root Run may also begin as a loop and discover, before its external-effect
commit boundary, that the objective genuinely decomposes. It may then propose a
WorkGraph. That transition is persisted and explainable.

## Child runs are workers, not new agents

A model-backed WorkGraph node executes as a bounded child Run (or an equivalent
child execution unit carrying the same semantics).

The child Run exists to reuse the machinery this document already owns:

* durable RunSteps
* budgets
* cancellation
* tool-call persistence
* progress
* workspace acquisition
* failure state

It does **not** create a new `RuntimeAgent`.

A child Run:

* belongs to the same RuntimeAgent perspective as its root Run
* has no connector identity of its own
* does not appear as a coworker in Slack
* does not become an Actor
* does not own private long-term memory
* receives a bounded subset of the root Run's context and tools
* may never have greater authority than its parent

This is how Mesh gets fresh specialist execution contexts without violating the
perspective boundary in [`../perspective.md`](/perspective).

V1 child Runs execute bounded loops; they do not recursively create private
WorkGraphs. A worker that discovers more decomposition requests a root-graph
revision. That keeps all fan-out visible to one scheduler and one budget.

## Acking that the run is working

Ingress enqueues the root Run and returns. It does not wait for the model.

The ack is a **reaction on the triggering message**, not a message of its own: ⏳
the instant the run is admitted, swapped for ✅ or ❌ when it settles. It goes on
immediately, for every run, because a reaction is unobtrusive enough that there
is no cost to being early and no need to predict whether the run will be slow.

Mesh previously acked with a threaded *on it, this will take a bit* **reply**,
posted by a timer at a five-second deadline. That reply is retired: the
reaction says the same thing sooner and without spending a message in the thread.
Its deadline reasoning is what survives and is worth keeping for any FUTURE
message-shaped progress signal — a deadline cannot be wrong the way a prediction
can, because it observes that a reply has not landed rather than guessing that it
won't. A threaded message earns its place only for work long enough that a human
would otherwise wonder whether the agent died, and that is the only thing it
should ever be used for.

Graph planning and graph execution are covered by the same root ack. Child Runs
do not emit their own conversational acks; they are implementation details of
one user-visible objective.

## Complexity classification

The fast-model classifier is an optimization on top of the deadline, not a
replacement for it. It has no bearing on the reaction, which is unconditional;
what it triages is the message-shaped progress signal described above and the
execution shape.

It runs in parallel with the root Run, and its jobs may include both expected
duration and execution-shape triage:

* obviously long work — set the deadline to zero and say something immediately
* obviously trivial work — suppress the message entirely
* plausibly decomposable work — allow graph planning
* anything else — leave the default deadline and loop execution alone

Typed output, so a bad classifier response degrades to the default behavior
instead of corrupting the run:

```go theme={null}
type Complexity struct {
    DurationBucket string  // instant | seconds | minutes | long
    NeedsTools     bool
    ExecutionShape string  // loop | graph_candidate
    AckText        string
    Confidence     float64
}
```

`graph_candidate` is deliberately not `graph`. The classifier may nominate a
request for planning, but the planner still has to produce a useful DAG and the
runtime still has to admit it.

`AckText` is the sharp part. The classifier writes the ack copy in the same call
it does triage, and the harness stashes it on the root Run. When the deadline
fires, delivering a specific, context-aware ack costs zero additional model
latency — the sentence is already sitting there.

If the classifier is slow, errors, or is not configured, the deadline mechanism
still works. That ordering is deliberate: the mechanism must be correct on its
own, and the model only makes it feel better.

## Model roles

Primary and fast remain the main per-agent text models:

* `model_primary` — reasoning, tool use, final replies, and initial WorkGraph planning
* `model_fast` — quick drafting, extraction, summaries and compatibility classification
* the independent installation **decision model** — fixed-choice classification
  through a native decision API; see [decision-models.md (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/decision-models.md)

This is a first-class part of the agent contract, not a classifier
implementation detail. The fast model is reused for:

* complexity triage and ack copy
* execution-shape candidacy
* conversation titles
* context compaction
* routing decisions

Use the decision role for existing engagement, continuation consent and memory
sensitivity classification. Keep generation and open-ended extraction on the fast
model; evaluate new bounded decision use cases independently.

Do not add `model_planner` merely because WorkGraph exists. Measure first. If
planning develops materially different capability/cost requirements, it can
become a model role later without changing the graph contract.

## Tool contract

```go theme={null}
type Tool struct {
    Name        string
    Description string
    InputSchema json.RawMessage
    Invoke      func(ctx context.Context, input json.RawMessage) (Observation, error)
}
```

Tools are registered per agent and gated per connector, because the same agent
in a private channel and in a shared one should not necessarily have the same
reach. [`tools.md`](/runtime/tools) is the normative contract for tool effect
classes, registration, gating, and the cloud services each class executes on.

Exec-class tools — shell, file writes, git, tests, builds — do not run in the
harness process. They run in a sandboxed workspace materialized by a pluggable
execution layer, and the loop above stays in the control plane.
[`sandboxed-execution.md`](/runtime/sandboxed-execution) is the normative contract for that boundary,
including the credential model, which is a grant per run rather than the agent's
full keyring. No tool names an execution backend.

For graph mode, an `agent_loop` node that needs exec-class tools runs as a child
Run and therefore gets its own workspace. Parallel WorkGraph nodes never share
one mutable working directory merely because they share a root Run. A later
`join` node owns reconciliation.

Every tool call and every observation is a persisted event. That keeps runs
explainable after the fact and keeps attribution intact: a tool ran because a
specific actor said a specific thing in a specific thread, through a specific
root Run and, when graph mode is active, a specific WorkNodeAttempt.

## Budgets

Runs must terminate legibly rather than hang:

* max iterations
* max tool calls
* max tokens
* max wall clock
* max compute-seconds (see [`sandboxed-execution.md`](/runtime/sandboxed-execution))

Exhausting a budget ends the current allowance, not the user's task. For
human-requested work, the final completion explains actual progress and what
remains, then asks whether to continue. An affirmative reply from the original
requester authorizes one fresh bounded run with the saved task context. Silence,
a refusal, or an ambiguous reply never authorizes more task work. Runtime
counters and internal state labels belong in the ledger, not the conversation.
See [human-authorized continuation (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/task-continuation.md) for the implemented
consent, checkpoint, and workspace lifecycle.

Graph mode adds composition rather than a new source of money. The root Run owns
an aggregate budget; WorkNodes receive bounded slices of it. Drawing six nodes
cannot turn one Run's budget into six Run budgets.

WorkGraph-specific ceilings include node count, depth, parallelism, revisions,
replans, and aggregate model/tool/compute spend. See [`work-graph.md`](/runtime/work-graph).

The budget that is not per-root-run — how many root Runs may exist at once, for
one agent and for one install — is admission control, and it lives with the
addressing/concurrency contract in [`addressing.md`](/runtime/addressing). Graph
fan-out is additionally constrained by those same per-agent and per-install
ceilings. Ten child Runs each perfectly within their node budget can still
saturate a host.

### Tool results are compacted within a run

The loop re-sends the whole conversation on every completion, so a run that
makes real tool calls — a checkout, a dozen test runs, several file reads — would
otherwise grow its request until the model's context window failed it, well
before any budget above fired. `internal/turn/compaction.go` bounds that: before
each completion, the newest few tool results stay verbatim (they are what the
model is reasoning about), older results are kept verbatim up to a byte budget,
and everything past it is replaced in place by a short excerpt — the result's
first and last lines, its size, and a pointer to the run ledger, where the full
text lives in `tool_observations`.

This is deliberately **mechanical**, not a model-written summary. It touches only
the current run's own tool output, never conversation history, memory, or
attribution, so it is not the compaction Track A3 gates on an attribution eval;
the elision is visible to the model the way the workspace tools' output clamps
are, and it is safe for the same reason. It also does not change the quadratic
shape `internal/run/budget.go` describes — tool calls, assistant turns, history,
and the excerpts themselves still accumulate — it shrinks the per-round delta
enough that a coding run finishes on its budget instead of at the provider.

## Progress emissions

Progress is emitted as events, like everything else outbound:

* the ack
* periodic progress while the run is long
* the terminal result

All emissions are idempotent and keyed on the logical request/root Run, so a
worker retry cannot spam a thread with duplicate acks.

Graph mode gives progress a much better substrate. Instead of inventing progress
copy from an opaque loop, Mesh can summarize durable WorkGraph facts: three of
five branches complete, one verifier running, one branch blocked. Child Runs do
not independently narrate themselves into the channel unless the root Run's
progress policy chooses to surface them.

Slack-specific: prefer a reaction lifecycle for medium-length work — eyes when
the run starts, brain while it is thinking, check when it lands. Post a real
threaded message only for runs long enough that a human would otherwise wonder
whether the agent died. Channel noise is a cost, and the reaction lane pays less
of it.

The ⏳/✅/❌ reaction covers every turn. Contextual progress starts automatically
for every active agent on new and resumed executions. No per-agent configuration
is needed; existing settings cannot disable it.

`internal/progress` supervises the active turn independently of model/tool
calls. It sends an initial update after 60 seconds of silence and subsequent
updates at three-minute intervals. Agent `report_progress` checkpoints can
arrive sooner for a plan, milestone, replan, or wait, with a one-minute minimum
gap after the previous attempt. The supervisor coalesces candidates and uses a
bounded fast-model summary when the agent has not supplied an update. A slow or
unavailable summarizer falls back to honest activity/waiting facts.

Workspace acquisition no longer sends a fixed sandbox message. Its first
successful acquisition feeds the same progress supervisor, which uses the
agent's checkpoint or the task objective and current activity to compose copy.
Repeated acquisitions do not create additional acknowledgements.

Progress is an additive `run.EffectProgressNotice`, with a durable reservation
per logical subject/sequence. Ordinary runs use their run ID; commitments share
cadence across continuation attempts. Progress never settles the task, consumes
the final reply boundary, or counts as completion evidence. Ambiguous delivery
is recorded and never blindly replayed. This is at-most-once application
dispatch, not a guarantee of exactly-once remote receipt.

The runtime stops and joins the supervisor before sending a result or failure
reply. Cancellation and ownership checks prevent stale attempts from sending;
a remote request already admitted before cancellation cannot be recalled.
Progress copy and receipts are retained in `progress_events` and the dashboard
trace, separately from ordinary assistant history and automatic memory
extraction. Follow-up context may include three recent progress records,
explicitly labeled and subject to live source/disclosure checks; their lineage
also propagates into any resulting reply.

See [the implementation plan (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/agent-progress-updates-plan.md) for the design,
acceptance criteria, and implementation boundaries.

## `awaiting_input` is the differentiator

Mid-run, an agent can ask a clarifying question and park the run in
`awaiting_input`.

In a single-user harness that is a modal prompt: the one person who asked has to
be the one who answers.

In Mesh it is a thread. **Any actor in that conversation can answer and resume
the run**, and because every message already carries an actor, the resumed run
records who unblocked it. The clarification is part of the conversation graph,
not a side channel.

Graph mode can also contain a `human_gate` WorkNode. The distinction is semantic:

* `awaiting_input` means the agent lacks information needed to continue
* `human_gate` means the runtime requires an attributed decision before a
  declared edge, usually because the next work is effectful or expensive to undo

Both pause durably. Neither becomes a private modal prompt outside the thread.

That is the capability the multi-actor model buys us, and no harness built
around a single conversation partner can express it.

## First slice

Keep the first implementation boring:

* one root Run
* execution shape `loop`
* one step, one model call, no tools
* ingress enqueues, a worker executes, the reply routes back as it does today
* the ack timer wired but with no classifier

That proves the durable Run contract before WorkGraph fan-out exists.

Then implement the graph substrate before broad exec-class tools:

* `loop | graph` execution shape
* WorkGraph proposal + deterministic admission
* DAG scheduling
* child Runs
* typed artifacts
* graph budgets and revisions

A graph made of unreliable loops is simply a more parallel failure mode. The
ordering is deliberate.

## Interrupt-driven turns

Everything above assumes a request stops arriving once the run starts. It does
not.

> The normative contract for this — the mailbox, the per-agent config surface,
> and the commit-boundary state machine that decides which behaviors are legal —
> lives in [`interrupt-model.md`](/runtime/interrupt-model). This section is the harness-level
> narrative: why the problem exists and what it costs. Where the two appear to
> disagree, [`interrupt-model.md`](/runtime/interrupt-model) is authoritative.

Today rapid-fire messages are processed serially: thirty seconds on the first,
then thirty seconds on the second. That is frustrating, and it is usually also
wrong. Messages in quick succession are normally one request — a follow-up
adding context, or a correction — not two independent ones. Answering the first
in isolation spends a full run producing a reply the person has already moved
past.

So the harness needs a position on what happens to in-flight work when the
conversation keeps going.

Graph mode does not create a second interrupt model. New conversational input
still targets the **root Run**. The root determines whether to amend context,
preempt speculative work, enqueue the input, or revise the WorkGraph. Child Runs
never independently consume messages from the conversation mailbox.

### Debounce before interrupt

The cheapest large win is not cancellation. It is a quiescence window before a
run starts.

Hold a newly-enqueued turn for a short quiescence interval and let messages that
arrive inside that window join the same run. Four rapid messages become one run
with all four in context. [`interrupt-model.md`](/runtime/interrupt-model) specifies the window as
per-agent config; a low-single-digit-seconds default is a reasonable starting
point for a conversational agent.

This is strictly better than interrupting, because there is nothing to
interrupt:

* no wasted inference
* no restart cost
* no cancellation semantics to get right

The window is configurable per agent (`interrupt.coalesce_window`) and can become
adaptive later — a message that ends mid-thought suggests more is coming — but
the fixed window already handles the majority of real cases. It ships first.

### Cancellation safety is a property of steps and attempts

This is the correctness property. Everything else in this section is
experience.

Whether in-flight work can be abandoned does not depend on the root Run alone.
It depends on what execution already did. Inference and pure reads are safe to
cancel and redo. A step or WorkNodeAttempt that already had a side effect —
posted a Slack message, created an issue, pushed a commit — cannot be silently
restarted, because restarting it duplicates it.

So `RunStep` carries a replayability classification, and graph mode extends the
same classification to `WorkNodeAttempt`. The root Run's first externally
effectful step or node attempt moves it from `buffering` to `committed`, and that
transition is what the interrupt policy actually consults — configured intent
is downgraded to what the run's state allows.
[`interrupt-model.md`](/runtime/interrupt-model) specifies the state machine and the downgrade
matrix.

Past the commit boundary a root Run has two legal outcomes:

* finish
* stop with a partial result the user is told about

It never silently re-runs an already observed effect, whether that effect came
from the root loop or one graph node.

This classification is also the first thing to build. It is a pure contract
change with no behavior change, and every interrupt behavior is unsafe without
it.

### Restart is resumption with more information

The obvious objection to interrupting is that the agent keeps starting over and
never converges. That is only true if restart means starting from zero.

It does not, because Runs, RunSteps, WorkGraphs, node attempts, and artifacts are
persisted. A superseding root Run can inherit completed speculative work that is
still semantically valid, or propose a graph revision around it. The second
attempt does not begin from nothing.

A cancelled root Run persists with status `cancelled` and a `superseded_by`
pointer to the Run that replaced it. That pointer is also the audit trail: it is
the answer to "why did this run stop?", which is the question
[`philosophy.md`](/philosophy) commits to being able to answer for every agent turn.

### Starvation bound

Interrupt-always means a continuously-typing human never gets an answer.

The policy needs a hard bound:

* a maximum number of supersessions per logical request
* and/or a wall-clock ceiling, after which the root Run stops accepting
  interrupts and finishes

Without a bound the feature has a trivial denial-of-service-by-chattiness
failure mode. Graph replanning needs its own revision/replan ceiling for the same
reason: a planner that redraws the graph forever has merely reinvented an
unbounded loop at a higher level.

### The fast model decides refine versus new

Blanket interrupt-on-any-message is the wrong default. Not every follow-up
belongs to the in-flight request.

This is one more typed question for the `model_fast` classifier described
above:

```go theme={null}
type Relation struct {
    Relation   string  // refines | corrects | independent | cancels
    Confidence float64
}
```

"build me a url shortener" followed by "actually, in rust" refines — interrupt
and merge. A third message of "also, what's the weather" is independent — do not
interrupt it into the first run.

In graph mode, a refinement may produce a graph revision rather than a wholesale
restart. Completed valid branches stay completed; pending topology can change.

`cancels` is worth calling out on its own. "never mind, stop" is a first-class
relation, not an edge case, and honoring it is the cheapest win in the whole
feature.

### Whose interrupt counts

This is the Mesh-specific question, and no single-user harness has to ask it.

In a shared thread the next message may come from a different actor than the one
who started the run. Does one person's message interrupt a run another person
triggered?

That needs an explicit per-agent policy, not a default that quietly picks a
side:

* `same_actor_only` — only the actor who triggered the run may supersede it
* `thread_scoped` — anyone in the conversation may refine it
* `addressed_actor_only` — only the actor the agent is currently answering

Attribution has to survive the merge. A root Run answering a composite of two
actors' messages records both as triggers, and the reply may need to address
both. The WorkGraph inherits that attribution through the root Run; worker nodes
do not invent a new conversational owner.

This is the same family of problem as `awaiting_input`: cases where the
conversation, not the process, owns the turn.

### Ack interaction

If the harness already said *on it, this will take a bit*, a supersession must
not say it again.

Ack state is inherited by the superseding root Run. Progress and ack emissions
are keyed on the logical request rather than child execution ids, so graph
revisions and node retries cannot re-ack work the user was already told about.

Slack message edits are the same problem wearing a different hat. An edited
message is a refinement of an in-flight request, and it should route through
this policy rather than a parallel one.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.