Skip to main content

Agent harness

The turn pipeline proves the loop: one message in, one model call, one reply out. That is call-and-response. It is not an agent. The agent harness is the layer that turns a single model call into a bounded, resumable, multi-iteration run that can use tools, report progress, ask for help, and, when the work genuinely decomposes, compile the objective into a durable WorkGraph without losing the thread. The relationship is simple:
The harness makes one loop reliable. The WorkGraph composes reliable loops.
See work-graph.md for the graph contract.

Why it exists

A useful agent turn is not one request. It is a loop:
  • read the request
  • decide whether a tool is needed
  • call the tool
  • read the observation
  • decide again
  • eventually produce a reply
Two things break when that loop lives entirely inside one in-memory function call. First, the run becomes invisible. There is no persisted state, so nothing can report where the run is, nothing can resume it, and a worker crash silently loses the whole turn. The user sees nothing and never learns why. Second, the run becomes unattributable. Tool calls and observations are the most consequential things an agent does, and if they only exist in a stack frame they cannot be audited. philosophy.md says every agent turn is explainable. An unpersisted tool-calling loop is the fastest way to make that untrue. The harness exists to make long agent work durable, legible, and bounded.

A precondition, stated once

Everything below makes a turn more capable and therefore more expensive: tools, iterations, subprocesses, sandboxes, and eventually WorkGraph fan-out. That is fine as long as one thing is true upstream of it:
A turn exists because someone addressed the agent, and there is at most one in flight per (agent, conversation).
That is not this document’s contract — it is addressing.md’s — but the harness is where its absence gets expensive. A harness this powerful, driven by a router that treats every message in a busy channel as a request, is a fan-out of tool-calling loops that ends with the host out of CPU and every co-tenanted agent dead. Budgets below bound one run. They do not bound the number of runs. WorkGraphs do not weaken this invariant. One conversation turn still owns one root Run. Graph nodes are child executions inside that Run, not additional conversation turns.

Runs and run steps

A turn becomes a Run. Each iteration of a loop becomes a RunStep. Run statuses:
  • queued — accepted, not yet picked up
  • running — a worker holds it
  • awaiting_tool — a tool call is outstanding
  • awaiting_input — the agent asked a question and is waiting on an actor
  • succeeded — terminal, produced a final reply
  • failed — terminal, produced an error
  • cancelled — terminal, stopped by an actor
  • budget_exhausted — terminal, hit a declared limit
Each iteration commits as one transaction: the step, its tool calls, its observations, and the new run status land together. A worker that dies mid-run resumes from the last committed step instead of starting over or disappearing. This is the root fix, and it is worth being blunt about why. The failure mode we are designing against is an agent that accepts a hard request and then goes quiet for four minutes. That is not a messaging bug. It is a state bug: a long turn held as one in-memory goroutine has no state to report progress from and no state to resume to. Persist the run and progress reporting becomes a read, not a guess. This fix has a hard scope, and its boundary is the subject of another contract. A durable Run makes one turn survive a crash. It does nothing for a promise that outlives the turn — an agent that says “on it, spinning up the implementation” and then reaches succeeded having created nothing has not crashed; the turn succeeded and the promise evaporated. That is a Commitment, not a Run, and it is owned by commitments.md (repository). Do not expect the durable Run to close it. The same rows are the audit substrate. What one run step must capture — the model call with its request, response, and token usage; the tool call with its arguments and observation — and how an operator reads it back is specified in observability.md. Capture ships with the ledger, not as a follow-up.

Execution shape: loop or graph

A root Run has an execution shape:
This is not a claim that every request deserves decomposition. A short, coherent, sequential request should stay a loop. Conceptually that is a one-node graph; the implementation does not need to materialize graph state for it. Graph mode earns its complexity when work splits into independent branches, needs explicit verification or fan-in, requires differently scoped execution contexts, or needs a human gate before an effect. A conservative classifier may use model_fast to decide whether graph planning is worth attempting. The classifier decides shape, not topology. If it is uncertain or fails, the safe fallback is loop. When shape is graph, model_primary (or a future dedicated planning role) may propose topology. The proposal is typed data and is handed to the deterministic admission layer described in work-graph.md. The model never turns its own proposal directly into scheduled work. A root Run may also begin as a loop and discover, before its external-effect commit boundary, that the objective genuinely decomposes. It may then propose a WorkGraph. That transition is persisted and explainable.

Child runs are workers, not new agents

A model-backed WorkGraph node executes as a bounded child Run (or an equivalent child execution unit carrying the same semantics). The child Run exists to reuse the machinery this document already owns:
  • durable RunSteps
  • budgets
  • cancellation
  • tool-call persistence
  • progress
  • workspace acquisition
  • failure state
It does not create a new RuntimeAgent. A child Run:
  • belongs to the same RuntimeAgent perspective as its root Run
  • has no connector identity of its own
  • does not appear as a coworker in Slack
  • does not become an Actor
  • does not own private long-term memory
  • receives a bounded subset of the root Run’s context and tools
  • may never have greater authority than its parent
This is how Mesh gets fresh specialist execution contexts without violating the perspective boundary in ../perspective.md. V1 child Runs execute bounded loops; they do not recursively create private WorkGraphs. A worker that discovers more decomposition requests a root-graph revision. That keeps all fan-out visible to one scheduler and one budget.

Acking that the run is working

Ingress enqueues the root Run and returns. It does not wait for the model. The ack is a reaction on the triggering message, not a message of its own: ⏳ the instant the run is admitted, swapped for ✅ or ❌ when it settles. It goes on immediately, for every run, because a reaction is unobtrusive enough that there is no cost to being early and no need to predict whether the run will be slow. Mesh previously acked with a threaded on it, this will take a bit reply, posted by a timer at a five-second deadline. That reply is retired: the reaction says the same thing sooner and without spending a message in the thread. Its deadline reasoning is what survives and is worth keeping for any FUTURE message-shaped progress signal — a deadline cannot be wrong the way a prediction can, because it observes that a reply has not landed rather than guessing that it won’t. A threaded message earns its place only for work long enough that a human would otherwise wonder whether the agent died, and that is the only thing it should ever be used for. Graph planning and graph execution are covered by the same root ack. Child Runs do not emit their own conversational acks; they are implementation details of one user-visible objective.

Complexity classification

The fast-model classifier is an optimization on top of the deadline, not a replacement for it. It has no bearing on the reaction, which is unconditional; what it triages is the message-shaped progress signal described above and the execution shape. It runs in parallel with the root Run, and its jobs may include both expected duration and execution-shape triage:
  • obviously long work — set the deadline to zero and say something immediately
  • obviously trivial work — suppress the message entirely
  • plausibly decomposable work — allow graph planning
  • anything else — leave the default deadline and loop execution alone
Typed output, so a bad classifier response degrades to the default behavior instead of corrupting the run:
graph_candidate is deliberately not graph. The classifier may nominate a request for planning, but the planner still has to produce a useful DAG and the runtime still has to admit it. AckText is the sharp part. The classifier writes the ack copy in the same call it does triage, and the harness stashes it on the root Run. When the deadline fires, delivering a specific, context-aware ack costs zero additional model latency — the sentence is already sitting there. If the classifier is slow, errors, or is not configured, the deadline mechanism still works. That ordering is deliberate: the mechanism must be correct on its own, and the model only makes it feel better.

Model roles

Primary and fast remain the main per-agent text models:
  • model_primary — reasoning, tool use, final replies, and initial WorkGraph planning
  • model_fast — quick drafting, extraction, summaries and compatibility classification
  • the independent installation decision model — fixed-choice classification through a native decision API; see decision-models.md (repository)
This is a first-class part of the agent contract, not a classifier implementation detail. The fast model is reused for:
  • complexity triage and ack copy
  • execution-shape candidacy
  • conversation titles
  • context compaction
  • routing decisions
Use the decision role for existing engagement, continuation consent and memory sensitivity classification. Keep generation and open-ended extraction on the fast model; evaluate new bounded decision use cases independently. Do not add model_planner merely because WorkGraph exists. Measure first. If planning develops materially different capability/cost requirements, it can become a model role later without changing the graph contract.

Tool contract

Tools are registered per agent and gated per connector, because the same agent in a private channel and in a shared one should not necessarily have the same reach. tools.md is the normative contract for tool effect classes, registration, gating, and the cloud services each class executes on. Exec-class tools — shell, file writes, git, tests, builds — do not run in the harness process. They run in a sandboxed workspace materialized by a pluggable execution layer, and the loop above stays in the control plane. sandboxed-execution.md is the normative contract for that boundary, including the credential model, which is a grant per run rather than the agent’s full keyring. No tool names an execution backend. For graph mode, an agent_loop node that needs exec-class tools runs as a child Run and therefore gets its own workspace. Parallel WorkGraph nodes never share one mutable working directory merely because they share a root Run. A later join node owns reconciliation. Every tool call and every observation is a persisted event. That keeps runs explainable after the fact and keeps attribution intact: a tool ran because a specific actor said a specific thing in a specific thread, through a specific root Run and, when graph mode is active, a specific WorkNodeAttempt.

Budgets

Runs must terminate legibly rather than hang: Exhausting a budget ends the current allowance, not the user’s task. For human-requested work, the final completion explains actual progress and what remains, then asks whether to continue. An affirmative reply from the original requester authorizes one fresh bounded run with the saved task context. Silence, a refusal, or an ambiguous reply never authorizes more task work. Runtime counters and internal state labels belong in the ledger, not the conversation. See human-authorized continuation (repository) for the implemented consent, checkpoint, and workspace lifecycle. Graph mode adds composition rather than a new source of money. The root Run owns an aggregate budget; WorkNodes receive bounded slices of it. Drawing six nodes cannot turn one Run’s budget into six Run budgets. WorkGraph-specific ceilings include node count, depth, parallelism, revisions, replans, and aggregate model/tool/compute spend. See work-graph.md. The budget that is not per-root-run — how many root Runs may exist at once, for one agent and for one install — is admission control, and it lives with the addressing/concurrency contract in addressing.md. Graph fan-out is additionally constrained by those same per-agent and per-install ceilings. Ten child Runs each perfectly within their node budget can still saturate a host.

Tool results are compacted within a run

The loop re-sends the whole conversation on every completion, so a run that makes real tool calls — a checkout, a dozen test runs, several file reads — would otherwise grow its request until the model’s context window failed it, well before any budget above fired. internal/turn/compaction.go bounds that: before each completion, the newest few tool results stay verbatim (they are what the model is reasoning about), older results are kept verbatim up to a byte budget, and everything past it is replaced in place by a short excerpt — the result’s first and last lines, its size, and a pointer to the run ledger, where the full text lives in tool_observations. This is deliberately mechanical, not a model-written summary. It touches only the current run’s own tool output, never conversation history, memory, or attribution, so it is not the compaction Track A3 gates on an attribution eval; the elision is visible to the model the way the workspace tools’ output clamps are, and it is safe for the same reason. It also does not change the quadratic shape internal/run/budget.go describes — tool calls, assistant turns, history, and the excerpts themselves still accumulate — it shrinks the per-round delta enough that a coding run finishes on its budget instead of at the provider.

Progress emissions

Progress is emitted as events, like everything else outbound:
  • the ack
  • periodic progress while the run is long
  • the terminal result
All emissions are idempotent and keyed on the logical request/root Run, so a worker retry cannot spam a thread with duplicate acks. Graph mode gives progress a much better substrate. Instead of inventing progress copy from an opaque loop, Mesh can summarize durable WorkGraph facts: three of five branches complete, one verifier running, one branch blocked. Child Runs do not independently narrate themselves into the channel unless the root Run’s progress policy chooses to surface them. Slack-specific: prefer a reaction lifecycle for medium-length work — eyes when the run starts, brain while it is thinking, check when it lands. Post a real threaded message only for runs long enough that a human would otherwise wonder whether the agent died. Channel noise is a cost, and the reaction lane pays less of it. The ⏳/✅/❌ reaction covers every turn. Contextual progress starts automatically for every active agent on new and resumed executions. No per-agent configuration is needed; existing settings cannot disable it. internal/progress supervises the active turn independently of model/tool calls. It sends an initial update after 60 seconds of silence and subsequent updates at three-minute intervals. Agent report_progress checkpoints can arrive sooner for a plan, milestone, replan, or wait, with a one-minute minimum gap after the previous attempt. The supervisor coalesces candidates and uses a bounded fast-model summary when the agent has not supplied an update. A slow or unavailable summarizer falls back to honest activity/waiting facts. Workspace acquisition no longer sends a fixed sandbox message. Its first successful acquisition feeds the same progress supervisor, which uses the agent’s checkpoint or the task objective and current activity to compose copy. Repeated acquisitions do not create additional acknowledgements. Progress is an additive run.EffectProgressNotice, with a durable reservation per logical subject/sequence. Ordinary runs use their run ID; commitments share cadence across continuation attempts. Progress never settles the task, consumes the final reply boundary, or counts as completion evidence. Ambiguous delivery is recorded and never blindly replayed. This is at-most-once application dispatch, not a guarantee of exactly-once remote receipt. The runtime stops and joins the supervisor before sending a result or failure reply. Cancellation and ownership checks prevent stale attempts from sending; a remote request already admitted before cancellation cannot be recalled. Progress copy and receipts are retained in progress_events and the dashboard trace, separately from ordinary assistant history and automatic memory extraction. Follow-up context may include three recent progress records, explicitly labeled and subject to live source/disclosure checks; their lineage also propagates into any resulting reply. See the implementation plan (repository) for the design, acceptance criteria, and implementation boundaries.

awaiting_input is the differentiator

Mid-run, an agent can ask a clarifying question and park the run in awaiting_input. In a single-user harness that is a modal prompt: the one person who asked has to be the one who answers. In Mesh it is a thread. Any actor in that conversation can answer and resume the run, and because every message already carries an actor, the resumed run records who unblocked it. The clarification is part of the conversation graph, not a side channel. Graph mode can also contain a human_gate WorkNode. The distinction is semantic:
  • awaiting_input means the agent lacks information needed to continue
  • human_gate means the runtime requires an attributed decision before a declared edge, usually because the next work is effectful or expensive to undo
Both pause durably. Neither becomes a private modal prompt outside the thread. That is the capability the multi-actor model buys us, and no harness built around a single conversation partner can express it.

First slice

Keep the first implementation boring:
  • one root Run
  • execution shape loop
  • one step, one model call, no tools
  • ingress enqueues, a worker executes, the reply routes back as it does today
  • the ack timer wired but with no classifier
That proves the durable Run contract before WorkGraph fan-out exists. Then implement the graph substrate before broad exec-class tools:
  • loop | graph execution shape
  • WorkGraph proposal + deterministic admission
  • DAG scheduling
  • child Runs
  • typed artifacts
  • graph budgets and revisions
A graph made of unreliable loops is simply a more parallel failure mode. The ordering is deliberate.

Interrupt-driven turns

Everything above assumes a request stops arriving once the run starts. It does not.
The normative contract for this — the mailbox, the per-agent config surface, and the commit-boundary state machine that decides which behaviors are legal — lives in interrupt-model.md. This section is the harness-level narrative: why the problem exists and what it costs. Where the two appear to disagree, interrupt-model.md is authoritative.
Today rapid-fire messages are processed serially: thirty seconds on the first, then thirty seconds on the second. That is frustrating, and it is usually also wrong. Messages in quick succession are normally one request — a follow-up adding context, or a correction — not two independent ones. Answering the first in isolation spends a full run producing a reply the person has already moved past. So the harness needs a position on what happens to in-flight work when the conversation keeps going. Graph mode does not create a second interrupt model. New conversational input still targets the root Run. The root determines whether to amend context, preempt speculative work, enqueue the input, or revise the WorkGraph. Child Runs never independently consume messages from the conversation mailbox.

Debounce before interrupt

The cheapest large win is not cancellation. It is a quiescence window before a run starts. Hold a newly-enqueued turn for a short quiescence interval and let messages that arrive inside that window join the same run. Four rapid messages become one run with all four in context. interrupt-model.md specifies the window as per-agent config; a low-single-digit-seconds default is a reasonable starting point for a conversational agent. This is strictly better than interrupting, because there is nothing to interrupt:
  • no wasted inference
  • no restart cost
  • no cancellation semantics to get right
The window is configurable per agent (interrupt.coalesce_window) and can become adaptive later — a message that ends mid-thought suggests more is coming — but the fixed window already handles the majority of real cases. It ships first.

Cancellation safety is a property of steps and attempts

This is the correctness property. Everything else in this section is experience. Whether in-flight work can be abandoned does not depend on the root Run alone. It depends on what execution already did. Inference and pure reads are safe to cancel and redo. A step or WorkNodeAttempt that already had a side effect — posted a Slack message, created an issue, pushed a commit — cannot be silently restarted, because restarting it duplicates it. So RunStep carries a replayability classification, and graph mode extends the same classification to WorkNodeAttempt. The root Run’s first externally effectful step or node attempt moves it from buffering to committed, and that transition is what the interrupt policy actually consults — configured intent is downgraded to what the run’s state allows. interrupt-model.md specifies the state machine and the downgrade matrix. Past the commit boundary a root Run has two legal outcomes:
  • finish
  • stop with a partial result the user is told about
It never silently re-runs an already observed effect, whether that effect came from the root loop or one graph node. This classification is also the first thing to build. It is a pure contract change with no behavior change, and every interrupt behavior is unsafe without it.

Restart is resumption with more information

The obvious objection to interrupting is that the agent keeps starting over and never converges. That is only true if restart means starting from zero. It does not, because Runs, RunSteps, WorkGraphs, node attempts, and artifacts are persisted. A superseding root Run can inherit completed speculative work that is still semantically valid, or propose a graph revision around it. The second attempt does not begin from nothing. A cancelled root Run persists with status cancelled and a superseded_by pointer to the Run that replaced it. That pointer is also the audit trail: it is the answer to “why did this run stop?”, which is the question philosophy.md commits to being able to answer for every agent turn.

Starvation bound

Interrupt-always means a continuously-typing human never gets an answer. The policy needs a hard bound:
  • a maximum number of supersessions per logical request
  • and/or a wall-clock ceiling, after which the root Run stops accepting interrupts and finishes
Without a bound the feature has a trivial denial-of-service-by-chattiness failure mode. Graph replanning needs its own revision/replan ceiling for the same reason: a planner that redraws the graph forever has merely reinvented an unbounded loop at a higher level.

The fast model decides refine versus new

Blanket interrupt-on-any-message is the wrong default. Not every follow-up belongs to the in-flight request. This is one more typed question for the model_fast classifier described above:
“build me a url shortener” followed by “actually, in rust” refines — interrupt and merge. A third message of “also, what’s the weather” is independent — do not interrupt it into the first run. In graph mode, a refinement may produce a graph revision rather than a wholesale restart. Completed valid branches stay completed; pending topology can change. cancels is worth calling out on its own. “never mind, stop” is a first-class relation, not an edge case, and honoring it is the cheapest win in the whole feature.

Whose interrupt counts

This is the Mesh-specific question, and no single-user harness has to ask it. In a shared thread the next message may come from a different actor than the one who started the run. Does one person’s message interrupt a run another person triggered? That needs an explicit per-agent policy, not a default that quietly picks a side:
  • same_actor_only — only the actor who triggered the run may supersede it
  • thread_scoped — anyone in the conversation may refine it
  • addressed_actor_only — only the actor the agent is currently answering
Attribution has to survive the merge. A root Run answering a composite of two actors’ messages records both as triggers, and the reply may need to address both. The WorkGraph inherits that attribution through the root Run; worker nodes do not invent a new conversational owner. This is the same family of problem as awaiting_input: cases where the conversation, not the process, owns the turn.

Ack interaction

If the harness already said on it, this will take a bit, a supersession must not say it again. Ack state is inherited by the superseding root Run. Progress and ack emissions are keyed on the logical request rather than child execution ids, so graph revisions and node retries cannot re-ack work the user was already told about. Slack message edits are the same problem wearing a different hat. An edited message is a refinement of an in-flight request, and it should route through this policy rather than a parallel one.