Agent harness
The turn pipeline proves the loop: one message in, one model call, one reply out. That is call-and-response. It is not an agent. The agent harness is the layer that turns a single model call into a bounded, resumable, multi-iteration run that can use tools, report progress, ask for help, and, when the work genuinely decomposes, compile the objective into a durable WorkGraph without losing the thread. The relationship is simple:The harness makes one loop reliable. The WorkGraph composes reliable loops.See
work-graph.md for the graph contract.
Why it exists
A useful agent turn is not one request. It is a loop:- read the request
- decide whether a tool is needed
- call the tool
- read the observation
- decide again
- eventually produce a reply
philosophy.md says every agent turn is explainable.
An unpersisted tool-calling loop is the fastest way to make that untrue.
The harness exists to make long agent work durable, legible, and bounded.
A precondition, stated once
Everything below makes a turn more capable and therefore more expensive: tools, iterations, subprocesses, sandboxes, and eventually WorkGraph fan-out. That is fine as long as one thing is true upstream of it:A turn exists because someone addressed the agent, and there is at most one in flight per (agent, conversation).That is not this document’s contract — it is
addressing.md’s —
but the harness is where its absence gets expensive. A harness this powerful, driven by a
router that treats every message in a busy channel as a request, is a fan-out of
tool-calling loops that ends with the host out of CPU and every co-tenanted agent
dead. Budgets below bound one run. They do not bound the number of runs.
WorkGraphs do not weaken this invariant. One conversation turn still owns one
root Run. Graph nodes are child executions inside that Run, not additional
conversation turns.
Runs and run steps
A turn becomes aRun. Each iteration of a loop becomes a RunStep.
Run statuses:
queued— accepted, not yet picked uprunning— a worker holds itawaiting_tool— a tool call is outstandingawaiting_input— the agent asked a question and is waiting on an actorsucceeded— terminal, produced a final replyfailed— terminal, produced an errorcancelled— terminal, stopped by an actorbudget_exhausted— terminal, hit a declared limit
succeeded having created nothing has not crashed; the turn
succeeded and the promise evaporated. That is a Commitment, not a Run, and it is
owned by commitments.md (repository). Do not expect the durable Run to
close it.
The same rows are the audit substrate. What one run step must capture — the
model call with its request, response, and token usage; the tool call with its
arguments and observation — and how an operator reads it back is specified in
observability.md. Capture ships with the ledger, not as a
follow-up.
Execution shape: loop or graph
A root Run has an execution shape:model_fast to decide whether graph planning
is worth attempting. The classifier decides shape, not topology. If it is
uncertain or fails, the safe fallback is loop.
When shape is graph, model_primary (or a future dedicated planning role) may
propose topology. The proposal is typed data and is handed to the deterministic
admission layer described in work-graph.md. The model never
turns its own proposal directly into scheduled work.
A root Run may also begin as a loop and discover, before its external-effect
commit boundary, that the objective genuinely decomposes. It may then propose a
WorkGraph. That transition is persisted and explainable.
Child runs are workers, not new agents
A model-backed WorkGraph node executes as a bounded child Run (or an equivalent child execution unit carrying the same semantics). The child Run exists to reuse the machinery this document already owns:- durable RunSteps
- budgets
- cancellation
- tool-call persistence
- progress
- workspace acquisition
- failure state
RuntimeAgent.
A child Run:
- belongs to the same RuntimeAgent perspective as its root Run
- has no connector identity of its own
- does not appear as a coworker in Slack
- does not become an Actor
- does not own private long-term memory
- receives a bounded subset of the root Run’s context and tools
- may never have greater authority than its parent
../perspective.md.
V1 child Runs execute bounded loops; they do not recursively create private
WorkGraphs. A worker that discovers more decomposition requests a root-graph
revision. That keeps all fan-out visible to one scheduler and one budget.
Acking that the run is working
Ingress enqueues the root Run and returns. It does not wait for the model. The ack is a reaction on the triggering message, not a message of its own: ⏳ the instant the run is admitted, swapped for ✅ or ❌ when it settles. It goes on immediately, for every run, because a reaction is unobtrusive enough that there is no cost to being early and no need to predict whether the run will be slow. Mesh previously acked with a threaded on it, this will take a bit reply, posted by a timer at a five-second deadline. That reply is retired: the reaction says the same thing sooner and without spending a message in the thread. Its deadline reasoning is what survives and is worth keeping for any FUTURE message-shaped progress signal — a deadline cannot be wrong the way a prediction can, because it observes that a reply has not landed rather than guessing that it won’t. A threaded message earns its place only for work long enough that a human would otherwise wonder whether the agent died, and that is the only thing it should ever be used for. Graph planning and graph execution are covered by the same root ack. Child Runs do not emit their own conversational acks; they are implementation details of one user-visible objective.Complexity classification
The fast-model classifier is an optimization on top of the deadline, not a replacement for it. It has no bearing on the reaction, which is unconditional; what it triages is the message-shaped progress signal described above and the execution shape. It runs in parallel with the root Run, and its jobs may include both expected duration and execution-shape triage:- obviously long work — set the deadline to zero and say something immediately
- obviously trivial work — suppress the message entirely
- plausibly decomposable work — allow graph planning
- anything else — leave the default deadline and loop execution alone
graph_candidate is deliberately not graph. The classifier may nominate a
request for planning, but the planner still has to produce a useful DAG and the
runtime still has to admit it.
AckText is the sharp part. The classifier writes the ack copy in the same call
it does triage, and the harness stashes it on the root Run. When the deadline
fires, delivering a specific, context-aware ack costs zero additional model
latency — the sentence is already sitting there.
If the classifier is slow, errors, or is not configured, the deadline mechanism
still works. That ordering is deliberate: the mechanism must be correct on its
own, and the model only makes it feel better.
Model roles
Primary and fast remain the main per-agent text models:model_primary— reasoning, tool use, final replies, and initial WorkGraph planningmodel_fast— quick drafting, extraction, summaries and compatibility classification- the independent installation decision model — fixed-choice classification through a native decision API; see decision-models.md (repository)
- complexity triage and ack copy
- execution-shape candidacy
- conversation titles
- context compaction
- routing decisions
model_planner merely because WorkGraph exists. Measure first. If
planning develops materially different capability/cost requirements, it can
become a model role later without changing the graph contract.
Tool contract
tools.md is the normative contract for tool effect
classes, registration, gating, and the cloud services each class executes on.
Exec-class tools — shell, file writes, git, tests, builds — do not run in the
harness process. They run in a sandboxed workspace materialized by a pluggable
execution layer, and the loop above stays in the control plane.
sandboxed-execution.md is the normative contract for that boundary,
including the credential model, which is a grant per run rather than the agent’s
full keyring. No tool names an execution backend.
For graph mode, an agent_loop node that needs exec-class tools runs as a child
Run and therefore gets its own workspace. Parallel WorkGraph nodes never share
one mutable working directory merely because they share a root Run. A later
join node owns reconciliation.
Every tool call and every observation is a persisted event. That keeps runs
explainable after the fact and keeps attribution intact: a tool ran because a
specific actor said a specific thing in a specific thread, through a specific
root Run and, when graph mode is active, a specific WorkNodeAttempt.
Budgets
Runs must terminate legibly rather than hang:- max iterations
- max tool calls
- max tokens
- max wall clock
- max compute-seconds (see
sandboxed-execution.md)
work-graph.md.
The budget that is not per-root-run — how many root Runs may exist at once, for
one agent and for one install — is admission control, and it lives with the
addressing/concurrency contract in addressing.md. Graph
fan-out is additionally constrained by those same per-agent and per-install
ceilings. Ten child Runs each perfectly within their node budget can still
saturate a host.
Tool results are compacted within a run
The loop re-sends the whole conversation on every completion, so a run that makes real tool calls — a checkout, a dozen test runs, several file reads — would otherwise grow its request until the model’s context window failed it, well before any budget above fired.internal/turn/compaction.go bounds that: before
each completion, the newest few tool results stay verbatim (they are what the
model is reasoning about), older results are kept verbatim up to a byte budget,
and everything past it is replaced in place by a short excerpt — the result’s
first and last lines, its size, and a pointer to the run ledger, where the full
text lives in tool_observations.
This is deliberately mechanical, not a model-written summary. It touches only
the current run’s own tool output, never conversation history, memory, or
attribution, so it is not the compaction Track A3 gates on an attribution eval;
the elision is visible to the model the way the workspace tools’ output clamps
are, and it is safe for the same reason. It also does not change the quadratic
shape internal/run/budget.go describes — tool calls, assistant turns, history,
and the excerpts themselves still accumulate — it shrinks the per-round delta
enough that a coding run finishes on its budget instead of at the provider.
Progress emissions
Progress is emitted as events, like everything else outbound:- the ack
- periodic progress while the run is long
- the terminal result
internal/progress supervises the active turn independently of model/tool
calls. It sends an initial update after 60 seconds of silence and subsequent
updates at three-minute intervals. Agent report_progress checkpoints can
arrive sooner for a plan, milestone, replan, or wait, with a one-minute minimum
gap after the previous attempt. The supervisor coalesces candidates and uses a
bounded fast-model summary when the agent has not supplied an update. A slow or
unavailable summarizer falls back to honest activity/waiting facts.
Workspace acquisition no longer sends a fixed sandbox message. Its first
successful acquisition feeds the same progress supervisor, which uses the
agent’s checkpoint or the task objective and current activity to compose copy.
Repeated acquisitions do not create additional acknowledgements.
Progress is an additive run.EffectProgressNotice, with a durable reservation
per logical subject/sequence. Ordinary runs use their run ID; commitments share
cadence across continuation attempts. Progress never settles the task, consumes
the final reply boundary, or counts as completion evidence. Ambiguous delivery
is recorded and never blindly replayed. This is at-most-once application
dispatch, not a guarantee of exactly-once remote receipt.
The runtime stops and joins the supervisor before sending a result or failure
reply. Cancellation and ownership checks prevent stale attempts from sending;
a remote request already admitted before cancellation cannot be recalled.
Progress copy and receipts are retained in progress_events and the dashboard
trace, separately from ordinary assistant history and automatic memory
extraction. Follow-up context may include three recent progress records,
explicitly labeled and subject to live source/disclosure checks; their lineage
also propagates into any resulting reply.
See the implementation plan (repository) for the design,
acceptance criteria, and implementation boundaries.
awaiting_input is the differentiator
Mid-run, an agent can ask a clarifying question and park the run in
awaiting_input.
In a single-user harness that is a modal prompt: the one person who asked has to
be the one who answers.
In Mesh it is a thread. Any actor in that conversation can answer and resume
the run, and because every message already carries an actor, the resumed run
records who unblocked it. The clarification is part of the conversation graph,
not a side channel.
Graph mode can also contain a human_gate WorkNode. The distinction is semantic:
awaiting_inputmeans the agent lacks information needed to continuehuman_gatemeans the runtime requires an attributed decision before a declared edge, usually because the next work is effectful or expensive to undo
First slice
Keep the first implementation boring:- one root Run
- execution shape
loop - one step, one model call, no tools
- ingress enqueues, a worker executes, the reply routes back as it does today
- the ack timer wired but with no classifier
loop | graphexecution shape- WorkGraph proposal + deterministic admission
- DAG scheduling
- child Runs
- typed artifacts
- graph budgets and revisions
Interrupt-driven turns
Everything above assumes a request stops arriving once the run starts. It does not.The normative contract for this — the mailbox, the per-agent config surface, and the commit-boundary state machine that decides which behaviors are legal — lives inToday rapid-fire messages are processed serially: thirty seconds on the first, then thirty seconds on the second. That is frustrating, and it is usually also wrong. Messages in quick succession are normally one request — a follow-up adding context, or a correction — not two independent ones. Answering the first in isolation spends a full run producing a reply the person has already moved past. So the harness needs a position on what happens to in-flight work when the conversation keeps going. Graph mode does not create a second interrupt model. New conversational input still targets the root Run. The root determines whether to amend context, preempt speculative work, enqueue the input, or revise the WorkGraph. Child Runs never independently consume messages from the conversation mailbox.interrupt-model.md. This section is the harness-level narrative: why the problem exists and what it costs. Where the two appear to disagree,interrupt-model.mdis authoritative.
Debounce before interrupt
The cheapest large win is not cancellation. It is a quiescence window before a run starts. Hold a newly-enqueued turn for a short quiescence interval and let messages that arrive inside that window join the same run. Four rapid messages become one run with all four in context.interrupt-model.md specifies the window as
per-agent config; a low-single-digit-seconds default is a reasonable starting
point for a conversational agent.
This is strictly better than interrupting, because there is nothing to
interrupt:
- no wasted inference
- no restart cost
- no cancellation semantics to get right
interrupt.coalesce_window) and can become
adaptive later — a message that ends mid-thought suggests more is coming — but
the fixed window already handles the majority of real cases. It ships first.
Cancellation safety is a property of steps and attempts
This is the correctness property. Everything else in this section is experience. Whether in-flight work can be abandoned does not depend on the root Run alone. It depends on what execution already did. Inference and pure reads are safe to cancel and redo. A step or WorkNodeAttempt that already had a side effect — posted a Slack message, created an issue, pushed a commit — cannot be silently restarted, because restarting it duplicates it. SoRunStep carries a replayability classification, and graph mode extends the
same classification to WorkNodeAttempt. The root Run’s first externally
effectful step or node attempt moves it from buffering to committed, and that
transition is what the interrupt policy actually consults — configured intent
is downgraded to what the run’s state allows.
interrupt-model.md specifies the state machine and the downgrade
matrix.
Past the commit boundary a root Run has two legal outcomes:
- finish
- stop with a partial result the user is told about
Restart is resumption with more information
The obvious objection to interrupting is that the agent keeps starting over and never converges. That is only true if restart means starting from zero. It does not, because Runs, RunSteps, WorkGraphs, node attempts, and artifacts are persisted. A superseding root Run can inherit completed speculative work that is still semantically valid, or propose a graph revision around it. The second attempt does not begin from nothing. A cancelled root Run persists with statuscancelled and a superseded_by
pointer to the Run that replaced it. That pointer is also the audit trail: it is
the answer to “why did this run stop?”, which is the question
philosophy.md commits to being able to answer for every agent turn.
Starvation bound
Interrupt-always means a continuously-typing human never gets an answer. The policy needs a hard bound:- a maximum number of supersessions per logical request
- and/or a wall-clock ceiling, after which the root Run stops accepting interrupts and finishes
The fast model decides refine versus new
Blanket interrupt-on-any-message is the wrong default. Not every follow-up belongs to the in-flight request. This is one more typed question for themodel_fast classifier described
above:
cancels is worth calling out on its own. “never mind, stop” is a first-class
relation, not an edge case, and honoring it is the cheapest win in the whole
feature.
Whose interrupt counts
This is the Mesh-specific question, and no single-user harness has to ask it. In a shared thread the next message may come from a different actor than the one who started the run. Does one person’s message interrupt a run another person triggered? That needs an explicit per-agent policy, not a default that quietly picks a side:same_actor_only— only the actor who triggered the run may supersede itthread_scoped— anyone in the conversation may refine itaddressed_actor_only— only the actor the agent is currently answering
awaiting_input: cases where the
conversation, not the process, owns the turn.