Skip to main content

Interrupt model

Implementation note: Live steering (repository) describes the shipped requester-scoped amendment and cancellation behavior. The broader modes below remain design context rather than a claim that every mode is implemented. Message handling is not a serial FIFO. That sentence is the whole document. Everything below is the consequence of taking it seriously. A serial queue treats every inbound message as an independent unit of work and processes them in arrival order, one at a time, to completion. It is the obvious design and it is wrong for conversation, because conversation is not a stream of independent requests. It is a stream of revisions. The second message is usually not a new question; it is the first question, said better. So a turn does not consume a message. A turn runs against a mailbox it can observe mid-flight.

The mailbox

Each agent has a per-conversation mailbox: the ordered set of inbound messages that are addressed to it and not yet accounted for by a completed turn. “Addressed to it” is doing real work in that sentence, and it is decided elsewhere: addressing.md owns the gate that separates messages the agent merely observes in a shared channel from messages that are actually a request. Without that gate the mailbox in a busy channel is every message anybody ever sent, which is not a mailbox. That document also names the invariant this one assumed silently — at most one in-flight turn per (agent, conversation) — and without it the modes below are decorative, because a second message opens a second turn instead of reaching this mailbox at all. The properties that matter:
  • A turn reads the mailbox, it does not drain it. Messages are marked accounted-for when a turn commits, not when a turn starts.
  • The mailbox is observable during a run. A running turn can ask “has anything else arrived?” and get a truthful answer. This is what makes any of the modes below possible; without it the only implementable policy is FIFO.
  • The mailbox is durable. It is derived from persisted MessageEvent rows (see data-model.md), not from in-memory channel state, so a worker crash does not lose the fact that three more messages arrived.
A turn’s trigger set is therefore plural. One turn may be triggered by four messages from two actors, and all four are recorded as triggers. Attribution does not get collapsed just because the work did — see “Attribution survives coalescing” below.

Two config knobs

Interrupt behavior is configured per agent, because the right answer genuinely differs by agent. A support agent answering customers should probably wait and fold. A pager-duty responder should probably never wait at all.

1. Coalescing window

A duration. Messages arriving within the window of each other fold into a single turn.
The window is a quiescence timer, not a fixed delay: each new message inside the window restarts it, so a person typing four messages in a row produces one turn that starts shortly after they stop, not four turns. Rough intent of the common settings: Coalescing is the cheapest large win in this entire document, and it is worth being explicit about why: there is nothing to interrupt. No inference is wasted, no side effect is half-applied, no cancellation semantics have to be correct. Four messages become one turn with all four in context, and the agent answers the question the person actually finished asking. Ship the window first. Everything after this is harder and buys less.

2. Mid-turn mode

The window only covers messages that arrive before the turn starts. Mid-turn mode covers messages that arrive after.
  • preempt — cancel the in-flight turn and restart with the fuller context. The new turn inherits the cancelled turn’s completed steps, so restart is resumption with more information rather than starting from zero. Best when later messages usually invalidate earlier ones.
  • enqueue — let the in-flight turn finish, then process the new messages as the next turn. The conservative mode. Always legal, never surprising, occasionally answers a question the user has moved past.
  • amend — inject the new messages into the running turn as a signal the agent observes at its next decision point. Nothing is cancelled; the run simply learns something new before its next step. The most powerful mode and the most dependent on the agent loop actually checking for signals between steps.

The commit boundary

Here is the part that makes the config honest. Config expresses intent. The commit boundary determines what is actually possible. A turn has a moment before which it has changed nothing outside Mesh, and after which it has. That moment is the commit boundary, and it is crossed by the turn’s first external side effect:
  • a message sent to Slack
  • a Linear issue created or updated
  • a commit pushed, a PR opened
  • any write to any system Mesh does not own
Before that point, the turn is speculative. Cancelling it costs inference and nothing else. After that point, the turn has told the world something. You cannot un-send a Slack message, and a “restart” that re-runs the send does not retry it — it duplicates it. So the turn is a two-state machine:
buffering is not “before any work.” A turn in buffering may have made a dozen model calls and read from six systems. Pure reads and inference do not cross the boundary. Only mutations do. committed is terminal for the purposes of this policy. A turn does not return to buffering.

The downgrade matrix

Requested mode meets actual state, and the actual state wins: A downgrade is a normal event, not an error. It is recorded on the turn so the audit trail explains why a preempt-configured agent finished a turn it was asked to abandon: the turn had already posted. “Enqueue-or-reconcile” is deliberately two options, because after a side effect there are exactly two honest things to do:
  • Enqueue. Finish the turn, then handle the new messages in the next turn.
  • Reconcile. Finish the turn, and have the next turn correct the record — an edited message, a follow-up that says “ignore that, you’d said X,” a closed Linear issue. Reconciliation is a compensating action, which is the only general way to undo something that has already been observed.
Reconcile is strictly better UX and strictly more work. Enqueue is the required baseline; reconcile is the upgrade, and it is per-effect: a Slack send is reconcilable via edit, a pushed commit generally is not.

Design the boundary first

The ordering claim is the most important engineering statement here. The commit boundary is not an implementation detail of the three modes. It is the substrate they hang off. If it is added afterward, every mode has already been written against an assumption that turns are freely cancellable, and every one of them is subtly wrong in production in a way that manifests as duplicate Slack messages and double-created issues — the failure class that erodes trust in an agent fastest. Concretely, this means the first implementable slice of this document is not a mode. It is:
  1. Classify every RunStep as pure or effectful (see agent-harness.md).
  2. Record on the run the transition to committed, with the step that caused it.
  3. Expose that state to the interrupt policy.
  4. Then implement modes, all of which consult it.
With steps 1–3 and no modes at all, the system behaves exactly as it does today (effectively enqueue), and it is now correct-by-construction to add the rest. That is the change worth making first.

The enumerated effect surface (built, B1)

Steps 1–3 above now exist in code as internal/run/effect.go. It is the closed catalog of every side effect a turn can produce today, each classified by what cancelling or repeating it costs: The distinction that matters: internal_durable effects write to Postgres, which Mesh owns, so no external observer ever sees them and the definition of the boundary — “a write to any system Mesh does not own” — does not apply. They are durable but retractable by compensation (a forget undoes a remember). The memory verbs are declared effect/effectful in the tool contract because a remember cannot be repeated safely, and they stay inside the boundary through an explicit tool.Definition.Internal flag rather than a reclassification: the default for an effectful non-query tool is to cross, and only a hand-registered Mesh-owned verb may opt out. Tool effects and the reply are two different crossings, and the state machine knows the difference. The run-level fact “work is no longer freely cancellable because some external effect happened” and the slot-level fact “this reply must not be sent twice” are separate concerns, and conflating them would make a correctly crossed tool call refuse the turn’s own reply with ErrAlreadyCommitted. So:
  • the run transitions to committed on its first external-observed effect — a tool call or the reply — and remains committed; committed_at is stamped once and never moves;
  • a tool effect (run.Effect.Tool) commits the run but does not occupy the reply slot: further tool effects are legal, and the reply is legal and takes over committed_by_effect, so a resume that reads the cause learns “the reply was attempted” rather than “a tool ran and the reply is pending”;
  • after the reply, nothing crosses — not a second reply, and not a tool effect, which is a resume executing past the send;
  • the per-operation attempt record is the tool_calls row: started before the handler runs, completed or failed after, so “which calls ran” is answerable from the ledger without a provider round-trip.
The crossing is written before the handler runs (turn.Engine.executeRecordedTool), the same write-ahead order the reply uses, and it is fail-closed: a crossing that cannot be recorded refuses the effect. It is deliberately coarse — tool.workspace, tool.browser, tool.effect, by class — because the harness cannot see whether a workspace_exec pushed a branch or only ran the tests, and an honest coarse crossing beats a precise one it cannot verify. RecordRunCommitBoundary enforces the same rules in SQL, handed the catalog’s tool-effect names so the row and the in-memory machine cannot disagree.

What a resume does with the record

The run worker (internal/app/run_worker.go) re-claims a run whose lease lapsed and drives it back through Runtime.ResumeRun. Before the resumed turn’s first model call, three things happen, in this order:
  1. The reply boundary is consulted. A run whose reply already crossed is finalized loud (ErrAlreadyCommitted) and nothing is re-sent. A run committed only by a tool still owes its reply and falls through.
  2. The interrupted attempt is read back (run.Repository.RecordedAttempt): every primary-model completion, each tool request it made with its validated arguments, and the outcome the ledger holds for it — succeeded with its observation, failed or rejected with the reason, or started and never completed, which is the call the crash landed inside. A write the ledger recorded as failed whose journal intent was never confirmed is read the same way — the remote system may have accepted it before the error came back — so it is rendered as unknown, never as “did not take effect”. The turn engine (turn.WithRecovery) renders that record as attributed history under the normal context budget, exactly where a paused continuation’s transcript goes, and adds one developer instruction: continue from where the attempt stopped, do not repeat a recorded success, and treat an unknown outcome as inspect first. A recorded MCP or web result re-enters the model’s context with its tool_source_evidence id carried into the turn’s provenance, exactly as the live path does, so a reply or memory write built on it keeps the audience and conversation restriction recorded at the time. A read that fails is returned without terminating the run, so the lease is released and the run re-claimed rather than resumed blind. The run is also returned from awaiting_tool to running (run.Repository.ReopenAfterInterrupt): a crash mid-tool leaves it awaiting a round that is now over, and a tool-call reservation is admitted only for a running run.
  3. The workspace’s fate is stated, and verified. A run holds one workspace. A ready row is recorded state, not proof: before the model is told its files were retained, the resume asks the backend once (workspace.Manager.VerifyForRun — the same ownership and kind checks reuse applies, then one bounded probe). A row that answers is reused with its files intact; anything else (released, failed, stuck provisioning, or ready but unreachable) cannot be replaced within the run, so the model is told not to reach for workspace tools and to report what remains instead.
The effect journal (commitment_effects, migration 0045, applied by commitment.Service.Wrap) is what turns “may have happened” into a refusal: every durable run journals its declared external writes keyed on the run and a hash of the call, so an identical repeat returns the recorded result and a call whose previous attempt has no recorded outcome is refused with ErrAmbiguousEffect. The engine feeds that refusal back to the model as an observation (“this was NOT run now; go and look”) rather than failing the turn, so the resumed model can do the inspection the refusal asks for. Workspace and browser commands are deliberately not journaled by exact invocation — a repeated go test is not a duplicate effect — and the recovery instruction says so. What this does not yet provide: a provider idempotency key per external operation (so a retry of one create_github_pull_request could be proven a no-op rather than refused), and restoring a released workspace’s files from a work-checkpoint snapshot (snapshot-lineage.md) — today a resume reuses a live sandbox or tells the model the tree is gone.

Interaction with the rest of the harness

Attribution survives coalescing

A turn triggered by four messages from two actors records all four as triggers and both as contributing actors. philosophy.md says every agent turn is explainable; a coalesced turn that remembers only its first trigger is not. The reply may need to address both actors, and the reply’s own attribution edge points at the trigger set, not a single message.

Acks are keyed on the logical request

If the harness already said on it, this will take a bit, a preempting turn must not say it again. Emissions in agent-harness.md are idempotent on run id. Interrupts make run id the wrong key — a superseded run and its successor are the same logical request to the human who asked. Ack state is inherited across supersession and keyed on the logical request.

Whose interrupt counts

In a shared thread the next message may come from a different actor than the one who triggered the run. This needs an explicit per-agent policy rather than a default that quietly picks a side:
  • same_actor_only — only the triggering actor may preempt or amend
  • thread_scoped — anyone in the conversation may
  • addressed_actor_only — only the actor the agent is currently answering
This is the Mesh-specific question. A single-user harness never has to ask it, and it is downstream of the same property that makes awaiting_input work: sometimes the conversation, not the process, owns the turn.

Relation classification decides whether to interrupt at all

Not every follow-up belongs to the in-flight request. “also, what’s the weather” arriving during a code-generation run is a new turn, not a refinement of the old one. That judgment is a typed call to model_fast (see agent-harness.md), and it gates mode selection: independent means enqueue regardless of configuration, and cancels — “never mind, stop” — is a first-class relation worth honoring on its own.

Starvation bound

preempt with no bound means a continuously-typing human never gets an answer. The policy needs a hard ceiling — a maximum number of supersessions per logical request, and/or a wall-clock deadline after which the turn stops accepting interrupts and finishes. Without one, the feature has a denial-of-service-by-chattiness failure mode.

Message edits route through here

An edited Slack message is a refinement of an in-flight request. It should be handled by this policy rather than a parallel one that reinvents the same decisions worse.

What this document does not yet decide

  • Whether the coalescing window should become adaptive (a message ending mid-thought suggests more is coming). The fixed window handles the majority of real cases and ships first.
  • Which effects are reconcilable per connector. That belongs with each connector’s doc, not here.
  • Whether amend needs a distinct signal channel in the tool loop or can ride on the existing observation path.
  • Reconciliation across turns — driving a promise to done after the turn that made it has ended — is not this document. “Enqueue-or-reconcile” here is a per-effect compensating action inside the turn’s lifetime; a promise that outlives the turn is a Commitment, owned by commitments.md (repository). A Commitment’s reconcile Runs cross this document’s commit boundary like any other Run and never silently re-run an already-observed effect.

Status

Built, in the order “Design the boundary first” prescribes: effect classification and the buffering → committed machine (internal/run/effect.go), the durable crossing on the run row (0020, RecordCommit / LoadCommitBoundary), and the tool half — effectful workspace, browser, and effect-class calls cross at the executor before they run, and the reply is admitted after them. Live steering now adds durable requester follow-ups, obsolete-proposal invalidation, and explicit stop cancellation across model/tool calls. See live-steering.md (repository) for its precise boundaries. General replacement modes and the coalescing window remain future work.