MESH_RUN_WORKER_ENABLED=true (the default), a replacement worker claims
the same run after its lease expires and resumes from that frontier. No external
workflow service is required. Recovery is bounded by the run’s existing attempt,
iteration, tool-call and token limits.
Durable boundaries
run_checkpoints has one versioned row per run, referring to the current
run_steps and model_calls audit rows. Its provider-neutral payload contains
the working transcript, normalized model decision, tool activations, source
lineage, committed memory operations, continuation identity and, when available,
the completed turn. The prompt and context reports are retained for accounting;
the saved prompt is never used as live instructions.
The steering mailbox revision accompanies a model decision. Recovery checks
for requester corrections before accepting its reply or executing its tools.
Every individual tool outcome is recorded before the round advances. If the
process dies between one tool’s completion and the next frontier, recovery
reconstructs that round from its existing call rows without reserving tool
budget again. Successful and failed outcomes are reused. An admitted call that
never started may execute after current validation and authority checks.
An interrupted inference can have incurred provider cost before its result was
recorded. Recovery cannot reconstruct those unknown tokens; the old attempt and
its iteration reservation remain visible, and only recorded usage is billed in
the ledger. A restart does not restore consumed budget. The existing worker
grants a resumed attempt its admitted wall-clock allowance again.
When no inference can finish, a deterministic fallback can close the interrupted
model/step rows and save its status atomically without inventing a model result.
Interrupted tools
A started call with no durable outcome is retried only when both its original and current declarations sayReplayable, their fingerprints match, and there
is no unconfirmed effect intent. The fingerprint covers schema, execution class,
replay policy, required capabilities and internal-effect classification.
Declarations must accurately describe handler behavior.
All other started calls become permanently outcome_unknown, with an explicit
observation telling the model that the action may have happened and must be
checked before a retry. Recovery does not automatically execute that invocation
again. Existing external-effect journals continue to reject an identical write
with an unconfirmed outcome. Shell and browser actions still require inspection
of the external state; this is not an exactly-once guarantee for remote effects.
Completed observations retain their source restrictions. Memory revision/entry
mapping and this turn’s mutation acknowledgements survive both a round checkpoint
and a crash immediately after an individual observation. Atomic remember
retains its revision identity alongside the memory effect and observation.
Ownership, continuation and delivery
Checkpoint transitions and tool settlement lock the acquired run owner and check its lease immediately before commit. A stale worker cannot settle a tool, move the frontier or terminate a run adopted by another worker. Remote calls already in flight cannot be fenced by a Postgres transaction. Persona, live context, tool registry, grants, configuration, model admission and audience checks are rebuilt on resume. Working transcript messages do not contain saved system/developer instructions or raw image bytes. Images are omitted with an explicit notice; their original audit/artifact storage is separate. Unrecognized checkpoint versions, malformed payloads and mismatched audit rows fail closed. Runs created before checkpoints existed retain the recorded-attempt recovery path. The continuation session is restored before task resolution so its consumed continuation and workspace lineage keep the same identity. Process shutdown or lease loss leaves the run’s workspace available for recovery; ordinary terminal cleanup and the workspace reaper still apply. A workspace that has expired or become unreachable is reported through the existing workspace recovery path. The existing reply commit boundary remains write-ahead of delivery. A crash around that boundary can leave a missing or uncertain reply; Mesh refuses blind resending.reply_ready avoids repeating inference before that boundary and
does not change the delivery guarantee.
Checkpoint payloads are bounded to 4 MiB. Oversized tool observations are visibly
excerpted in the recovery copy; their full bodies remain in the tool ledger and
the live request is unchanged. The continuation’s working transcript is stored
once and rehydrated from that copy. Requester input and model decisions are never
silently shortened to fit. Encoding failures close the opened audit rows.
Terminal-run checkpoints expire under
the configured model_call_body_days retention window, in bounded batches;
active runs retain their recovery frontier. A zero window keeps checkpoints.
The process-exit integration test in
internal/app/checkpoint_restart_integration_test.go kills a subprocess without
unwinding defers at eight execution boundaries and resumes through a fresh run
worker against Postgres. It checks model/tool counts, stable budget accounting,
unknown write outcomes and one final delivery. Lease, attribution and retention
tests cover the accompanying safety boundaries.