Skip to main content
Mesh stores an explicit execution frontier in Postgres for each active run. With MESH_RUN_WORKER_ENABLED=true (the default), a replacement worker claims the same run after its lease expires and resumes from that frontier. No external workflow service is required. Recovery is bounded by the run’s existing attempt, iteration, tool-call and token limits.

Durable boundaries

run_checkpoints has one versioned row per run, referring to the current run_steps and model_calls audit rows. Its provider-neutral payload contains the working transcript, normalized model decision, tool activations, source lineage, committed memory operations, continuation identity and, when available, the completed turn. The prompt and context reports are retained for accounting; the saved prompt is never used as live instructions. The steering mailbox revision accompanies a model decision. Recovery checks for requester corrections before accepting its reply or executing its tools. Every individual tool outcome is recorded before the round advances. If the process dies between one tool’s completion and the next frontier, recovery reconstructs that round from its existing call rows without reserving tool budget again. Successful and failed outcomes are reused. An admitted call that never started may execute after current validation and authority checks. An interrupted inference can have incurred provider cost before its result was recorded. Recovery cannot reconstruct those unknown tokens; the old attempt and its iteration reservation remain visible, and only recorded usage is billed in the ledger. A restart does not restore consumed budget. The existing worker grants a resumed attempt its admitted wall-clock allowance again. When no inference can finish, a deterministic fallback can close the interrupted model/step rows and save its status atomically without inventing a model result.

Interrupted tools

A started call with no durable outcome is retried only when both its original and current declarations say Replayable, their fingerprints match, and there is no unconfirmed effect intent. The fingerprint covers schema, execution class, replay policy, required capabilities and internal-effect classification. Declarations must accurately describe handler behavior. All other started calls become permanently outcome_unknown, with an explicit observation telling the model that the action may have happened and must be checked before a retry. Recovery does not automatically execute that invocation again. Existing external-effect journals continue to reject an identical write with an unconfirmed outcome. Shell and browser actions still require inspection of the external state; this is not an exactly-once guarantee for remote effects. Completed observations retain their source restrictions. Memory revision/entry mapping and this turn’s mutation acknowledgements survive both a round checkpoint and a crash immediately after an individual observation. Atomic remember retains its revision identity alongside the memory effect and observation.

Ownership, continuation and delivery

Checkpoint transitions and tool settlement lock the acquired run owner and check its lease immediately before commit. A stale worker cannot settle a tool, move the frontier or terminate a run adopted by another worker. Remote calls already in flight cannot be fenced by a Postgres transaction. Persona, live context, tool registry, grants, configuration, model admission and audience checks are rebuilt on resume. Working transcript messages do not contain saved system/developer instructions or raw image bytes. Images are omitted with an explicit notice; their original audit/artifact storage is separate. Unrecognized checkpoint versions, malformed payloads and mismatched audit rows fail closed. Runs created before checkpoints existed retain the recorded-attempt recovery path. The continuation session is restored before task resolution so its consumed continuation and workspace lineage keep the same identity. Process shutdown or lease loss leaves the run’s workspace available for recovery; ordinary terminal cleanup and the workspace reaper still apply. A workspace that has expired or become unreachable is reported through the existing workspace recovery path. The existing reply commit boundary remains write-ahead of delivery. A crash around that boundary can leave a missing or uncertain reply; Mesh refuses blind resending. reply_ready avoids repeating inference before that boundary and does not change the delivery guarantee. Checkpoint payloads are bounded to 4 MiB. Oversized tool observations are visibly excerpted in the recovery copy; their full bodies remain in the tool ledger and the live request is unchanged. The continuation’s working transcript is stored once and rehydrated from that copy. Requester input and model decisions are never silently shortened to fit. Encoding failures close the opened audit rows. Terminal-run checkpoints expire under the configured model_call_body_days retention window, in bounded batches; active runs retain their recovery frontier. A zero window keeps checkpoints. The process-exit integration test in internal/app/checkpoint_restart_integration_test.go kills a subprocess without unwinding defers at eight execution boundaries and resumes through a fresh run worker against Postgres. It checks model/tool counts, stable budget accounting, unknown write outcomes and one final delivery. Lease, attribution and retention tests cover the accompanying safety boundaries.