> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Turn checkpoints

> Resume interrupted runs from durable boundaries without blindly repeating effects.

Mesh stores an explicit execution frontier in Postgres for each active run.
With `MESH_RUN_WORKER_ENABLED=true` (the default), a replacement worker claims
the same run after its lease expires and resumes from that frontier. No external
workflow service is required. Recovery is bounded by the run's existing attempt,
iteration, tool-call and token limits.

## Durable boundaries

`run_checkpoints` has one versioned row per run, referring to the current
`run_steps` and `model_calls` audit rows. Its provider-neutral payload contains
the working transcript, normalized model decision, tool activations, source
lineage, committed memory operations, continuation identity and, when available,
the completed turn. The prompt and context reports are retained for accounting;
the saved prompt is never used as live instructions.
The steering mailbox revision accompanies a model decision. Recovery checks
for requester corrections before accepting its reply or executing its tools.

| Phase | What is durable | Recovery |
| - | - | - |
| `model_pending` | Input, iteration reservation and an open model call | Close the interrupted attempt and reserve a new inference against the remaining budget. |
| `tools_pending` | Model result, usage and exact requested calls, committed in one transaction | Reuse the decision and the same durable tool-call IDs. |
| `model_ready` | Closed tool round and its resulting working transcript | Make the next inference without repeating that round. |
| `reply_ready` | Completed turn, accounting and continuation state | Reuse the turn and pass it through the normal live delivery checks. |

Every individual tool outcome is recorded before the round advances. If the
process dies between one tool's completion and the next frontier, recovery
reconstructs that round from its existing call rows without reserving tool
budget again. Successful and failed outcomes are reused. An admitted call that
never started may execute after current validation and authority checks.

An interrupted inference can have incurred provider cost before its result was
recorded. Recovery cannot reconstruct those unknown tokens; the old attempt and
its iteration reservation remain visible, and only recorded usage is billed in
the ledger. A restart does not restore consumed budget. The existing worker
grants a resumed attempt its admitted wall-clock allowance again.
When no inference can finish, a deterministic fallback can close the interrupted
model/step rows and save its status atomically without inventing a model result.

## Interrupted tools

A started call with no durable outcome is retried only when both its original
and current declarations say `Replayable`, their fingerprints match, and there
is no unconfirmed effect intent. The fingerprint covers schema, execution class,
replay policy, required capabilities and internal-effect classification.
Declarations must accurately describe handler behavior.

All other started calls become permanently `outcome_unknown`, with an explicit
observation telling the model that the action may have happened and must be
checked before a retry. Recovery does not automatically execute that invocation
again. Existing external-effect journals continue to reject an identical write
with an unconfirmed outcome. Shell and browser actions still require inspection
of the external state; this is not an exactly-once guarantee for remote effects.

Completed observations retain their source restrictions. Memory revision/entry
mapping and this turn's mutation acknowledgements survive both a round checkpoint
and a crash immediately after an individual observation. Atomic `remember`
retains its revision identity alongside the memory effect and observation.

## Ownership, continuation and delivery

Checkpoint transitions and tool settlement lock the acquired run owner and check
its lease immediately before commit. A stale worker cannot settle a tool, move
the frontier or terminate a run adopted by another worker. Remote calls already
in flight cannot be fenced by a Postgres transaction.

Persona, live context, tool registry, grants, configuration, model admission and
audience checks are rebuilt on resume. Working transcript messages do not contain
saved system/developer instructions or raw image bytes. Images are omitted with
an explicit notice; their original audit/artifact storage is separate.
Unrecognized checkpoint versions, malformed payloads and mismatched audit rows
fail closed. Runs created before checkpoints existed retain the recorded-attempt
recovery path.

The continuation session is restored before task resolution so its consumed
continuation and workspace lineage keep the same identity. Process shutdown or
lease loss leaves the run's workspace available for recovery; ordinary terminal
cleanup and the workspace reaper still apply. A workspace that has expired or
become unreachable is reported through the existing workspace recovery path.

The existing reply commit boundary remains write-ahead of delivery. A crash
around that boundary can leave a missing or uncertain reply; Mesh refuses blind
resending. `reply_ready` avoids repeating inference before that boundary and
does not change the delivery guarantee.

Checkpoint payloads are bounded to 4 MiB. Oversized tool observations are visibly
excerpted in the recovery copy; their full bodies remain in the tool ledger and
the live request is unchanged. The continuation's working transcript is stored
once and rehydrated from that copy. Requester input and model decisions are never
silently shortened to fit. Encoding failures close the opened audit rows.
Terminal-run checkpoints expire under
the configured `model_call_body_days` retention window, in bounded batches;
active runs retain their recovery frontier. A zero window keeps checkpoints.

The process-exit integration test in
`internal/app/checkpoint_restart_integration_test.go` kills a subprocess without
unwinding defers at eight execution boundaries and resumes through a fresh run
worker against Postgres. It checks model/tool counts, stable budget accounting,
unknown write outcomes and one final delivery. Lease, attribution and retention
tests cover the accompanying safety boundaries.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.