> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Snapshot lineage

# Snapshot isolation and lineage

Warm project snapshots save clone, install, and build time. They must not turn
one task's mutable checkout into another task's starting state.

This document defines the control-plane policy above the provider-neutral
snapshot operations in
[`warm-workspace-snapshotting.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/warm-workspace-snapshotting.md). Modal images,
Fly snapshots, and future backend primitives may differ, but they all implement
the same rule:

> **New work starts from a clean repository baseline. Continuing work resumes
> only from that work's own checkpoint lineage.**

A snapshot is immutable filesystem state. Lineage is Mesh metadata deciding
which work may restore it. The backend owns capture, restore, and deletion; the
workspace manager owns classification, selection, and lifecycle.

## The failure this prevents

A workspace remains isolated per Run, so concurrent Runs do not directly share
a mutable directory. That is necessary but insufficient.

A single "latest snapshot for this agent and repository" slot creates an
indirect collision:

1. Run A restores the repository, creates `feature/a`, and leaves tracked or
   untracked changes.
2. Run B independently starts from an earlier repository snapshot.
3. Run A releases and publishes its dirty tree as the repository's latest
   snapshot.
4. Run C, unrelated to A, restores that snapshot and inherits A's branch and
   files.
5. Run B later releases and replaces the slot again.

No two Runs shared a live workspace, but unrelated work still contaminated each
other through the mutable meaning of "latest." Concurrent release order also
made the next starting state nondeterministic.

This was the implementation gap until migration `0082_snapshot_lineage.sql`:
`project_snapshots` had one ready row per `(agent_id, repo, backend_kind,
base_image)`, and release capture replaced it without proving that the tree was
a clean canonical baseline or binding it to the work that produced it. The Modal
adapter correctly implemented immutable snapshot mechanics; the missing boundary
was selection and lineage above it.

### What is built

The **baseline** half of this contract is implemented:

* `project_snapshots` carries `class`, `canonical_ref`, `validation`,
  `retired_at` and `deleted_at`, plus the checkpoint columns
  (`work_context_id`, `sequence`, `predecessor_id`) with the class-shape CHECK
  and the one-ready-head-per-lineage index, so the checkpoint writer is a code
  change against a schema that already fits it.
* Every pre-lineage row was **retired**, not promoted: none of them was
  validated, so none is a baseline. The first validated checkout of each
  repository publishes the first real one.
* The checkout (`internal/coding.Checkout.Perform`) gathers the baseline
  evidence (`execution.BaselineEvidence`) after fetching and before cutting the
  task branch and offers the tree to `workspace.Manager.PublishBaseline`, which
  is the **only** writer of a baseline. Release never captures.
* A workspace seeded from a baseline is **normalized** to the requested ref
  before use (reset onto the fetched remote ref, other local branches dropped,
  untracked files removed, ignored caches kept), and the checkout learns that a
  baseline seeded it through `workspace.WarmHint.Restored`.
* Every restore is attributed in `snapshot_restores`, and the workspace sweeper
  reclaims the artifacts of superseded, retired and failed rows past a
  seven-day retention window (`Reaper.SweepSnapshots`), each through the
  backend that minted it (resolved by `backend_kind`, like release; see
  [sandboxed-execution.md](/runtime/sandboxed-execution)).

The **work checkpoint** half is implemented too:

* Every turn names its **work context** (`workspace.WithWorkContext`): the root
  run of a paused-and-resumed continuation, so every run that continues the
  same task shares one lineage across run ids; otherwise the run itself. It is
  derived from durable task state, never from the conversation. A commitment
  that continues in a later run without a paused continuation starts a new
  lineage and restores the baseline; the resumed agent inspects durable
  external state before repeating writes, as the restore decision requires.
* `Acquire` restores the head of the work's own lineage first
  (`GetCheckpointHead`), then the repository baseline, then cold-starts. A
  missing lineage never falls to another work context's checkpoint. A
  checkpoint restore is reported as `RestoredCheckpoint` and the checkout
  leaves the tree exactly as it is — it is this work's own state.
* The continuation carries repository/directory routing across process restarts,
  so an exec-first resume can restore its checkpoint and capture another without
  requiring the agent to repeat checkout.
* Paused work on an idle-persistent backend is held for fourteen days. Fly
  Sprites uses this live-filesystem path: its adapter still does not support
  project snapshots. Other backends retain the one-hour handoff window and use
  their checkpoints after it expires.
* A paused turn appends a checkpoint (`Manager.Checkpoint`) after the
  continuation is saved and **before** the paused reply is delivered. The
  ordering is the serialization: nothing makes the continuation consumable
  until the delivered message carries its stamp, so no "yes" can hand the live
  workspace to a child run while the snapshot is in flight. The append is
  fenced on the head observed before capture — a head that moved *or vanished*
  loses the fence — and advances sequence and predecessor atomically
  (`AppendCheckpoint`); a lost fence reclaims the artifact.
* A restore that fails walks down the ladder: a checkpoint that will not
  restore falls to the repository baseline, a baseline that will not restore
  falls to cold. A requester who declines the paused task ends its lineage.
* A resumed turn reports whether the **original task** is complete separately
  from its reply text. An answer to a side question or an unstructured legacy
  response keeps the continuation pending; answering a message does not retire
  the task. A classified unrelated request leaves the earlier continuation alone.
* A turn that explicitly completes the original task retires the lineage
  (`Manager.RetireLineage`); the workspace sweep retires heads neither captured
  nor restored within fourteen days (`Reaper.RetireStaleCheckpoints`) and
  reclaims retired artifacts after the seven-day window.

Not built: a final checkpoint at a non-terminal **release** (the pause-time
checkpoint is the recovery path; a workspace released by the reaper after its
hold expired has already been checkpointed at the pause), and an explicit,
authorized retry that forks a new work context from a selected checkpoint.

## Two snapshot classes

Mesh recognizes two classes with different eligibility and lookup keys.

### Repository baseline

A repository baseline is reusable input for **new, unrelated work**. It is keyed
by:

* owning RuntimeAgent
* repository identity
* backend kind
* immutable base-image identity
* canonical upstream ref, normally the repository's default branch

A baseline is publishable only after deterministic validation establishes that:

* the remote-tracking ref has been fetched successfully
* `HEAD` equals the resolved commit of the canonical upstream ref
* the canonical branch is checked out, not a task branch or detached task commit
* the index has no staged changes
* tracked files have no modifications
* there are no non-ignored untracked files
* no merge, rebase, cherry-pick, or bisect is in progress
* the captured subtree is the repository project-state root and excludes grants,
  credentials, and unrelated repositories

Ignored dependency and build-cache directories may remain: preserving them is
part of the latency win. They must be reproducible cache state, not the only
copy of a task artifact.

A baseline's commit may be behind the remote by the time it is restored. That is
safe because restore is followed by fetch and deterministic normalization to the
requested canonical ref before a task branch is created. Once validated and
captured, a baseline is immutable. Publishing a newer baseline supersedes it by
an atomic compare-and-swap or equivalent transaction; an older concurrent
publisher must not replace a newer commit merely because it released later.

### Work checkpoint

A work checkpoint preserves **one in-progress work context**, including its
branch, local commits, staged changes, untracked files, dependencies, and build
cache. Dirty state is expected here. A checkpoint is not eligible to seed new
work.

Every independent unit of work receives a durable `work_context_id` before its
first workspace acquisition. A WorkGraph child that mutates a repository gets
its own context unless it is the explicit join owner. The exact parent object
may be a Run, continuation, Commitment, or WorkNode; conversation identity alone
is not sufficient because one conversation can contain several independent
objectives and one objective can span several Runs.

A lineage is keyed by:

* owning RuntimeAgent
* `work_context_id`
* repository identity
* backend kind
* immutable base-image identity

Each checkpoint records its sequence and predecessor, plus the baseline or
earlier checkpoint from which the lineage began. Only one head is current for a
lineage. Publication must be fenced by the workspace lease and advance the head
atomically, so a stale or concurrently resumed worker cannot replace a newer
checkpoint.

## Restore decision

Workspace acquisition must make the distinction explicitly:

```text theme={null}
if this execution continues an existing work_context_id:
    restore the newest usable checkpoint from that exact lineage
    otherwise restore its recorded baseline and reconstruct/verify external state
else:
    create and durably persist a new work_context_id
    restore the newest validated repository baseline
    otherwise cold-start
    fetch and normalize to the canonical upstream ref
    optionally publish the refreshed clean baseline
    create the task branch for that work context
```

Selection never falls back from a missing lineage checkpoint to another work
context's checkpoint. It may fall back to that lineage's recorded clean baseline
or to a cold checkout, but the resumed agent must then inspect durable external
state before repeating writes. A snapshot is filesystem evidence, not proof that
an external effect did or did not happen.

The baseline restored for new work is copy-in state. Ten concurrent tasks may
restore the same immutable baseline, but each receives a fresh workspace and a
new lineage. Their later checkpoints cannot affect baseline selection or one
another's restore selection.

A user asking a new question in the same conversation does not implicitly mean
"continue the last workspace." Continuation must come from durable task/run
state that names the work context. Conversely, a Commitment continuing in a
later Run keeps its work context even though the Run id changes.

## Capture policy

Capture at boundaries that improve recovery without changing classification:

* After a cold checkout or baseline restore has fetched and passed baseline
  validation, Mesh may publish a repository baseline before task mutation.
* At each completed turn or other durable pause of non-terminal work, Mesh may
  append a work checkpoint and atomically advance that lineage's head.
* Before releasing a workspace that still owns non-terminal work, Mesh should
  attempt a final checkpoint. Capture failure is visible but does not relabel or
  publish partial state as a baseline.
* An arbitrary workspace release must never update the repository baseline.
  Baseline promotion is a separate validated operation, not a side effect of
  release order.

If a live workspace is safely held for immediate continuation, it remains the
cheapest resume path. The latest durable checkpoint is the recovery path when
that lease expires, a worker dies, or the provider restarts. Holding a workspace
and snapshotting it are complementary optimizations, not aliases.

## Completion and retirement

Once work reaches a verified terminal state and its durable outputs exist
outside the workspace, Mesh does not need another checkpoint merely to preserve
a disposable local tree. It should:

1. mark the lineage terminal
2. release the live workspace
3. retire its work checkpoints from restore eligibility
4. asynchronously delete provider artifacts under retention policy

Terminal checkpoints may be retained briefly for audit or operator recovery,
but they must never become repository baselines automatically. If completed work
has merged into the canonical branch, a later clean baseline refresh will
capture it from upstream through the normal validation path.

Cancellation and failure follow the same eligibility rule: their checkpoints
may be retained according to policy, but no unrelated work may restore them.
An explicit, authorized retry may continue the same lineage or fork a new work
context from a selected checkpoint; that decision is recorded rather than
inferred from recency.

## Persistence shape

The schema may evolve, but these concepts must be representable and enforced by
constraints rather than naming conventions:

* snapshot class: `baseline` or `checkpoint`
* immutable provider reference and backend/base-image scope
* repository and owning RuntimeAgent
* canonical ref and resolved commit for a baseline
* `work_context_id`, predecessor, and monotonic sequence for a checkpoint
* source Run/workspace attribution
* lifecycle state: capturing, ready, superseded/retired, failed, deleting
* deterministic validation evidence for baseline publication
* restore attribution: which snapshot seeded which workspace

There is at most one ready baseline head per baseline key and one ready
checkpoint head per lineage key. Those are separate uniqueness domains. A
checkpoint can never satisfy a baseline lookup, even if its Git status happens
to be clean.

Migration `0082` gives `project_snapshots` exactly this shape: `class`, the
`(agent_id, repo, backend_kind, base_image, canonical_ref)` baseline key with a
partial unique index over ready baselines, `work_context_id`/`sequence`/
`predecessor_id` with a partial unique index over ready checkpoint heads, a
class-shape CHECK so a row cannot claim to be both, `validation` (the evidence
as JSON), `retired_at`, `deleted_at` (the artifact reclaim), and the
`snapshot_restores` table for restore attribution. Provider refs remain opaque;
no lineage field belongs inside a Modal or Fly reference.

## Provider mapping

The policy is provider-independent:

* Modal may represent each capture as an immutable image mounted into a new
  sandbox.
* Fly may represent it as a volume or machine snapshot when its adapter can
  prove the same subtree, writability, isolation, and deletion semantics.
* A backend without snapshot support cold-starts when no held workspace remains.
  This loses local-only work; live retention is not an independent backup.

Providers do not choose whether a ref is a baseline or checkpoint. They mint an
opaque immutable artifact; Mesh records what the artifact means and where it is
eligible to restore.

## Relationship to other runtime state

* A **Run** is one execution turn and owns a live workspace lease.
* A **work context** is the isolation identity for mutable filesystem work and
  may span Runs when continuation is explicit.
* A **Commitment** may carry a work context across autonomous attempts.
* A **WorkGraph node** that mutates code gets an isolated work context; a join
  node owns reconciliation rather than allowing siblings to share a directory.
* A **conversation** supplies attributed context but is not a workspace lineage.
* A **snapshot** preserves filesystem state; it does not replace transcript,
  effect-journal, provenance, or authorization state.

## Design guardrails

* do not restore "the latest snapshot" without first selecting a snapshot class
  and exact eligibility key
* do not publish a task branch or dirty tree as a repository baseline
* do not use conversation id alone as work identity
* do not let unrelated work restore a checkpoint lineage
* do not let concurrent release order choose the next baseline
* do not let a stale worker advance a checkpoint lineage without a lease fence
* do not promote a checkpoint to baseline because Git status is clean
* do not infer external-effect outcomes from restored filesystem state
* do not share a mutable workspace between parallel work contexts
* do not put Mesh lineage metadata inside provider-opaque references
* do not require a final checkpoint after verified terminal completion solely to
  preserve disposable local state

The important branching is not a provider feature and not a Git metaphor. It is
a control-plane invariant: **immutable baseline fan-out for independent work,
and isolated checkpoint chains for continuation.**

## Paused-work durability limits

Live holds are bounded, not backups. Fly's `IdlePersistence` capability keeps an
unreleased Sprite's filesystem available while compute sleeps; `Release` still
destroys it. Explicit cancellation, retention expiry, credential revocation, or
provider-side loss can remove that tree. Native project snapshots remain a
separate capability, currently supported by Modal. Neither path replaces pushing
commits and saving final artifacts outside a workspace.

The fourteen-day hold applies to newly saved pauses. Existing one-hour holds are
not retroactively extended. Autonomous commitments without a paused continuation
still start separate workspace lineages; this change concerns human-authorized
continuations and does not make a conversation a shared mutable workspace.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.