Skip to main content

Snapshot isolation and lineage

Warm project snapshots save clone, install, and build time. They must not turn one task’s mutable checkout into another task’s starting state. This document defines the control-plane policy above the provider-neutral snapshot operations in warm-workspace-snapshotting.md (repository). Modal images, Fly snapshots, and future backend primitives may differ, but they all implement the same rule:
New work starts from a clean repository baseline. Continuing work resumes only from that work’s own checkpoint lineage.
A snapshot is immutable filesystem state. Lineage is Mesh metadata deciding which work may restore it. The backend owns capture, restore, and deletion; the workspace manager owns classification, selection, and lifecycle.

The failure this prevents

A workspace remains isolated per Run, so concurrent Runs do not directly share a mutable directory. That is necessary but insufficient. A single “latest snapshot for this agent and repository” slot creates an indirect collision:
  1. Run A restores the repository, creates feature/a, and leaves tracked or untracked changes.
  2. Run B independently starts from an earlier repository snapshot.
  3. Run A releases and publishes its dirty tree as the repository’s latest snapshot.
  4. Run C, unrelated to A, restores that snapshot and inherits A’s branch and files.
  5. Run B later releases and replaces the slot again.
No two Runs shared a live workspace, but unrelated work still contaminated each other through the mutable meaning of “latest.” Concurrent release order also made the next starting state nondeterministic. This was the implementation gap until migration 0082_snapshot_lineage.sql: project_snapshots had one ready row per (agent_id, repo, backend_kind, base_image), and release capture replaced it without proving that the tree was a clean canonical baseline or binding it to the work that produced it. The Modal adapter correctly implemented immutable snapshot mechanics; the missing boundary was selection and lineage above it.

What is built

The baseline half of this contract is implemented:
  • project_snapshots carries class, canonical_ref, validation, retired_at and deleted_at, plus the checkpoint columns (work_context_id, sequence, predecessor_id) with the class-shape CHECK and the one-ready-head-per-lineage index, so the checkpoint writer is a code change against a schema that already fits it.
  • Every pre-lineage row was retired, not promoted: none of them was validated, so none is a baseline. The first validated checkout of each repository publishes the first real one.
  • The checkout (internal/coding.Checkout.Perform) gathers the baseline evidence (execution.BaselineEvidence) after fetching and before cutting the task branch and offers the tree to workspace.Manager.PublishBaseline, which is the only writer of a baseline. Release never captures.
  • A workspace seeded from a baseline is normalized to the requested ref before use (reset onto the fetched remote ref, other local branches dropped, untracked files removed, ignored caches kept), and the checkout learns that a baseline seeded it through workspace.WarmHint.Restored.
  • Every restore is attributed in snapshot_restores, and the workspace sweeper reclaims the artifacts of superseded, retired and failed rows past a seven-day retention window (Reaper.SweepSnapshots), each through the backend that minted it (resolved by backend_kind, like release; see sandboxed-execution.md).
The work checkpoint half is implemented too:
  • Every turn names its work context (workspace.WithWorkContext): the root run of a paused-and-resumed continuation, so every run that continues the same task shares one lineage across run ids; otherwise the run itself. It is derived from durable task state, never from the conversation. A commitment that continues in a later run without a paused continuation starts a new lineage and restores the baseline; the resumed agent inspects durable external state before repeating writes, as the restore decision requires.
  • Acquire restores the head of the work’s own lineage first (GetCheckpointHead), then the repository baseline, then cold-starts. A missing lineage never falls to another work context’s checkpoint. A checkpoint restore is reported as RestoredCheckpoint and the checkout leaves the tree exactly as it is — it is this work’s own state.
  • The continuation carries repository/directory routing across process restarts, so an exec-first resume can restore its checkpoint and capture another without requiring the agent to repeat checkout.
  • Paused work on an idle-persistent backend is held for fourteen days. Fly Sprites uses this live-filesystem path: its adapter still does not support project snapshots. Other backends retain the one-hour handoff window and use their checkpoints after it expires.
  • A paused turn appends a checkpoint (Manager.Checkpoint) after the continuation is saved and before the paused reply is delivered. The ordering is the serialization: nothing makes the continuation consumable until the delivered message carries its stamp, so no “yes” can hand the live workspace to a child run while the snapshot is in flight. The append is fenced on the head observed before capture — a head that moved or vanished loses the fence — and advances sequence and predecessor atomically (AppendCheckpoint); a lost fence reclaims the artifact.
  • A restore that fails walks down the ladder: a checkpoint that will not restore falls to the repository baseline, a baseline that will not restore falls to cold. A requester who declines the paused task ends its lineage.
  • A resumed turn reports whether the original task is complete separately from its reply text. An answer to a side question or an unstructured legacy response keeps the continuation pending; answering a message does not retire the task. A classified unrelated request leaves the earlier continuation alone.
  • A turn that explicitly completes the original task retires the lineage (Manager.RetireLineage); the workspace sweep retires heads neither captured nor restored within fourteen days (Reaper.RetireStaleCheckpoints) and reclaims retired artifacts after the seven-day window.
Not built: a final checkpoint at a non-terminal release (the pause-time checkpoint is the recovery path; a workspace released by the reaper after its hold expired has already been checkpointed at the pause), and an explicit, authorized retry that forks a new work context from a selected checkpoint.

Two snapshot classes

Mesh recognizes two classes with different eligibility and lookup keys.

Repository baseline

A repository baseline is reusable input for new, unrelated work. It is keyed by:
  • owning RuntimeAgent
  • repository identity
  • backend kind
  • immutable base-image identity
  • canonical upstream ref, normally the repository’s default branch
A baseline is publishable only after deterministic validation establishes that:
  • the remote-tracking ref has been fetched successfully
  • HEAD equals the resolved commit of the canonical upstream ref
  • the canonical branch is checked out, not a task branch or detached task commit
  • the index has no staged changes
  • tracked files have no modifications
  • there are no non-ignored untracked files
  • no merge, rebase, cherry-pick, or bisect is in progress
  • the captured subtree is the repository project-state root and excludes grants, credentials, and unrelated repositories
Ignored dependency and build-cache directories may remain: preserving them is part of the latency win. They must be reproducible cache state, not the only copy of a task artifact. A baseline’s commit may be behind the remote by the time it is restored. That is safe because restore is followed by fetch and deterministic normalization to the requested canonical ref before a task branch is created. Once validated and captured, a baseline is immutable. Publishing a newer baseline supersedes it by an atomic compare-and-swap or equivalent transaction; an older concurrent publisher must not replace a newer commit merely because it released later.

Work checkpoint

A work checkpoint preserves one in-progress work context, including its branch, local commits, staged changes, untracked files, dependencies, and build cache. Dirty state is expected here. A checkpoint is not eligible to seed new work. Every independent unit of work receives a durable work_context_id before its first workspace acquisition. A WorkGraph child that mutates a repository gets its own context unless it is the explicit join owner. The exact parent object may be a Run, continuation, Commitment, or WorkNode; conversation identity alone is not sufficient because one conversation can contain several independent objectives and one objective can span several Runs. A lineage is keyed by:
  • owning RuntimeAgent
  • work_context_id
  • repository identity
  • backend kind
  • immutable base-image identity
Each checkpoint records its sequence and predecessor, plus the baseline or earlier checkpoint from which the lineage began. Only one head is current for a lineage. Publication must be fenced by the workspace lease and advance the head atomically, so a stale or concurrently resumed worker cannot replace a newer checkpoint.

Restore decision

Workspace acquisition must make the distinction explicitly:
Selection never falls back from a missing lineage checkpoint to another work context’s checkpoint. It may fall back to that lineage’s recorded clean baseline or to a cold checkout, but the resumed agent must then inspect durable external state before repeating writes. A snapshot is filesystem evidence, not proof that an external effect did or did not happen. The baseline restored for new work is copy-in state. Ten concurrent tasks may restore the same immutable baseline, but each receives a fresh workspace and a new lineage. Their later checkpoints cannot affect baseline selection or one another’s restore selection. A user asking a new question in the same conversation does not implicitly mean “continue the last workspace.” Continuation must come from durable task/run state that names the work context. Conversely, a Commitment continuing in a later Run keeps its work context even though the Run id changes.

Capture policy

Capture at boundaries that improve recovery without changing classification:
  • After a cold checkout or baseline restore has fetched and passed baseline validation, Mesh may publish a repository baseline before task mutation.
  • At each completed turn or other durable pause of non-terminal work, Mesh may append a work checkpoint and atomically advance that lineage’s head.
  • Before releasing a workspace that still owns non-terminal work, Mesh should attempt a final checkpoint. Capture failure is visible but does not relabel or publish partial state as a baseline.
  • An arbitrary workspace release must never update the repository baseline. Baseline promotion is a separate validated operation, not a side effect of release order.
If a live workspace is safely held for immediate continuation, it remains the cheapest resume path. The latest durable checkpoint is the recovery path when that lease expires, a worker dies, or the provider restarts. Holding a workspace and snapshotting it are complementary optimizations, not aliases.

Completion and retirement

Once work reaches a verified terminal state and its durable outputs exist outside the workspace, Mesh does not need another checkpoint merely to preserve a disposable local tree. It should:
  1. mark the lineage terminal
  2. release the live workspace
  3. retire its work checkpoints from restore eligibility
  4. asynchronously delete provider artifacts under retention policy
Terminal checkpoints may be retained briefly for audit or operator recovery, but they must never become repository baselines automatically. If completed work has merged into the canonical branch, a later clean baseline refresh will capture it from upstream through the normal validation path. Cancellation and failure follow the same eligibility rule: their checkpoints may be retained according to policy, but no unrelated work may restore them. An explicit, authorized retry may continue the same lineage or fork a new work context from a selected checkpoint; that decision is recorded rather than inferred from recency.

Persistence shape

The schema may evolve, but these concepts must be representable and enforced by constraints rather than naming conventions:
  • snapshot class: baseline or checkpoint
  • immutable provider reference and backend/base-image scope
  • repository and owning RuntimeAgent
  • canonical ref and resolved commit for a baseline
  • work_context_id, predecessor, and monotonic sequence for a checkpoint
  • source Run/workspace attribution
  • lifecycle state: capturing, ready, superseded/retired, failed, deleting
  • deterministic validation evidence for baseline publication
  • restore attribution: which snapshot seeded which workspace
There is at most one ready baseline head per baseline key and one ready checkpoint head per lineage key. Those are separate uniqueness domains. A checkpoint can never satisfy a baseline lookup, even if its Git status happens to be clean. Migration 0082 gives project_snapshots exactly this shape: class, the (agent_id, repo, backend_kind, base_image, canonical_ref) baseline key with a partial unique index over ready baselines, work_context_id/sequence/ predecessor_id with a partial unique index over ready checkpoint heads, a class-shape CHECK so a row cannot claim to be both, validation (the evidence as JSON), retired_at, deleted_at (the artifact reclaim), and the snapshot_restores table for restore attribution. Provider refs remain opaque; no lineage field belongs inside a Modal or Fly reference.

Provider mapping

The policy is provider-independent:
  • Modal may represent each capture as an immutable image mounted into a new sandbox.
  • Fly may represent it as a volume or machine snapshot when its adapter can prove the same subtree, writability, isolation, and deletion semantics.
  • A backend without snapshot support cold-starts when no held workspace remains. This loses local-only work; live retention is not an independent backup.
Providers do not choose whether a ref is a baseline or checkpoint. They mint an opaque immutable artifact; Mesh records what the artifact means and where it is eligible to restore.

Relationship to other runtime state

  • A Run is one execution turn and owns a live workspace lease.
  • A work context is the isolation identity for mutable filesystem work and may span Runs when continuation is explicit.
  • A Commitment may carry a work context across autonomous attempts.
  • A WorkGraph node that mutates code gets an isolated work context; a join node owns reconciliation rather than allowing siblings to share a directory.
  • A conversation supplies attributed context but is not a workspace lineage.
  • A snapshot preserves filesystem state; it does not replace transcript, effect-journal, provenance, or authorization state.

Design guardrails

  • do not restore “the latest snapshot” without first selecting a snapshot class and exact eligibility key
  • do not publish a task branch or dirty tree as a repository baseline
  • do not use conversation id alone as work identity
  • do not let unrelated work restore a checkpoint lineage
  • do not let concurrent release order choose the next baseline
  • do not let a stale worker advance a checkpoint lineage without a lease fence
  • do not promote a checkpoint to baseline because Git status is clean
  • do not infer external-effect outcomes from restored filesystem state
  • do not share a mutable workspace between parallel work contexts
  • do not put Mesh lineage metadata inside provider-opaque references
  • do not require a final checkpoint after verified terminal completion solely to preserve disposable local state
The important branching is not a provider feature and not a Git metaphor. It is a control-plane invariant: immutable baseline fan-out for independent work, and isolated checkpoint chains for continuation.

Paused-work durability limits

Live holds are bounded, not backups. Fly’s IdlePersistence capability keeps an unreleased Sprite’s filesystem available while compute sleeps; Release still destroys it. Explicit cancellation, retention expiry, credential revocation, or provider-side loss can remove that tree. Native project snapshots remain a separate capability, currently supported by Modal. Neither path replaces pushing commits and saving final artifacts outside a workspace. The fourteen-day hold applies to newly saved pauses. Existing one-hour holds are not retroactively extended. Autonomous commitments without a paused continuation still start separate workspace lineages; this change concerns human-authorized continuations and does not make a conversation a shared mutable workspace.