Snapshot isolation and lineage
Warm project snapshots save clone, install, and build time. They must not turn one task’s mutable checkout into another task’s starting state. This document defines the control-plane policy above the provider-neutral snapshot operations inwarm-workspace-snapshotting.md (repository). Modal images,
Fly snapshots, and future backend primitives may differ, but they all implement
the same rule:
New work starts from a clean repository baseline. Continuing work resumes only from that work’s own checkpoint lineage.A snapshot is immutable filesystem state. Lineage is Mesh metadata deciding which work may restore it. The backend owns capture, restore, and deletion; the workspace manager owns classification, selection, and lifecycle.
The failure this prevents
A workspace remains isolated per Run, so concurrent Runs do not directly share a mutable directory. That is necessary but insufficient. A single “latest snapshot for this agent and repository” slot creates an indirect collision:- Run A restores the repository, creates
feature/a, and leaves tracked or untracked changes. - Run B independently starts from an earlier repository snapshot.
- Run A releases and publishes its dirty tree as the repository’s latest snapshot.
- Run C, unrelated to A, restores that snapshot and inherits A’s branch and files.
- Run B later releases and replaces the slot again.
0082_snapshot_lineage.sql:
project_snapshots had one ready row per (agent_id, repo, backend_kind, base_image), and release capture replaced it without proving that the tree was
a clean canonical baseline or binding it to the work that produced it. The Modal
adapter correctly implemented immutable snapshot mechanics; the missing boundary
was selection and lineage above it.
What is built
The baseline half of this contract is implemented:project_snapshotscarriesclass,canonical_ref,validation,retired_atanddeleted_at, plus the checkpoint columns (work_context_id,sequence,predecessor_id) with the class-shape CHECK and the one-ready-head-per-lineage index, so the checkpoint writer is a code change against a schema that already fits it.- Every pre-lineage row was retired, not promoted: none of them was validated, so none is a baseline. The first validated checkout of each repository publishes the first real one.
- The checkout (
internal/coding.Checkout.Perform) gathers the baseline evidence (execution.BaselineEvidence) after fetching and before cutting the task branch and offers the tree toworkspace.Manager.PublishBaseline, which is the only writer of a baseline. Release never captures. - A workspace seeded from a baseline is normalized to the requested ref
before use (reset onto the fetched remote ref, other local branches dropped,
untracked files removed, ignored caches kept), and the checkout learns that a
baseline seeded it through
workspace.WarmHint.Restored. - Every restore is attributed in
snapshot_restores, and the workspace sweeper reclaims the artifacts of superseded, retired and failed rows past a seven-day retention window (Reaper.SweepSnapshots), each through the backend that minted it (resolved bybackend_kind, like release; see sandboxed-execution.md).
- Every turn names its work context (
workspace.WithWorkContext): the root run of a paused-and-resumed continuation, so every run that continues the same task shares one lineage across run ids; otherwise the run itself. It is derived from durable task state, never from the conversation. A commitment that continues in a later run without a paused continuation starts a new lineage and restores the baseline; the resumed agent inspects durable external state before repeating writes, as the restore decision requires. Acquirerestores the head of the work’s own lineage first (GetCheckpointHead), then the repository baseline, then cold-starts. A missing lineage never falls to another work context’s checkpoint. A checkpoint restore is reported asRestoredCheckpointand the checkout leaves the tree exactly as it is — it is this work’s own state.- The continuation carries repository/directory routing across process restarts, so an exec-first resume can restore its checkpoint and capture another without requiring the agent to repeat checkout.
- Paused work on an idle-persistent backend is held for fourteen days. Fly Sprites uses this live-filesystem path: its adapter still does not support project snapshots. Other backends retain the one-hour handoff window and use their checkpoints after it expires.
- A paused turn appends a checkpoint (
Manager.Checkpoint) after the continuation is saved and before the paused reply is delivered. The ordering is the serialization: nothing makes the continuation consumable until the delivered message carries its stamp, so no “yes” can hand the live workspace to a child run while the snapshot is in flight. The append is fenced on the head observed before capture — a head that moved or vanished loses the fence — and advances sequence and predecessor atomically (AppendCheckpoint); a lost fence reclaims the artifact. - A restore that fails walks down the ladder: a checkpoint that will not restore falls to the repository baseline, a baseline that will not restore falls to cold. A requester who declines the paused task ends its lineage.
- A resumed turn reports whether the original task is complete separately from its reply text. An answer to a side question or an unstructured legacy response keeps the continuation pending; answering a message does not retire the task. A classified unrelated request leaves the earlier continuation alone.
- A turn that explicitly completes the original task retires the lineage
(
Manager.RetireLineage); the workspace sweep retires heads neither captured nor restored within fourteen days (Reaper.RetireStaleCheckpoints) and reclaims retired artifacts after the seven-day window.
Two snapshot classes
Mesh recognizes two classes with different eligibility and lookup keys.Repository baseline
A repository baseline is reusable input for new, unrelated work. It is keyed by:- owning RuntimeAgent
- repository identity
- backend kind
- immutable base-image identity
- canonical upstream ref, normally the repository’s default branch
- the remote-tracking ref has been fetched successfully
HEADequals the resolved commit of the canonical upstream ref- the canonical branch is checked out, not a task branch or detached task commit
- the index has no staged changes
- tracked files have no modifications
- there are no non-ignored untracked files
- no merge, rebase, cherry-pick, or bisect is in progress
- the captured subtree is the repository project-state root and excludes grants, credentials, and unrelated repositories
Work checkpoint
A work checkpoint preserves one in-progress work context, including its branch, local commits, staged changes, untracked files, dependencies, and build cache. Dirty state is expected here. A checkpoint is not eligible to seed new work. Every independent unit of work receives a durablework_context_id before its
first workspace acquisition. A WorkGraph child that mutates a repository gets
its own context unless it is the explicit join owner. The exact parent object
may be a Run, continuation, Commitment, or WorkNode; conversation identity alone
is not sufficient because one conversation can contain several independent
objectives and one objective can span several Runs.
A lineage is keyed by:
- owning RuntimeAgent
work_context_id- repository identity
- backend kind
- immutable base-image identity
Restore decision
Workspace acquisition must make the distinction explicitly:Capture policy
Capture at boundaries that improve recovery without changing classification:- After a cold checkout or baseline restore has fetched and passed baseline validation, Mesh may publish a repository baseline before task mutation.
- At each completed turn or other durable pause of non-terminal work, Mesh may append a work checkpoint and atomically advance that lineage’s head.
- Before releasing a workspace that still owns non-terminal work, Mesh should attempt a final checkpoint. Capture failure is visible but does not relabel or publish partial state as a baseline.
- An arbitrary workspace release must never update the repository baseline. Baseline promotion is a separate validated operation, not a side effect of release order.
Completion and retirement
Once work reaches a verified terminal state and its durable outputs exist outside the workspace, Mesh does not need another checkpoint merely to preserve a disposable local tree. It should:- mark the lineage terminal
- release the live workspace
- retire its work checkpoints from restore eligibility
- asynchronously delete provider artifacts under retention policy
Persistence shape
The schema may evolve, but these concepts must be representable and enforced by constraints rather than naming conventions:- snapshot class:
baselineorcheckpoint - immutable provider reference and backend/base-image scope
- repository and owning RuntimeAgent
- canonical ref and resolved commit for a baseline
work_context_id, predecessor, and monotonic sequence for a checkpoint- source Run/workspace attribution
- lifecycle state: capturing, ready, superseded/retired, failed, deleting
- deterministic validation evidence for baseline publication
- restore attribution: which snapshot seeded which workspace
0082 gives project_snapshots exactly this shape: class, the
(agent_id, repo, backend_kind, base_image, canonical_ref) baseline key with a
partial unique index over ready baselines, work_context_id/sequence/
predecessor_id with a partial unique index over ready checkpoint heads, a
class-shape CHECK so a row cannot claim to be both, validation (the evidence
as JSON), retired_at, deleted_at (the artifact reclaim), and the
snapshot_restores table for restore attribution. Provider refs remain opaque;
no lineage field belongs inside a Modal or Fly reference.
Provider mapping
The policy is provider-independent:- Modal may represent each capture as an immutable image mounted into a new sandbox.
- Fly may represent it as a volume or machine snapshot when its adapter can prove the same subtree, writability, isolation, and deletion semantics.
- A backend without snapshot support cold-starts when no held workspace remains. This loses local-only work; live retention is not an independent backup.
Relationship to other runtime state
- A Run is one execution turn and owns a live workspace lease.
- A work context is the isolation identity for mutable filesystem work and may span Runs when continuation is explicit.
- A Commitment may carry a work context across autonomous attempts.
- A WorkGraph node that mutates code gets an isolated work context; a join node owns reconciliation rather than allowing siblings to share a directory.
- A conversation supplies attributed context but is not a workspace lineage.
- A snapshot preserves filesystem state; it does not replace transcript, effect-journal, provenance, or authorization state.
Design guardrails
- do not restore “the latest snapshot” without first selecting a snapshot class and exact eligibility key
- do not publish a task branch or dirty tree as a repository baseline
- do not use conversation id alone as work identity
- do not let unrelated work restore a checkpoint lineage
- do not let concurrent release order choose the next baseline
- do not let a stale worker advance a checkpoint lineage without a lease fence
- do not promote a checkpoint to baseline because Git status is clean
- do not infer external-effect outcomes from restored filesystem state
- do not share a mutable workspace between parallel work contexts
- do not put Mesh lineage metadata inside provider-opaque references
- do not require a final checkpoint after verified terminal completion solely to preserve disposable local state
Paused-work durability limits
Live holds are bounded, not backups. Fly’sIdlePersistence capability keeps an
unreleased Sprite’s filesystem available while compute sleeps; Release still
destroys it. Explicit cancellation, retention expiry, credential revocation, or
provider-side loss can remove that tree. Native project snapshots remain a
separate capability, currently supported by Modal. Neither path replaces pushing
commits and saving final artifacts outside a workspace.
The fourteen-day hold applies to newly saved pauses. Existing one-hour holds are
not retroactively extended. Autonomous commitments without a paused continuation
still start separate workspace lineages; this change concerns human-authorized
continuations and does not make a conversation a shared mutable workspace.