Skip to main content
Durable intake records accepted connector work before acknowledging it, so a process exit does not silently discard an accepted message. It is separate from run checkpoints, which recover execution after a run has started.

Connector coverage

Slack, Telegram, and phone connectors commit a versioned normalized envelope before returning HTTP 200. An unsuccessful or timed-out insert returns a retryable failure. A duplicate source identity reuses the original accepted payload. Generic ingress still uses its in-process handoff. Its acknowledged intake does not have this queue guarantee. Once a run has started, its execution uses the common run ledger and recovery path. Database-less test paths also use in-process execution.

Work lifecycle

Capture continues while an earlier turn runs. Ready triggers in one conversation coalesce to the newest durable receipt; older messages remain in the graph. Steering can move eligible requester follow-ups into an active run instead of starting another one. Workers re-resolve the current route and reject disabled agents, stopped installs, and missing or reassigned routes. They do not execute an old message as a new owner. A saved engagement classification is reused during capture recovery. Inference can repeat if a process dies before its classification result commits.

Queue ownership and run ownership

Work-item leases govern intake/capture. Run leases govern model/tool execution. The queue never takes over a run by simply resetting its work item.
  • Work-item leases last 60 seconds and renew every 20 seconds. Owner, epoch, phase, and database-time expiry fence transitions.
  • Run creation and work-item binding commit together with a run lease. A live worker renews it; an expired lease permits recovery by the run worker.
  • Expired capture or unstarted execution can be reacquired up to three times per phase. Once a run is bound, queue recovery reconciles its observed status and leaves execution recovery to the run lease.
  • MESH_RUN_WORKER_ENABLED defaults to true. The enabled worker resumes checkpoints; disabling it uses the legacy path that closes interrupted runs instead of resuming their execution.
Capacity defaults to 4 active turns per agent and 16 per installation, controlled by MESH_MAX_TURNS_PER_AGENT and MESH_MAX_TURNS. Use consistent settings across processes. Conversation locks and the active-run uniqueness constraint prevent overlapping execution in one agent/conversation. Current ownership checks cannot retract a remote call already in flight. Queue leases and checkpoints do not provide exactly-once external effects or prove that every aspect of a deployment is ready for multiple replicas.

Bounds and inspection

Each agent can have 256 nonterminal work items; excess new work gets a retryable failure instead of a successful acknowledgement. Duplicates still succeed when the backlog is full. An envelope is limited to 1 MiB. The five-minute trigger staleness window starts at durable receipt. Stale work is retired explicitly after graph capture. Terminal envelope copies are scrubbed after seven days; deduplication identities, references, and outcomes remain. Conversation history and run-body retention follow their separate policies. There is no public queue retry API. Operators can inspect content-free status:
Do not reset a bound job to ready: its run may already have performed effects. Inspect the run and repair the cause before requesting new work.

Deployment requirements

At least one runtime must stay running with CPU available between requests. Queue pickup, lease renewal, schedules, and recovery are background processes; a request-lifetime function or a scale-to-zero service cannot supply them. Keep the database and required master key through upgrades. When upgrading from an old binary that predates durable intake, drain it and transfer connector endpoints exclusively to the migrated runtime. The legacy binary cannot consume this queue and must not share its traffic during that cutover. Implementation: internal/workqueue (repository), internal/app/durable_queue.go (repository), and internal/app/run_worker.go (repository).