Connector coverage
Slack, Telegram, and phone connectors commit a versioned normalized envelope before returning HTTP 200. An unsuccessful or timed-out insert returns a retryable failure. A duplicate source identity reuses the original accepted payload. Generic ingress still uses its in-process handoff. Its acknowledged intake does not have this queue guarantee. Once a run has started, its execution uses the common run ledger and recovery path. Database-less test paths also use in-process execution.Work lifecycle
Queue ownership and run ownership
Work-item leases govern intake/capture. Run leases govern model/tool execution. The queue never takes over a run by simply resetting its work item.- Work-item leases last 60 seconds and renew every 20 seconds. Owner, epoch, phase, and database-time expiry fence transitions.
- Run creation and work-item binding commit together with a run lease. A live worker renews it; an expired lease permits recovery by the run worker.
- Expired capture or unstarted execution can be reacquired up to three times per phase. Once a run is bound, queue recovery reconciles its observed status and leaves execution recovery to the run lease.
MESH_RUN_WORKER_ENABLEDdefaults totrue. The enabled worker resumes checkpoints; disabling it uses the legacy path that closes interrupted runs instead of resuming their execution.
MESH_MAX_TURNS_PER_AGENT and MESH_MAX_TURNS. Use consistent settings across
processes. Conversation locks and the active-run uniqueness constraint prevent
overlapping execution in one agent/conversation.
Current ownership checks cannot retract a remote call already in flight. Queue
leases and checkpoints do not provide exactly-once external effects or prove
that every aspect of a deployment is ready for multiple replicas.
Bounds and inspection
Each agent can have 256 nonterminal work items; excess new work gets a retryable failure instead of a successful acknowledgement. Duplicates still succeed when the backlog is full. An envelope is limited to 1 MiB. The five-minute trigger staleness window starts at durable receipt. Stale work is retired explicitly after graph capture. Terminal envelope copies are scrubbed after seven days; deduplication identities, references, and outcomes remain. Conversation history and run-body retention follow their separate policies. There is no public queue retry API. Operators can inspect content-free status:ready: its run may already have performed effects.
Inspect the run and repair the cause before requesting new work.
Deployment requirements
At least one runtime must stay running with CPU available between requests. Queue pickup, lease renewal, schedules, and recovery are background processes; a request-lifetime function or a scale-to-zero service cannot supply them. Keep the database and required master key through upgrades. When upgrading from an old binary that predates durable intake, drain it and transfer connector endpoints exclusively to the migrated runtime. The legacy binary cannot consume this queue and must not share its traffic during that cutover. Implementation:internal/workqueue (repository),
internal/app/durable_queue.go (repository),
and internal/app/run_worker.go (repository).