Egress reliability
A human co-worker never types “LLM request failed.” Neither should Mesh. The moment a user sees an internal error string in a conversation, the illusion that they are talking to a reliable colleague collapses, and it does not come back cheaply — trust erodes far faster than it rebuilds. So Mesh treats the boundary between “what the runtime knows” and “what a conversation sees” as a contract, not a hope. This document is the normative statement of that contract. It defines what may cross the egress boundary, how transient failures are absorbed before they get near it, and what a user sees when Mesh genuinely cannot answer.
Status: specification. This describes the target contract and the layers
that enforce it. Where a layer is not yet implemented, that is remaining work,
not a description of current behaviour. See
../roadmap/milestones.md (repository) for sequencing.
The failure taxonomy
The mistake is to treat “the reply looks bad” as one problem and reach for one tool. It is two problems with two costs, and conflating them buys the expensive tool for the cheap problem.Class 1 — transport / infra failures
LLM request failed, OpenRouter 5xx/429, request timeout, empty completion,
context-window overflow, malformed tool-call JSON, connector send error.
The defining property: Mesh already knows these failed. We are holding an
exception or a non-200 status code. The classification is deterministic and free.
Spending a fast-model round-trip to read the string “LLM request failed” and
decide whether it is garbage is paying an LLM to re-derive an error code we
already have in hand.
Class 1 is the overwhelming majority of what leaks today, and none of it needs
inference to detect.
Class 2 — semantic / quality failures
Syntactically valid output that is nonetheless wrong to send: a refusal, a hallucinated fact, leaked chain-of-thought, a stray system-prompt fragment, badly wrong tone. This is the only class where a model-backed judge earns its cost, because the badness is not visible in a status code — it is in the meaning of the text. Even here the judge does not run on every message; it is gated behind cheap heuristics (see The egress quality gate). Keeping these classes separate is the whole design. Class 1 is solved by determinism and retries at zero happy-path latency. Class 2 is solved by a narrow, gated, opt-in judge. Neither pays for the other’s failure mode.The egress contract boundary
This is the structural win, and it is one Mesh can make and a single-user harness cannot. Every outgoing message already flows through one place: the outbound send path in the owning perspective (SendReplyCommitting today; see
turn-pipeline.md outbound flow). Because that single chokepoint
already exists, the guarantee can live there as a typed invariant instead of
scattered across every try/catch in the codebase.
Persistence order is a design decision this spec does not silently assume. The current write path persists the outbound event after delivery (seeThe invariant:../persistence.mdfor why). A typed egress boundary is orthogonal to that ordering — it can enforce “noEgressErrorreaches a user” whether the durable row is written before or after the wire send. If a future revision moves persistence before delivery (to get retry/idempotency off a durable record), that is target behaviour and must define failure-state, retry, idempotency, and duplicate-handling semantics before it is adopted — persist-then-send and send-then-persist have materially different partial-delivery semantics. This spec does not depend on that change; it only requires the typed boundary to sit on the one send path that already exists.
An egress payload is either a deliverable agent message or a structured
error event. A structured error event is never rendered to a user channel.
It is logged, attributed, retried, or escalated — but a connector’s user-facing
send path only ever accepts the first variant.
The two variants are distinct types, not one type with a success flag that a
caller can forget to check. The connector’s user-facing delivery entry point
accepts only the message variant. A structured error cannot be passed to it,
because it does not typecheck — the compiler, not code review, enforces that a
raw internal failure has no path to a user channel.
Deliver is the pre-render boundary. It takes an
AgentMessage carrying canonical Markdown and is responsible for calling Render
and sending the resulting Rendered parts through the connector’s Send — it
never ships canonical Markdown straight to the wire, and it keeps the existing
per-part chunking and size-limit behaviour. The type gate constrains what may
enter the send path; it does not change how a valid message is rendered and
chunked.
This is why OpenClaw leaks these and Mesh does not have to. In OpenClaw the
guarantee that no internal string reaches a user lives in the discipline of every
error path independently choosing not to render — one missed catch that falls
through to a generic “post the error to the channel” handler defeats it. Here the
guarantee is a property of the boundary’s type, so a code path that wants to leak
an error has to be written on purpose, and it will not compile against the
delivery API by accident.
The invariant is also the natural place for the regression net: a test that a
constructed EgressError has no route to Deliver, in either direction, fails
the day someone adds a bypass.
Layer A — classify, retry, fall back at the model-call layer
This is where Class 1 is absorbed, before it ever reaches the egress boundary as anything but a successfully-produced message. Around every model call:- Classify the failure as retryable or terminal.
- Retryable: 5xx, 429 (honour
Retry-After), connection reset, request timeout, empty completion, transient JSON-parse on a tool call. (A text-protocol tool decision that fails to parse first gets one in-attempt repair from the connector, which re-asks with the model’s reply and a correction; only if that also fails does the transient error reach this loop. Seedocs/connectors/model.md.) - Terminal: 400 with a real validation error, auth failure, a hard context overflow that a retry cannot fix without changing the request.
- Retryable: 5xx, 429 (honour
- Retry retryable failures with bounded attempts and exponential backoff
plus jitter. Bounded on three axes, not one:
- a maximum attempt count;
- a maximum per-wait delay — a provider-supplied
Retry-Afteris honoured but clamped to this ceiling.Retry-Afteris provider-controlled; an un-clamped large value lets a provider pin an interactive turn and a worker slot open for as long as it likes, which is a denial-of-service the contract hands to the upstream. Clamp it; - a total recovery deadline for the turn. When the deadline is reached, recovery stops and proceeds to Layer C regardless of remaining attempts. An unbounded retry loop — in attempts or in wall-clock — is a new outage, not a fix.
- Fall back across routes when retries on the primary are exhausted, or as soon as a primary attempt times out (a stalled route): another OpenRouter route, another provider, or a smaller/cheaper model for a degraded-but-real answer. Fallback order is policy, not hardcoded.
- Only when retries and fallbacks are exhausted does the failure become a
structured
EgressErrorand enter Layer C.
observability.md).
Shipped model deadline policy
internal/model/retry.go enforces deadlines on the in-flight completion, not
only between attempts. Process defaults are three attempts per route,
MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MS=60000,
MESH_MODEL_RETRY_TOTAL_DEADLINE_MS=180000, and
MESH_MODEL_RETRY_MAX_DELAY_MS=20000. The total includes attempts and backoff.
Setting attempts to one disables retries, not these deadlines.
The attempt timeout is the one knob whose disable is a reachable value:
MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MS=0 turns off the per-attempt bound, leaving
recovery to the route share, the total deadline, and the connector transport
timeout. Unset, empty, unparseable, or negative values restore the 60000
default — only a literal 0 disables, so a stray -1 can never silently drop
the bound. Keep the attempt timeout below the total deadline; if it is set at or
above the total, Mesh logs a single boot warning and clamps the effective
attempt timeout down to the total, since otherwise one attempt could consume the
whole budget and the three-attempt contract would collapse to one.
Remaining recovery time is divided among remaining routes. A route that exhausts
its share yields to the next configured fallback; unused time remains available
to later routes. Each attempt is capped by the attempt timeout, route share, and
caller deadline. Whenever the total deadline is at least the attempt timeout
times the number of routes (true for the defaults with up to two fallbacks),
that split leaves every fallback at least one full attempt timeout. Caller
cancellation, run deadline, or ownership loss stops recovery.
A timed-out attempt is treated as a stalled route, not a momentary error:
- With a later route configured, the route is abandoned at its first attempt timeout and the next fallback starts immediately; there is no second attempt on the stalled route. This was chosen over “one more try” because production stalls come in bursts: without it, a stalled primary spent attempt after attempt at one attempt timeout each (~3 minutes at the defaults) and the chain was never reached.
- With no later route (no fallback configured, or the last fallback), a timed-out attempt is retried on the same route like any retryable failure, bounded by the attempt count, route share, and total deadline. There is nothing else to try, and failing earlier would not give the user a better outcome.
- Non-timeout retryable failures (429, 5xx, transport resets) always retry the same route with backoff first; they are usually momentary and the primary is the preferred model.
model fallback started with route_index,
previous_route_index, and reason (attempt_timeout, route_budget_spent,
or retries_exhausted), so operators can see the chain advance and why.
None of this takes effect until MESH_MODEL_FALLBACK_MODELS is set: with
the default empty list the primary is the only route and a stall still costs up
to the total deadline. All provider
connectors must honor context cancellation; no abandoned request goroutines are
spawned to simulate a timeout.
This trades some slow-model tolerance for earlier recovery: a healthy completion
that needs over 60 seconds now requires an operator to raise the attempt timeout
and total budget together. Connector HTTP timeouts remain an additional ceiling
(OpenRouter: 120 seconds). With many fallback routes, route shares can be shorter
than the configured attempt timeout. Fallback configuration is opt-in via
MESH_MODEL_FALLBACK_MODELS; it uses this agent’s connector and credentials.
Separate-provider credentials/routing are not configured by this change.
Only the model request is retried, never completed tool calls. Prepared primary
requests retain their captured body; fallback requests are encoded for the fallback
model. Structured context-correlated logs report failed attempts and retry/fallback
recovery without request bodies or provider error text. The run ledger still
captures the primary request and aggregate model step; per-attempt durable ledger
rows and fallback wire-body capture are not implemented here.
Layer B — the egress contract boundary
Layer A’s residue — the genuinely unrecoverable failure — arrives at the boundary described in The egress contract boundary as a structuredEgressError. The boundary’s job is simply to refuse to render it and
hand it to Layer C. It is a short layer precisely because Layer A did the work; by
the time a failure is here it is real, and the only remaining decision is what the
user sees instead of the truth.
Layer C — graceful degradation policy
What a user sees when Mesh truly cannot answer depends on the shape of the conversation, because the right human behaviour does too.- Async / proactive contexts (a scheduled digest, an unprompted heads-up, a
background job posting a result): silence beats a fallback. A human who is
heads-down and hits a wall does not send “something went wrong” — they just
don’t send. A failed proactive message is logged and retried on the next cycle,
not announced. This composes with Mesh’s existing
observe/engagemodel: not emitting is already a first-class, explainable outcome. - Interactive contexts (a user is waiting on a direct reply): a warm holding
line — “give me a sec on this, hitting a snag” — or an explicit handoff, never
the raw internal string. The holding line is a real product decision with a
follow-up commitment attached (see
commitments.md(repository)), not a reworded error. - Never, in any context, the internal cause. The full detail lives in logs
and error-tracking (see
../error-tracking.md), attributed to the perspective and run, where an operator can act on it.
AgentMessage — a deliberate, attributed,
logged product surface — not an EgressError that slipped through. The
distinction matters: “give me a sec” is Mesh choosing what to say; “LLM request
failed” is Mesh failing to choose.
The egress quality gate
Class 2 is the only place a model-backed judge belongs, and it is deliberately constrained:-
Heuristic-gated, not blanket. Cheap deterministic checks run first — regex
and simple classifiers for known leak shapes (a “As an AI language model”
refusal opener, a leaked
<thinking>block, an apology-only reply to a substantive ask). Only a candidate flagged by a heuristic is escalated to the fast-model judge. The judge never runs on all traffic. -
Bounded and non-fatal, but fail-safe on high-confidence leaks. The judge is
a model call and can itself fail. Its failure mode is defined and depends on the
confidence of the heuristic that flagged the candidate:
- Low-confidence candidates (tone, a possible-but-ambiguous refusal): a judge timeout or error does not block the send — it degrades to “send the candidate, flag it for offline review.” The cost of a false hold outweighs the cost of shipping a marginal message.
- High-confidence deterministic matches (a leaked
<thinking>block, a verbatim system-prompt fragment): these are routed to Layer C or a safe sanitization path, never sent on judge failure. Sending the candidate because the judge timed out would leak exactly the content the gate exists to block — the fail-open default must not apply to the cases we already know are unsafe. For these, the judge only ever downgrades a block (confirms it is benign); it is never the thing standing between a known leak and the user.
- Opt-in per install/agent. The gate has a latency and cost profile; it is a policy an operator turns on for surfaces where quality matters more than the last 200–800ms, not a global default.
Relationship to the ingress classifier
The egress quality gate and the ingress engagement classifier are siblings on the same fast-model plumbing, pointed in opposite directions. Ingress asks “is this addressed to me / how complex is it?” before a turn (seeaddressing.md); egress asks “is this fit to send?” after one.
Both are cheap-heuristic-first, model-second, and both keep the expensive judgment
off the common path. Building one should reuse the other’s fast-model routing,
budget accounting, and failure handling rather than standing up a parallel stack.
What this is not
- It is not an LLM reading every outgoing message. That is the expensive tool for the cheap problem; Layer A handles the common case deterministically.
- It is not a global gate that adds latency to healthy replies. Layers A and B cost nothing on the happy path; the Class 2 gate is opt-in and heuristic-gated.
- It is not a promise that Mesh always has an answer. It is a promise that when Mesh does not, the user sees a human-shaped silence or a warm holding line — never the machine’s own error text.
Design summary
The one-line version: make “no internal error reaches a user” a property of the
boundary’s type, absorb transient failures with retries that cost nothing when
nothing is wrong, and reserve inference for the one failure class a status code
cannot see.