> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Egress reliability

# Egress reliability

A human co-worker never types "LLM request failed." Neither should Mesh. The
moment a user sees an internal error string in a conversation, the illusion that
they are talking to a reliable colleague collapses, and it does not come back
cheaply — trust erodes far faster than it rebuilds. So Mesh treats the boundary
between "what the runtime knows" and "what a conversation sees" as a contract, not
a hope.

This document is the normative statement of that contract. It defines what may
cross the egress boundary, how transient failures are absorbed before they get
near it, and what a user sees when Mesh genuinely cannot answer.

> **Status:** specification. This describes the target contract and the layers
> that enforce it. Where a layer is not yet implemented, that is remaining work,
> not a description of current behaviour. See
> [`../roadmap/milestones.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/roadmap/milestones.md) for sequencing.

## The failure taxonomy

The mistake is to treat "the reply looks bad" as one problem and reach for one
tool. It is two problems with two costs, and conflating them buys the expensive
tool for the cheap problem.

### Class 1 — transport / infra failures

`LLM request failed`, OpenRouter 5xx/429, request timeout, empty completion,
context-window overflow, malformed tool-call JSON, connector send error.

The defining property: **Mesh already knows these failed.** We are holding an
exception or a non-200 status code. The classification is deterministic and free.
Spending a fast-model round-trip to read the string "LLM request failed" and
decide whether it is garbage is paying an LLM to re-derive an error code we
already have in hand.

Class 1 is the overwhelming majority of what leaks today, and none of it needs
inference to detect.

### Class 2 — semantic / quality failures

Syntactically valid output that is nonetheless wrong to send: a refusal, a
hallucinated fact, leaked chain-of-thought, a stray system-prompt fragment, badly
wrong tone.

This is the only class where a model-backed judge earns its cost, because the
badness is not visible in a status code — it is in the meaning of the text. Even
here the judge does not run on every message; it is gated behind cheap heuristics
(see [The egress quality gate](#the-egress-quality-gate)).

Keeping these classes separate is the whole design. Class 1 is solved by
determinism and retries at zero happy-path latency. Class 2 is solved by a
narrow, gated, opt-in judge. Neither pays for the other's failure mode.

## The egress contract boundary

This is the structural win, and it is one Mesh can make and a single-user harness
cannot.

Every outgoing message already flows through one place: the outbound send path
in the owning perspective (`SendReplyCommitting` today; see
[`turn-pipeline.md`](/runtime/turn-pipeline) outbound flow). Because that single chokepoint
already exists, the guarantee can live there as a **typed invariant** instead of
scattered across every `try/catch` in the codebase.

> **Persistence order is a design decision this spec does not silently assume.**
> The *current* write path persists the outbound event **after** delivery (see
> [`../persistence.md`](/persistence) for why). A typed egress boundary is
> orthogonal to that ordering — it can enforce "no `EgressError` reaches a user"
> whether the durable row is written before or after the wire send. If a future
> revision moves persistence *before* delivery (to get retry/idempotency off a
> durable record), that is **target behaviour** and must define failure-state,
> retry, idempotency, and duplicate-handling semantics before it is adopted —
> persist-then-send and send-then-persist have materially different partial-delivery
> semantics. This spec does not depend on that change; it only requires the typed
> boundary to sit on the one send path that already exists.

The invariant:

> An egress payload is **either** a deliverable agent message **or** a structured
> error event. A structured error event is **never** rendered to a user channel.
> It is logged, attributed, retried, or escalated — but a connector's user-facing
> `send` path only ever accepts the first variant.

The two variants are distinct types, not one type with a `success` flag that a
caller can forget to check. The connector's user-facing delivery entry point
accepts only the message variant. A structured error cannot be passed to it,
because it does not typecheck — the compiler, not code review, enforces that a
raw internal failure has no path to a user channel.

```go theme={null}
// Illustrative shape; the point is that these are distinct types and the
// user-facing delivery path only accepts the first.
type Egress interface{ isEgress() }

type AgentMessage struct {
    Body        string       // canonical Markdown, pre-render
    Attribution Attribution  // which RuntimeAgent perspective, which conversation
    // …
}

type EgressError struct {
    Class      FailureClass // Transport | Semantic
    Retryable  bool
    Cause      error        // full internal detail, for logs/errtrack only
    Attempt    int
    // …
}

func (AgentMessage) isEgress() {}
func (EgressError) isEgress()  {}

// Deliver is the typed PRE-RENDER boundary. It accepts only AgentMessage; an
// EgressError does not typecheck here and is routed to logs / error-tracking /
// the retry loop instead. Deliver does not bypass the existing connector
// contract: it renders canonical Markdown to Rendered parts and sends each part
// through connector.Outbound.Send, preserving the connector's chunking and
// size-limit handling exactly as SendReplyCommitting does today. The type gate
// is added *around* the existing Render+Send path, not in place of it.
func (c Connector) Deliver(ctx context.Context, msg AgentMessage) (Receipt, error) {
    rendered := c.Render(msg.Body)          // canonical Markdown -> Rendered parts
    // ... cross commit boundary, then Send each rendered part in order,
    //     identical to the current SendReplyCommitting semantics.
}
```

The load-bearing detail: `Deliver` is the *pre-render* boundary. It takes an
`AgentMessage` carrying canonical Markdown and is responsible for calling `Render`
and sending the resulting `Rendered` parts through the connector's `Send` — it
never ships canonical Markdown straight to the wire, and it keeps the existing
per-part chunking and size-limit behaviour. The type gate constrains *what may
enter* the send path; it does not change *how* a valid message is rendered and
chunked.

This is why OpenClaw leaks these and Mesh does not have to. In OpenClaw the
guarantee that no internal string reaches a user lives in the discipline of every
error path independently choosing not to render — one missed `catch` that falls
through to a generic "post the error to the channel" handler defeats it. Here the
guarantee is a property of the boundary's type, so a code path that wants to leak
an error has to be written on purpose, and it will not compile against the
delivery API by accident.

The invariant is also the natural place for the regression net: a test that a
constructed `EgressError` has no route to `Deliver`, in either direction, fails
the day someone adds a bypass.

## Layer A — classify, retry, fall back at the model-call layer

This is where Class 1 is absorbed, before it ever reaches the egress boundary as
anything but a successfully-produced message.

Around every model call:

1. **Classify** the failure as retryable or terminal.
   * *Retryable:* 5xx, 429 (honour `Retry-After`), connection reset, request
     timeout, empty completion, transient JSON-parse on a tool call. (A
     text-protocol tool decision that fails to parse first gets one in-attempt
     repair from the connector, which re-asks with the model's reply and a
     correction; only if that also fails does the transient error reach this
     loop. See `docs/connectors/model.md`.)
   * *Terminal:* 400 with a real validation error, auth failure, a hard context
     overflow that a retry cannot fix without changing the request.
2. **Retry** retryable failures with bounded attempts and exponential backoff
   plus jitter. Bounded on three axes, not one:
   * a **maximum attempt count**;
   * a **maximum per-wait delay** — a provider-supplied `Retry-After` is
     *honoured but clamped* to this ceiling. `Retry-After` is provider-controlled;
     an un-clamped large value lets a provider pin an interactive turn and a worker
     slot open for as long as it likes, which is a denial-of-service the contract
     hands to the upstream. Clamp it;
   * a **total recovery deadline** for the turn. When the deadline is reached,
     recovery stops and proceeds to Layer C regardless of remaining attempts.
     An unbounded retry loop — in attempts *or* in wall-clock — is a new outage, not
     a fix.
3. **Fall back** across routes when retries on the primary are exhausted, or as
   soon as a primary attempt times out (a stalled route): another
   OpenRouter route, another provider, or a smaller/cheaper model for a
   degraded-but-real answer. Fallback order is policy, not hardcoded.
4. Only when retries **and** fallbacks are exhausted does the failure become a
   structured `EgressError` and enter Layer C.

The cost profile is the point: **this costs zero happy-path latency.** Every step
here fires only on a path that has *already* failed. A healthy call returns on the
first attempt and never touches the retry machinery. That is the argument against
reaching for an egress judge first — the judge taxes the 99% healthy path to catch
the 1%; Layer A taxes only the path that already broke.

Retryable/terminal classification, attempt counts, and which route ultimately
answered are recorded on the run ledger so a turn's reliability is
reconstructable after the fact (see [`observability.md`](/runtime/observability)).

### Shipped model deadline policy

`internal/model/retry.go` enforces deadlines on the in-flight completion, not
only between attempts. Process defaults are three attempts per route,
`MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MS=60000`,
`MESH_MODEL_RETRY_TOTAL_DEADLINE_MS=180000`, and
`MESH_MODEL_RETRY_MAX_DELAY_MS=20000`. The total includes attempts and backoff.
Setting attempts to one disables retries, **not** these deadlines.

The attempt timeout is the one knob whose disable is a reachable value:
`MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MS=0` turns off the per-attempt bound, leaving
recovery to the route share, the total deadline, and the connector transport
timeout. Unset, empty, unparseable, or negative values restore the 60000
default — only a literal `0` disables, so a stray `-1` can never silently drop
the bound. Keep the attempt timeout below the total deadline; if it is set at or
above the total, Mesh logs a single boot warning and clamps the effective
attempt timeout down to the total, since otherwise one attempt could consume the
whole budget and the three-attempt contract would collapse to one.

Remaining recovery time is divided among remaining routes. A route that exhausts
its share yields to the next configured fallback; unused time remains available
to later routes. Each attempt is capped by the attempt timeout, route share, and
caller deadline. Whenever the total deadline is at least the attempt timeout
times the number of routes (true for the defaults with up to two fallbacks),
that split leaves every fallback at least one full attempt timeout. Caller
cancellation, run deadline, or ownership loss stops recovery.

A timed-out attempt is treated as a **stalled route**, not a momentary error:

* **With a later route configured**, the route is abandoned at its first
  attempt timeout and the next fallback starts immediately; there is no second
  attempt on the stalled route. This was chosen over "one more try" because
  production stalls come in bursts: without it, a stalled primary spent attempt
  after attempt at one attempt timeout each (\~3 minutes at the defaults) and the
  chain was never reached.
* **With no later route** (no fallback configured, or the last fallback), a
  timed-out attempt is retried on the same route like any retryable failure,
  bounded by the attempt count, route share, and total deadline. There is
  nothing else to try, and failing earlier would not give the user a better
  outcome.
* **Non-timeout retryable failures** (429, 5xx, transport resets) always retry
  the same route with backoff first; they are usually momentary and the primary
  is the preferred model.

Each route change logs `model fallback started` with `route_index`,
`previous_route_index`, and `reason` (`attempt_timeout`, `route_budget_spent`,
or `retries_exhausted`), so operators can see the chain advance and why.
**None of this takes effect until `MESH_MODEL_FALLBACK_MODELS` is set**: with
the default empty list the primary is the only route and a stall still costs up
to the total deadline. All provider
connectors must honor context cancellation; no abandoned request goroutines are
spawned to simulate a timeout.

This trades some slow-model tolerance for earlier recovery: a healthy completion
that needs over 60 seconds now requires an operator to raise the attempt timeout
and total budget together. Connector HTTP timeouts remain an additional ceiling
(OpenRouter: 120 seconds). With many fallback routes, route shares can be shorter
than the configured attempt timeout. Fallback configuration is opt-in via
`MESH_MODEL_FALLBACK_MODELS`; it uses this agent's connector and credentials.
Separate-provider credentials/routing are **not** configured by this change.

Only the model request is retried, never completed tool calls. Prepared primary
requests retain their captured body; fallback requests are encoded for the fallback
model. Structured context-correlated logs report failed attempts and retry/fallback
recovery without request bodies or provider error text. The run ledger still
captures the primary request and aggregate model step; per-attempt durable ledger
rows and fallback wire-body capture are not implemented here.

## Layer B — the egress contract boundary

Layer A's residue — the genuinely unrecoverable failure — arrives at the boundary
described in [The egress contract boundary](#the-egress-contract-boundary) as a
structured `EgressError`. The boundary's job is simply to refuse to render it and
hand it to Layer C. It is a short layer precisely because Layer A did the work; by
the time a failure is here it is real, and the only remaining decision is what the
user sees instead of the truth.

## Layer C — graceful degradation policy

What a user sees when Mesh truly cannot answer depends on the shape of the
conversation, because the right human behaviour does too.

* **Async / proactive contexts** (a scheduled digest, an unprompted heads-up, a
  background job posting a result): **silence beats a fallback.** A human who is
  heads-down and hits a wall does not send "something went wrong" — they just
  don't send. A failed proactive message is logged and retried on the next cycle,
  not announced. This composes with Mesh's existing `observe`/`engage` model: not
  emitting is already a first-class, explainable outcome.
* **Interactive contexts** (a user is waiting on a direct reply): a warm holding
  line — "give me a sec on this, hitting a snag" — or an explicit handoff, never
  the raw internal string. The holding line is a real product decision with a
  follow-up commitment attached (see [`commitments.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/commitments.md)), not a
  reworded error.
* **Never**, in any context, the internal cause. The full detail lives in logs
  and error-tracking (see [`../error-tracking.md`](/error-tracking)), attributed
  to the perspective and run, where an operator can act on it.

The degradation message is itself an `AgentMessage` — a deliberate, attributed,
logged product surface — not an `EgressError` that slipped through. The
distinction matters: "give me a sec" is Mesh choosing what to say; "LLM request
failed" is Mesh failing to choose.

## The egress quality gate

Class 2 is the only place a model-backed judge belongs, and it is deliberately
constrained:

* **Heuristic-gated, not blanket.** Cheap deterministic checks run first — regex
  and simple classifiers for known leak shapes (a "As an AI language model"
  refusal opener, a leaked `<thinking>` block, an apology-only reply to a
  substantive ask). Only a candidate flagged by a heuristic is escalated to the
  fast-model judge. The judge never runs on all traffic.
* **Bounded and non-fatal, but fail-safe on high-confidence leaks.** The judge is
  a model call and can itself fail. Its failure mode is defined and depends on the
  *confidence of the heuristic that flagged the candidate*:

  * **Low-confidence candidates** (tone, a possible-but-ambiguous refusal): a judge
    timeout or error does **not** block the send — it degrades to "send the
    candidate, flag it for offline review." The cost of a false hold outweighs the
    cost of shipping a marginal message.
  * **High-confidence deterministic matches** (a leaked `<thinking>` block, a
    verbatim system-prompt fragment): these are routed to Layer C or a safe
    sanitization path, **never sent on judge failure.** Sending the candidate
    because the judge timed out would leak exactly the content the gate exists to
    block — the fail-open default must not apply to the cases we already know are
    unsafe. For these, the judge only ever *downgrades* a block (confirms it is
    benign); it is never the thing standing between a known leak and the user.

  Either way the judge's own failure is bounded and never becomes a second
  user-facing error. *Who guards the guard* is answered by making the guard's
  failure a no-op for the ambiguous case and a fail-*closed* for the unambiguous
  one, not a second outage in either.
* **Opt-in per install/agent.** The gate has a latency and cost profile; it is a
  policy an operator turns on for surfaces where quality matters more than the
  last 200–800ms, not a global default.

A post-send retract (edit or delete the message after the fact) is explicitly
**not** the primary mechanism. It dodges the pre-send latency tax but trades it
for a visible flicker — the user sees the bad message before it vanishes — which
is its own trust hit. Retract is a last-resort remediation for something that got
past the gate, not the gate itself.

## Relationship to the ingress classifier

The egress quality gate and the ingress engagement classifier are siblings on the
same fast-model plumbing, pointed in opposite directions. Ingress asks "is this
addressed to me / how complex is it?" before a turn (see
[`addressing.md`](/runtime/addressing)); egress asks "is this fit to send?" after one.
Both are cheap-heuristic-first, model-second, and both keep the expensive judgment
off the common path. Building one should reuse the other's fast-model routing,
budget accounting, and failure handling rather than standing up a parallel stack.

## What this is not

* It is **not** an LLM reading every outgoing message. That is the expensive tool
  for the cheap problem; Layer A handles the common case deterministically.
* It is **not** a global gate that adds latency to healthy replies. Layers A and B
  cost nothing on the happy path; the Class 2 gate is opt-in and heuristic-gated.
* It is **not** a promise that Mesh always has an answer. It is a promise that when
  Mesh does not, the user sees a human-shaped silence or a warm holding line —
  never the machine's own error text.

## Design summary

| Layer | Handles | Mechanism | Happy-path cost |
| - | - | - | - |
| A — model-call | Class 1 transport/infra | deterministic classify → bounded retry → route fallback | zero (fires only after failure) |
| B — egress boundary | leak prevention | typed invariant: `EgressError` cannot reach user `send` | zero (type-level) |
| C — degradation | unrecoverable failures | silence (async) / holding line (interactive) | zero (fires only after failure) |
| Quality gate | Class 2 semantic | heuristic filter → gated fast-model judge, non-fatal | opt-in, candidates only |

The one-line version: **make "no internal error reaches a user" a property of the
boundary's type, absorb transient failures with retries that cost nothing when
nothing is wrong, and reserve inference for the one failure class a status code
cannot see.**


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.