> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Error tracking

# Error tracking

Mesh reports unhandled failures to a pluggable error-tracking backend. It is
opt-in, off by default, and selected by name at runtime. Sentry is the first and
currently only adapter.

The layer lives in `internal/errtrack` (provider-neutral, zero vendor imports)
plus one leaf subpackage per vendor (`internal/errtrack/sentry`). Its boot,
failure, and shutdown behaviour deliberately mirror `internal/telemetry`: opt-in,
never fatal, a zero value that is a true no-op.

## Turning it on

There are two ways, and which one wins is not a matter of ordering luck — see
[Precedence](#-precedence-the-environment-always-wins) below.

### From the settings page

`/settings` → **Error tracking**. Paste a DSN, pick a sample rate, tick *Send
reports*, save. The change takes effect on the running process: no redeploy, no
restart.

The section also carries an **Enable tracing** toggle (with an optional *Traces
sample rate*). It is a per-install opt-in to the provider's OWN APM tracing,
separate from error reports and **off by default** — the posture every install
has had. OTEL stays Mesh's primary latency lens (see
[OTEL stays environment-only](#otel-stays-environment-only)); Sentry tracing is a
supported opt-in for teams that already page on Sentry and want spans in the same
place. It is a per-install setting rather than a hardcode or an infra decision
because Mesh is vendor-neutral and open-source-first: whether a second tracing
pipeline earns its cost is the operator's call, not the binary's.

That is the whole reason this exists. Error tracking's value is being
already-configured when something breaks, and requiring a deploy to turn it on
means the operator who just noticed a silent failure cannot have reports until
after a release — by which time the incident is over and the evidence is gone.

The page also offers:

* **Send a test report.** One synthetic issue through the configuration being
  edited, so a DSN can be checked *before* it goes live. It writes nothing. With
  the field blank it tests whatever is currently in force, which is how an
  env-managed install answers "is my deploy manifest's DSN actually reaching the
  project?".
* **Clear stored DSN.** Deletes the ciphertext (not a tombstone — see 0017) and
  turns reporting off, because a provider with no credential reports nowhere.

🚨 **The DSN is write-only.** No endpoint returns it, from either source. The page
shows only whether one is stored and a redacted hint (`https://…@host/project`),
and the input is always empty — leaving it blank keeps the stored value. A
settings page that echoed the credential would put it in the DOM, in the browser's
autofill store, and in every screenshot.

### From the environment

| Variable | Meaning |
| - | - |
| `MESH_ERROR_TRACKER` | Provider name. **Empty (the default) disables error tracking entirely.** Currently only `sentry`. |
| `MESH_ERROR_TRACKER_DSN` | The provider's endpoint-and-credential string. Required when a provider is named. |
| `MESH_ERROR_TRACKER_SAMPLE_RATE` | Fraction of reports to send, in `(0,1]`. Unset means send everything. |
| `MESH_ERROR_TRACKER_TRACING` | Opt into the provider's own APM tracing. **Unset (the default) is off.** |
| `MESH_ERROR_TRACKER_TRACES_SAMPLE_RATE` | Fraction of traces to send when tracing is on, in `(0,1]`. Unset means send everything. Inert when tracing is off. |

```sh theme={null}
MESH_ERROR_TRACKER=sentry
MESH_ERROR_TRACKER_DSN=https://<key>@<org>.ingest.sentry.io/<project>
# optional
MESH_ERROR_TRACKER_SAMPLE_RATE=1
# optional — off by default; OTEL remains the primary trace lens
MESH_ERROR_TRACKER_TRACING=true
MESH_ERROR_TRACKER_TRACES_SAMPLE_RATE=1
```

Tracing follows the section's precedence like every other field: when the
environment claims the section (see below), it owns the tracing toggle and its
rate too, and the stored row is not consulted — precedence is per SECTION, so
tracing rides the provider and DSN rather than mixing sources.

### 🚨 Precedence: the environment always wins

| `MESH_ERROR_TRACKER` set? | Stored settings? | Effective source | Editable in the UI? |
| - | - | - | - |
| yes | anything | `environment` | **no** — inputs disabled, a `PATCH` is refused with `409 settings_env_managed` |
| no | yes | `database` | yes |
| no | no | `unset` — reporting off | yes |

The rule runs that way round because a variable in a deploy manifest is an
explicit deploy-time decision, and a click in a dashboard must not silently
override it. A setting that can be overridden invisibly is worse than one that
cannot be set at all.

Two consequences worth stating plainly:

* **It applies per SECTION, not per field.** With `MESH_ERROR_TRACKER` set, the
  environment owns the provider *and* the DSN *and* the sample rate — there is no
  mixing, because a half-environment/half-database configuration is one nobody can
  reason about during an incident.
* **The refusal is loud, and the UI agrees with it.** The server answers a
  `PATCH` with `409`, and the settings page disables the inputs and names the
  variables that own the section. A form that accepted an edit the resolver would
  then discard is precisely the dishonesty this design exists to prevent — which
  is why both halves exist: the server refusing is the guarantee, the disabled
  form is the honesty, and each is close to worthless without the other.

One resolver implements this (`settings.ResolveErrorTracking`) and the boot path,
the save path, and the `GET` response all call it. Three copies of a precedence
rule is three answers to "which one is live".

### Where each piece is stored

| Value | Home | Why |
| - | - | - |
| provider, sample rate, enabled, tracing, traces sample rate | `install_settings` row, `value` jsonb | Non-secret. Plaintext, dumped by every backup. A row from an older build lacks the tracing keys and decodes as tracing off. |
| the DSN | `secrets`, `owner_type='install'`, `kind='error_tracker_dsn'` | A credential. Encrypted with `MESH_AGENT_MASTER_KEY`, which must not live in this database. |

🚨 **No credential may ever be written to `install_settings.value`.** The split is
the point of the table, not an implementation detail — `db/migrations/0017_install_settings.sql`
argues it at length, and an integration test asserts the stored document contains
no DSN.

An install-owned secret uses the nil uuid as its `owner_id`, enforced by a CHECK
so two writers cannot pick two sentinels and leave the install with two DSNs, one
of which nothing reads.

### What a save actually does

The reporter is swapped on the running process: build the replacement, publish it
with one atomic store, then flush and close the reporter it replaced — in that
order, so reports captured seconds before the save still reach the vendor.
`internal/app/errtrack_reload.go` explains why an atomic pointer rather than a
mutable field (a torn interface value is a crash, and it would be a crash inside
the panic handler).

If the rows commit but the new reporter cannot be built, the response is **`200`
with `status: "saved_restart_required"`**, not a `500`:

* The write already happened, and telling a client it failed invites a retry of
  something that is already done.
* 🚨 But the value is **stored and not in force** — reporting is still using the
  previous configuration. The settings page says so prominently rather than
  showing "Saved", because for the one feature whose job is telling an operator
  something is wrong, claiming reporting is on when it is not is the worst
  available answer.

A DSN that will not construct is rejected *before* it is persisted (`422`, with the
adapter's own explanation echoed), because failing in the form beats failing during
a real incident.

### OTEL stays environment-only

The settings page shows traces and logs read-only, and that is not an unfinished
form. The OTEL SDK reads `OTEL_*` at exporter **construction** and installs
process-global providers, so changing an endpoint means building new exporters and
re-installing globals under live traffic — genuinely restart-shaped. Offering
inputs would be lying about the size of the change.

Error tracking can hot-swap because its reporter is one interface value behind an
atomic pointer that Mesh owns end to end. That is a property of this layer, not a
difference in ambition.

The error tracker's own **tracing** toggle is not the same thing as OTEL, and
turning it on does not make Sentry Mesh's trace backend. OTEL remains the primary
latency lens; the Sentry-tracing opt-in exists so an install that already pages on
Sentry can co-locate a sampled set of APM spans with its issues, and it defaults
off precisely because a second tracing pipeline is a duplicate cost and a second
attribute policy to police. Because the reporter is rebuilt on save, flipping the
toggle takes effect on the running process — unlike the OTEL exporters above, which
are restart-shaped.

`MESH_ENV` supplies the environment tag, so an issue and a trace agree about
which install they describe. The release tag comes from the binary's embedded
`vcs.revision`, which is what makes "this regressed in build X" work without the
deployment passing a build argument.

An unrecognized provider name is **rejected by name** rather than silently
defaulting or silently disabling — the same rule as an unsupported OTLP protocol.
A deployment that believes it is paging on panics while reporting nowhere is the
worst outcome this feature has, and a typo is the likely cause.

A named provider with no DSN is likewise reported at boot, not treated as a
request to be disabled: naming the provider is an explicit opt-in.

## Why this is not the trace backend

Reasonable question, since both are "send failures somewhere". They answer
different questions and neither substitutes for the other:

| | OTEL traces/logs | Error tracking |
| - | - | - |
| Question | how long did this take, in what order | is this failure new, how often, in which build |
| Shape | one span or record per event | one **issue** with a count, first-seen, and regression marker |
| Grouping | none; you query | fingerprint and dedupe are the product |
| Stack traces | not carried | the primary artifact |
| Alerting | via a backend you configure separately | the reason the layer exists |
| Retention | short, sampled | long, per issue |

The concrete failure that motivated this: ingress acks Slack in milliseconds, so
a turn that fails afterwards produces **no HTTP error anywhere**. A user asked a
question, the agent silently did not answer, and the only trace was one
`slog.Error` line in a stream nobody alerts on. A trace backend cannot answer "is
this new?", and that is the question an operator is woken up by.
`docs/runtime/observability.md` makes the same point from the other direction: the
trace backend is a latency lens, **not the system of record**.

Both stay wired. `docs/runtime/observability.md`'s rule is that no opt-in backend
may be the only witness to anything, so every capture point below also logs.

## 🚨 Why a vendor adapter does not violate the no-vendor rule

`docs/telemetry.md` states, correctly, that no endpoint, vendor name, API key, or
backend-specific field appears anywhere in this repository. `internal/errtrack/sentry`
imports `github.com/getsentry/sentry-go`. That needs reconciling explicitly
rather than by assumption.

**The rule's subject is the neutral core.** What that document is preventing is a
vendor leaking into shared code: a `if backend == "..."` branch in the pipeline, a
vendor attribute name on a span, an endpoint baked into config. None of that is
possible here:

* `internal/errtrack` has **zero vendor imports** and its `Event` carries
  Mesh-owned identifiers only.
* The selection switch (`internal/app.errtrackProviders`) knows a name and a
  factory. It does not know what Sentry is.
* The vendor adapter is a **leaf**: nothing imports it except that switch.
  Deleting the directory removes the vendor from the module and the neutral core
  still compiles.

**OTEL gets to be vendor-free by different means, not by a different standard.**
OTLP is a wire protocol every backend speaks, so "no vendor" costs nothing there —
an endpoint change is a config change. Error tracking has **no such protocol**:
grouping semantics, fingerprints, and stack-trace formats are per vendor. So
pluggability has to be a Go interface plus one adapter per vendor. That is exactly
the shape Mesh already uses where the vendor is unavoidable: `internal/connector/slack`
and `internal/model/openrouter`. Slack is named in a subpackage; Slack is not named
in `internal/event`.

The properties that make it acceptable, and that a second adapter must preserve:

* off by default; nothing runs unless `MESH_ERROR_TRACKER` names it
* selected by name at runtime, never by a build tag or a compile-time swap
* a leaf package, imported only by the provider table
* the vendor never sees an unredacted value
* switching or removing providers is an environment change plus a directory,
  never a change to the neutral core

## Adding a second provider

Rollbar, Bugsnag, GlitchTip, and Honeybadger all fit the same seam:

1. Add `internal/errtrack/<provider>/` implementing `errtrack.Reporter`
   (`CaptureError`, `Flush`, `Close`) and exporting `Name` and `Factory`.
   The vendor SDK may be imported **only** there.
2. Add one `case` to `errtrackProviders` in `internal/app/errtrack.go`.
3. Add a row to the table above.
4. Test against a stub transport. No adapter test may require network.

### 🚨 Registry versus switch, and why this is a switch

`errtrackProviders` is an explicit `switch`, not an `init()`-populated registry
that adapters add themselves to. That is a deliberate choice against the more
"extensible" option, for three reasons:

1. **It is the local convention.** `admin.Server` carries one verifier field per
   model provider and its handler switches on the request's provider name — "a
   second provider adds a second field here and a case in the handler's provider
   switch, not a registry". Telemetry's protocol selection is the same shape. Mesh
   is not registry-happy.
2. **A registry makes "unsupported provider" ambiguous.** With one,
   `MESH_ERROR_TRACKER=sentry` failing means *either* the name is wrong *or*
   someone dropped the blank import that pulls the adapter in. Two very different
   bugs behind one message, and the second is invisible in review because the
   missing line is the absence of a line.
3. **Import-time side effects make the selection untraceable.** With a switch,
   "which providers can this binary build" is answered by reading one function;
   with `init()`, by auditing every import in the module.

The registry's real advantage — adding a provider without touching a central
file — is worth less to us than a selection an operator can grep for. If that list
ever grows long enough that the switch is the annoying part, a registry becomes
the right call; the trigger should be real providers, not anticipated ones.

## Why install-level configuration, not per-agent

Error tracking is configured per **install** — the environment variables above,
or one `install_settings` row plus one install-owned secret. It is never in
`runtime_agents` or in the per-agent `secrets` rows.

That is not an exception to `internal/config`'s rule that agent-level
configuration never belongs in the process environment. It follows from it: an
error tracker reports on the **binary**. A panic in the admin mux belongs to no
agent at all, and neither does a failure to load the agent registry. There is no
agent whose row could own this credential.

Dashboard settings made it settable at runtime, which changes the **delivery**
mechanism and not the altitude: it is still one error tracker per binary. What it
stops requiring is that a process-shaped fact can only arrive as an environment
variable — a redeploy to turn on the feature whose entire value is being
already-configured when something breaks.

The mechanical argument is stronger still: per-agent routing would require the
panic path and the boot path to **resolve an agent before they may report**, which
is exactly the ordering that loses the report. The reporter has to exist before
the first thing capable of failing does, which is why `App.Run` installs it before
the database pool opens.

## What gets captured

| Capture point | Component tag | What it replaced |
| - | - | - |
| HTTP handler panic (whole mux) | `http` | net/http aborting the connection with a stderr line and no 500 |
| Turn execution failure | `slack.ingress` | `slog.Error` only, after the 200 was already sent |
| Reply delivery failure | `slack.delivery` | same, and indistinguishable from the above |
| Inbound persist failure | `slack.ingress` | `slog.Error` only — history loss, worse than a lost reply |
| Telemetry degraded at boot | `boot` (warning) | `log.Printf` |
| Agent-load failure at boot | `boot` (fatal) | a returned error and a restart loop |
| Dashboard errors via `POST /api/client-errors` | `dashboard` | nothing; browser errors were invisible |

Delivery failures are classified separately from execution failures
(`app.ErrDeliveryFailed`) because they are different incidents with different
fixes — a revoked bot token is not a model outage — and grouping them would
average two unrelated stories into one useless one.

**No call site branches on whether error tracking is enabled.** With it off,
`errtrack.Provider.Reporter()` returns a no-op reporter, so instrumentation is
inert rather than conditional — the same trick `telemetry.Tracer()` plays with
OpenTelemetry's global no-op tracer. A capture site that grows an
`if errorTrackingEnabled` branch is a bug: it will be wrong in exactly the
deployment that has reporting on.

Successful turns, dropped bot echoes, Slack retries, and observed messages report
nothing. An error tracker that fires on success is an error tracker an operator
mutes, which costs exactly the failures it exists for.

## The browser gets no vendor SDK

The dashboard's `ErrorBoundary` posts to `POST /api/client-errors` and Mesh
forwards the report through the same `Reporter` a server panic goes to. No
JavaScript error SDK is bundled. Why, given that shipping one is the more common
industry choice:

* **Bundle cost.** A hundred-odd kilobytes on the critical path of an operator
  dashboard, for a code path that ideally never runs.
* **It would put the vendor in the frontend too.** A hardcoded browser SDK is a
  second vendor coupling that no `MESH_ERROR_TRACKER` value can switch off, and it
  would make "change providers" a frontend rebuild.
* **A browser DSN is a public credential.** Shipping one means a token in a public
  bundle that anyone can post arbitrary events to. Forwarding through an
  authenticated endpoint means only a logged-in operator can create an issue in
  the operator's project.
* **Correlation.** A dashboard error and the API failure that caused it land in
  one project, with one release and one environment, instead of two that have to
  be joined by hand.

The cost, stated honestly: no source maps, no automatic breadcrumbs, no offline
queue, and an error thrown before the session cookie exists is not reportable. For
a single-page operator dashboard that is an acceptable trade — the component stack
the boundary already has is the part that identifies the bug.

The endpoint is authenticated (session cookie plus the API's Origin check — see
`internal/admin/session.go`), size-limited well below the API's default, and
rate-limited per session. It accepts exactly `{message, name, componentStack,
route, userAgent}`; `DisallowUnknownFields` makes anything else a 400 at the door,
because that field list **is** the list of things forwarded to a third party.

## 🚨 An error report leaves the process

Same trust boundary `docs/telemetry.md` draws for span attributes and
`docs/runtime/observability.md` draws when it says bodies live in Postgres only.
Reports carry **identifiers only**:

| Carried | Never carried |
| - | - |
| agent slug | Slack message bodies, thread text |
| conversation key | prompt and response bodies |
| connector type, event id, turn id | bot tokens, signing secrets, model keys |
| component, severity, route pattern | `MESH_AGENT_MASTER_KEY` |
| panic/component stacks | the error tracker's own DSN |

`errtrack.Event` has no `Body`, `Payload`, or `Detail` field, and that absence is
the design. But the shape of a struct is not enough, because the realistic leak is
not a new field — review catches that — it is a debugging line:

```go theme={null}
fmt.Errorf("model call failed for %q: %w", normalized.Body, err)   // don't
event.Tags["authorization"] = r.Header.Get("Authorization")        // don't
```

Both arrive at the boundary as a string, so the defence lives where strings become
reports (`internal/errtrack/redact.go`), in three layers:

1. **Exact-literal scrubbing** of credentials the process knows —
   `MESH_AGENT_MASTER_KEY` and the tracker's own DSN. Precise rather than
   heuristic, which is the only technique that works on a value with no
   recognizable shape.
2. **Credential-shape scrubbing** for `xox*`/`xapp-` Slack tokens, `sk-` provider
   keys, `Bearer` headers, URL userinfo, and `token=`/`secret:` pairs. Per-agent
   credentials are loaded from the database long after boot and can never be on
   the literal list, so shape is the honest tool. Deliberately aggressive: a
   mangled error message costs five minutes, a leaked bot token costs a workspace.
3. **Hard length caps** — 1 KiB per message, 256 bytes per tag, 8 KiB per stack.
   This is the backstop against an interpolated payload: a thread or a prompt does
   not fit, so it arrives truncated and the truncation marker is visible in the
   issue, which is how the mistake gets noticed and fixed.

The original error object never reaches the adapter, only a redacted wrapper
carrying a message and the innermost cause's type name. A vendor SDK walks error
chains and may serialize wrapped values; it is not allowed to be the component
deciding what leaves the process. The type name is preserved because without it
every issue in the project would be titled identically and redaction would have
cost exactly the grouping it was protecting.

`internal/errtrack/redact_test.go` enforces all of this, and it is written to fail
for a future change rather than to describe the current one. If it fails, the fix
is almost never to relax the assertion.

## Shutdown

Reports drain on shutdown, after in-flight turns finish, inside a small bounded
slice of the shutdown grace period — the same treatment telemetry gets and for the
same reason (the platform SIGKILLs well before the nominal grace
expires, so a wedged flush is a lost shutdown, not just lost reports).

Both optional signals **share one flush sub-budget** rather than getting one each.
Two budgets would let an operator who enables telemetry and error tracking spend
twice as long flushing, and that time comes out of reply drain — while the ceiling
that matters is the platform's, which does not care how many signals are
configured. A delivered reply matters more than a complete issue list. See
`shutdownBudgets` in `internal/app/app.go`.

A failed flush is logged, never reported through the reporter that just failed,
and never fatal. Error tracking is not the product.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.