Error tracking
Mesh reports unhandled failures to a pluggable error-tracking backend. It is opt-in, off by default, and selected by name at runtime. Sentry is the first and currently only adapter. The layer lives ininternal/errtrack (provider-neutral, zero vendor imports)
plus one leaf subpackage per vendor (internal/errtrack/sentry). Its boot,
failure, and shutdown behaviour deliberately mirror internal/telemetry: opt-in,
never fatal, a zero value that is a true no-op.
Turning it on
There are two ways, and which one wins is not a matter of ordering luck — see Precedence below.From the settings page
/settings → Error tracking. Paste a DSN, pick a sample rate, tick Send
reports, save. The change takes effect on the running process: no redeploy, no
restart.
The section also carries an Enable tracing toggle (with an optional Traces
sample rate). It is a per-install opt-in to the provider’s OWN APM tracing,
separate from error reports and off by default — the posture every install
has had. OTEL stays Mesh’s primary latency lens (see
OTEL stays environment-only); Sentry tracing is a
supported opt-in for teams that already page on Sentry and want spans in the same
place. It is a per-install setting rather than a hardcode or an infra decision
because Mesh is vendor-neutral and open-source-first: whether a second tracing
pipeline earns its cost is the operator’s call, not the binary’s.
That is the whole reason this exists. Error tracking’s value is being
already-configured when something breaks, and requiring a deploy to turn it on
means the operator who just noticed a silent failure cannot have reports until
after a release — by which time the incident is over and the evidence is gone.
The page also offers:
- Send a test report. One synthetic issue through the configuration being edited, so a DSN can be checked before it goes live. It writes nothing. With the field blank it tests whatever is currently in force, which is how an env-managed install answers “is my deploy manifest’s DSN actually reaching the project?”.
- Clear stored DSN. Deletes the ciphertext (not a tombstone — see 0017) and turns reporting off, because a provider with no credential reports nowhere.
https://…@host/project),
and the input is always empty — leaving it blank keeps the stored value. A
settings page that echoed the credential would put it in the DOM, in the browser’s
autofill store, and in every screenshot.
From the environment
🚨 Precedence: the environment always wins
The rule runs that way round because a variable in a deploy manifest is an
explicit deploy-time decision, and a click in a dashboard must not silently
override it. A setting that can be overridden invisibly is worse than one that
cannot be set at all.
Two consequences worth stating plainly:
- It applies per SECTION, not per field. With
MESH_ERROR_TRACKERset, the environment owns the provider and the DSN and the sample rate — there is no mixing, because a half-environment/half-database configuration is one nobody can reason about during an incident. - The refusal is loud, and the UI agrees with it. The server answers a
PATCHwith409, and the settings page disables the inputs and names the variables that own the section. A form that accepted an edit the resolver would then discard is precisely the dishonesty this design exists to prevent — which is why both halves exist: the server refusing is the guarantee, the disabled form is the honesty, and each is close to worthless without the other.
settings.ResolveErrorTracking) and the boot path,
the save path, and the GET response all call it. Three copies of a precedence
rule is three answers to “which one is live”.
Where each piece is stored
🚨 No credential may ever be written to
install_settings.value. The split is
the point of the table, not an implementation detail — db/migrations/0017_install_settings.sql
argues it at length, and an integration test asserts the stored document contains
no DSN.
An install-owned secret uses the nil uuid as its owner_id, enforced by a CHECK
so two writers cannot pick two sentinels and leave the install with two DSNs, one
of which nothing reads.
What a save actually does
The reporter is swapped on the running process: build the replacement, publish it with one atomic store, then flush and close the reporter it replaced — in that order, so reports captured seconds before the save still reach the vendor.internal/app/errtrack_reload.go explains why an atomic pointer rather than a
mutable field (a torn interface value is a crash, and it would be a crash inside
the panic handler).
If the rows commit but the new reporter cannot be built, the response is 200
with status: "saved_restart_required", not a 500:
- The write already happened, and telling a client it failed invites a retry of something that is already done.
- 🚨 But the value is stored and not in force — reporting is still using the previous configuration. The settings page says so prominently rather than showing “Saved”, because for the one feature whose job is telling an operator something is wrong, claiming reporting is on when it is not is the worst available answer.
422, with the
adapter’s own explanation echoed), because failing in the form beats failing during
a real incident.
OTEL stays environment-only
The settings page shows traces and logs read-only, and that is not an unfinished form. The OTEL SDK readsOTEL_* at exporter construction and installs
process-global providers, so changing an endpoint means building new exporters and
re-installing globals under live traffic — genuinely restart-shaped. Offering
inputs would be lying about the size of the change.
Error tracking can hot-swap because its reporter is one interface value behind an
atomic pointer that Mesh owns end to end. That is a property of this layer, not a
difference in ambition.
The error tracker’s own tracing toggle is not the same thing as OTEL, and
turning it on does not make Sentry Mesh’s trace backend. OTEL remains the primary
latency lens; the Sentry-tracing opt-in exists so an install that already pages on
Sentry can co-locate a sampled set of APM spans with its issues, and it defaults
off precisely because a second tracing pipeline is a duplicate cost and a second
attribute policy to police. Because the reporter is rebuilt on save, flipping the
toggle takes effect on the running process — unlike the OTEL exporters above, which
are restart-shaped.
MESH_ENV supplies the environment tag, so an issue and a trace agree about
which install they describe. The release tag comes from the binary’s embedded
vcs.revision, which is what makes “this regressed in build X” work without the
deployment passing a build argument.
An unrecognized provider name is rejected by name rather than silently
defaulting or silently disabling — the same rule as an unsupported OTLP protocol.
A deployment that believes it is paging on panics while reporting nowhere is the
worst outcome this feature has, and a typo is the likely cause.
A named provider with no DSN is likewise reported at boot, not treated as a
request to be disabled: naming the provider is an explicit opt-in.
Why this is not the trace backend
Reasonable question, since both are “send failures somewhere”. They answer different questions and neither substitutes for the other:
The concrete failure that motivated this: ingress acks Slack in milliseconds, so
a turn that fails afterwards produces no HTTP error anywhere. A user asked a
question, the agent silently did not answer, and the only trace was one
slog.Error line in a stream nobody alerts on. A trace backend cannot answer “is
this new?”, and that is the question an operator is woken up by.
docs/runtime/observability.md makes the same point from the other direction: the
trace backend is a latency lens, not the system of record.
Both stay wired. docs/runtime/observability.md’s rule is that no opt-in backend
may be the only witness to anything, so every capture point below also logs.
🚨 Why a vendor adapter does not violate the no-vendor rule
docs/telemetry.md states, correctly, that no endpoint, vendor name, API key, or
backend-specific field appears anywhere in this repository. internal/errtrack/sentry
imports github.com/getsentry/sentry-go. That needs reconciling explicitly
rather than by assumption.
The rule’s subject is the neutral core. What that document is preventing is a
vendor leaking into shared code: a if backend == "..." branch in the pipeline, a
vendor attribute name on a span, an endpoint baked into config. None of that is
possible here:
internal/errtrackhas zero vendor imports and itsEventcarries Mesh-owned identifiers only.- The selection switch (
internal/app.errtrackProviders) knows a name and a factory. It does not know what Sentry is. - The vendor adapter is a leaf: nothing imports it except that switch. Deleting the directory removes the vendor from the module and the neutral core still compiles.
internal/connector/slack
and internal/model/openrouter. Slack is named in a subpackage; Slack is not named
in internal/event.
The properties that make it acceptable, and that a second adapter must preserve:
- off by default; nothing runs unless
MESH_ERROR_TRACKERnames it - selected by name at runtime, never by a build tag or a compile-time swap
- a leaf package, imported only by the provider table
- the vendor never sees an unredacted value
- switching or removing providers is an environment change plus a directory, never a change to the neutral core
Adding a second provider
Rollbar, Bugsnag, GlitchTip, and Honeybadger all fit the same seam:- Add
internal/errtrack/<provider>/implementingerrtrack.Reporter(CaptureError,Flush,Close) and exportingNameandFactory. The vendor SDK may be imported only there. - Add one
casetoerrtrackProvidersininternal/app/errtrack.go. - Add a row to the table above.
- Test against a stub transport. No adapter test may require network.
🚨 Registry versus switch, and why this is a switch
errtrackProviders is an explicit switch, not an init()-populated registry
that adapters add themselves to. That is a deliberate choice against the more
“extensible” option, for three reasons:
- It is the local convention.
admin.Servercarries one verifier field per model provider and its handler switches on the request’s provider name — “a second provider adds a second field here and a case in the handler’s provider switch, not a registry”. Telemetry’s protocol selection is the same shape. Mesh is not registry-happy. - A registry makes “unsupported provider” ambiguous. With one,
MESH_ERROR_TRACKER=sentryfailing means either the name is wrong or someone dropped the blank import that pulls the adapter in. Two very different bugs behind one message, and the second is invisible in review because the missing line is the absence of a line. - Import-time side effects make the selection untraceable. With a switch,
“which providers can this binary build” is answered by reading one function;
with
init(), by auditing every import in the module.
Why install-level configuration, not per-agent
Error tracking is configured per install — the environment variables above, or oneinstall_settings row plus one install-owned secret. It is never in
runtime_agents or in the per-agent secrets rows.
That is not an exception to internal/config’s rule that agent-level
configuration never belongs in the process environment. It follows from it: an
error tracker reports on the binary. A panic in the admin mux belongs to no
agent at all, and neither does a failure to load the agent registry. There is no
agent whose row could own this credential.
Dashboard settings made it settable at runtime, which changes the delivery
mechanism and not the altitude: it is still one error tracker per binary. What it
stops requiring is that a process-shaped fact can only arrive as an environment
variable — a redeploy to turn on the feature whose entire value is being
already-configured when something breaks.
The mechanical argument is stronger still: per-agent routing would require the
panic path and the boot path to resolve an agent before they may report, which
is exactly the ordering that loses the report. The reporter has to exist before
the first thing capable of failing does, which is why App.Run installs it before
the database pool opens.
What gets captured
Delivery failures are classified separately from execution failures
(
app.ErrDeliveryFailed) because they are different incidents with different
fixes — a revoked bot token is not a model outage — and grouping them would
average two unrelated stories into one useless one.
No call site branches on whether error tracking is enabled. With it off,
errtrack.Provider.Reporter() returns a no-op reporter, so instrumentation is
inert rather than conditional — the same trick telemetry.Tracer() plays with
OpenTelemetry’s global no-op tracer. A capture site that grows an
if errorTrackingEnabled branch is a bug: it will be wrong in exactly the
deployment that has reporting on.
Successful turns, dropped bot echoes, Slack retries, and observed messages report
nothing. An error tracker that fires on success is an error tracker an operator
mutes, which costs exactly the failures it exists for.
The browser gets no vendor SDK
The dashboard’sErrorBoundary posts to POST /api/client-errors and Mesh
forwards the report through the same Reporter a server panic goes to. No
JavaScript error SDK is bundled. Why, given that shipping one is the more common
industry choice:
- Bundle cost. A hundred-odd kilobytes on the critical path of an operator dashboard, for a code path that ideally never runs.
- It would put the vendor in the frontend too. A hardcoded browser SDK is a
second vendor coupling that no
MESH_ERROR_TRACKERvalue can switch off, and it would make “change providers” a frontend rebuild. - A browser DSN is a public credential. Shipping one means a token in a public bundle that anyone can post arbitrary events to. Forwarding through an authenticated endpoint means only a logged-in operator can create an issue in the operator’s project.
- Correlation. A dashboard error and the API failure that caused it land in one project, with one release and one environment, instead of two that have to be joined by hand.
internal/admin/session.go), size-limited well below the API’s default, and
rate-limited per session. It accepts exactly {message, name, componentStack, route, userAgent}; DisallowUnknownFields makes anything else a 400 at the door,
because that field list is the list of things forwarded to a third party.
🚨 An error report leaves the process
Same trust boundarydocs/telemetry.md draws for span attributes and
docs/runtime/observability.md draws when it says bodies live in Postgres only.
Reports carry identifiers only:
errtrack.Event has no Body, Payload, or Detail field, and that absence is
the design. But the shape of a struct is not enough, because the realistic leak is
not a new field — review catches that — it is a debugging line:
internal/errtrack/redact.go), in three layers:
- Exact-literal scrubbing of credentials the process knows —
MESH_AGENT_MASTER_KEYand the tracker’s own DSN. Precise rather than heuristic, which is the only technique that works on a value with no recognizable shape. - Credential-shape scrubbing for
xox*/xapp-Slack tokens,sk-provider keys,Bearerheaders, URL userinfo, andtoken=/secret:pairs. Per-agent credentials are loaded from the database long after boot and can never be on the literal list, so shape is the honest tool. Deliberately aggressive: a mangled error message costs five minutes, a leaked bot token costs a workspace. - Hard length caps — 1 KiB per message, 256 bytes per tag, 8 KiB per stack. This is the backstop against an interpolated payload: a thread or a prompt does not fit, so it arrives truncated and the truncation marker is visible in the issue, which is how the mistake gets noticed and fixed.
internal/errtrack/redact_test.go enforces all of this, and it is written to fail
for a future change rather than to describe the current one. If it fails, the fix
is almost never to relax the assertion.
Shutdown
Reports drain on shutdown, after in-flight turns finish, inside a small bounded slice of the shutdown grace period — the same treatment telemetry gets and for the same reason (the platform SIGKILLs well before the nominal grace expires, so a wedged flush is a lost shutdown, not just lost reports). Both optional signals share one flush sub-budget rather than getting one each. Two budgets would let an operator who enables telemetry and error tracking spend twice as long flushing, and that time comes out of reply drain — while the ceiling that matters is the platform’s, which does not care how many signals are configured. A delivered reply matters more than a complete issue list. SeeshutdownBudgets in internal/app/app.go.
A failed flush is logged, never reported through the reporter that just failed,
and never fatal. Error tracking is not the product.