Skip to main content

Error tracking

Mesh reports unhandled failures to a pluggable error-tracking backend. It is opt-in, off by default, and selected by name at runtime. Sentry is the first and currently only adapter. The layer lives in internal/errtrack (provider-neutral, zero vendor imports) plus one leaf subpackage per vendor (internal/errtrack/sentry). Its boot, failure, and shutdown behaviour deliberately mirror internal/telemetry: opt-in, never fatal, a zero value that is a true no-op.

Turning it on

There are two ways, and which one wins is not a matter of ordering luck — see Precedence below.

From the settings page

/settings → Error tracking. Paste a DSN, pick a sample rate, tick Send reports, save. The change takes effect on the running process: no redeploy, no restart. The section also carries an Enable tracing toggle (with an optional Traces sample rate). It is a per-install opt-in to the provider’s OWN APM tracing, separate from error reports and off by default — the posture every install has had. OTEL stays Mesh’s primary latency lens (see OTEL stays environment-only); Sentry tracing is a supported opt-in for teams that already page on Sentry and want spans in the same place. It is a per-install setting rather than a hardcode or an infra decision because Mesh is vendor-neutral and open-source-first: whether a second tracing pipeline earns its cost is the operator’s call, not the binary’s. That is the whole reason this exists. Error tracking’s value is being already-configured when something breaks, and requiring a deploy to turn it on means the operator who just noticed a silent failure cannot have reports until after a release — by which time the incident is over and the evidence is gone. The page also offers:
  • Send a test report. One synthetic issue through the configuration being edited, so a DSN can be checked before it goes live. It writes nothing. With the field blank it tests whatever is currently in force, which is how an env-managed install answers “is my deploy manifest’s DSN actually reaching the project?”.
  • Clear stored DSN. Deletes the ciphertext (not a tombstone — see 0017) and turns reporting off, because a provider with no credential reports nowhere.
🚨 The DSN is write-only. No endpoint returns it, from either source. The page shows only whether one is stored and a redacted hint (https://…@host/project), and the input is always empty — leaving it blank keeps the stored value. A settings page that echoed the credential would put it in the DOM, in the browser’s autofill store, and in every screenshot.

From the environment

Tracing follows the section’s precedence like every other field: when the environment claims the section (see below), it owns the tracing toggle and its rate too, and the stored row is not consulted — precedence is per SECTION, so tracing rides the provider and DSN rather than mixing sources.

🚨 Precedence: the environment always wins

The rule runs that way round because a variable in a deploy manifest is an explicit deploy-time decision, and a click in a dashboard must not silently override it. A setting that can be overridden invisibly is worse than one that cannot be set at all. Two consequences worth stating plainly:
  • It applies per SECTION, not per field. With MESH_ERROR_TRACKER set, the environment owns the provider and the DSN and the sample rate — there is no mixing, because a half-environment/half-database configuration is one nobody can reason about during an incident.
  • The refusal is loud, and the UI agrees with it. The server answers a PATCH with 409, and the settings page disables the inputs and names the variables that own the section. A form that accepted an edit the resolver would then discard is precisely the dishonesty this design exists to prevent — which is why both halves exist: the server refusing is the guarantee, the disabled form is the honesty, and each is close to worthless without the other.
One resolver implements this (settings.ResolveErrorTracking) and the boot path, the save path, and the GET response all call it. Three copies of a precedence rule is three answers to “which one is live”.

Where each piece is stored

🚨 No credential may ever be written to install_settings.value. The split is the point of the table, not an implementation detail — db/migrations/0017_install_settings.sql argues it at length, and an integration test asserts the stored document contains no DSN. An install-owned secret uses the nil uuid as its owner_id, enforced by a CHECK so two writers cannot pick two sentinels and leave the install with two DSNs, one of which nothing reads.

What a save actually does

The reporter is swapped on the running process: build the replacement, publish it with one atomic store, then flush and close the reporter it replaced — in that order, so reports captured seconds before the save still reach the vendor. internal/app/errtrack_reload.go explains why an atomic pointer rather than a mutable field (a torn interface value is a crash, and it would be a crash inside the panic handler). If the rows commit but the new reporter cannot be built, the response is 200 with status: "saved_restart_required", not a 500:
  • The write already happened, and telling a client it failed invites a retry of something that is already done.
  • 🚨 But the value is stored and not in force — reporting is still using the previous configuration. The settings page says so prominently rather than showing “Saved”, because for the one feature whose job is telling an operator something is wrong, claiming reporting is on when it is not is the worst available answer.
A DSN that will not construct is rejected before it is persisted (422, with the adapter’s own explanation echoed), because failing in the form beats failing during a real incident.

OTEL stays environment-only

The settings page shows traces and logs read-only, and that is not an unfinished form. The OTEL SDK reads OTEL_* at exporter construction and installs process-global providers, so changing an endpoint means building new exporters and re-installing globals under live traffic — genuinely restart-shaped. Offering inputs would be lying about the size of the change. Error tracking can hot-swap because its reporter is one interface value behind an atomic pointer that Mesh owns end to end. That is a property of this layer, not a difference in ambition. The error tracker’s own tracing toggle is not the same thing as OTEL, and turning it on does not make Sentry Mesh’s trace backend. OTEL remains the primary latency lens; the Sentry-tracing opt-in exists so an install that already pages on Sentry can co-locate a sampled set of APM spans with its issues, and it defaults off precisely because a second tracing pipeline is a duplicate cost and a second attribute policy to police. Because the reporter is rebuilt on save, flipping the toggle takes effect on the running process — unlike the OTEL exporters above, which are restart-shaped. MESH_ENV supplies the environment tag, so an issue and a trace agree about which install they describe. The release tag comes from the binary’s embedded vcs.revision, which is what makes “this regressed in build X” work without the deployment passing a build argument. An unrecognized provider name is rejected by name rather than silently defaulting or silently disabling — the same rule as an unsupported OTLP protocol. A deployment that believes it is paging on panics while reporting nowhere is the worst outcome this feature has, and a typo is the likely cause. A named provider with no DSN is likewise reported at boot, not treated as a request to be disabled: naming the provider is an explicit opt-in.

Why this is not the trace backend

Reasonable question, since both are “send failures somewhere”. They answer different questions and neither substitutes for the other: The concrete failure that motivated this: ingress acks Slack in milliseconds, so a turn that fails afterwards produces no HTTP error anywhere. A user asked a question, the agent silently did not answer, and the only trace was one slog.Error line in a stream nobody alerts on. A trace backend cannot answer “is this new?”, and that is the question an operator is woken up by. docs/runtime/observability.md makes the same point from the other direction: the trace backend is a latency lens, not the system of record. Both stay wired. docs/runtime/observability.md’s rule is that no opt-in backend may be the only witness to anything, so every capture point below also logs.

🚨 Why a vendor adapter does not violate the no-vendor rule

docs/telemetry.md states, correctly, that no endpoint, vendor name, API key, or backend-specific field appears anywhere in this repository. internal/errtrack/sentry imports github.com/getsentry/sentry-go. That needs reconciling explicitly rather than by assumption. The rule’s subject is the neutral core. What that document is preventing is a vendor leaking into shared code: a if backend == "..." branch in the pipeline, a vendor attribute name on a span, an endpoint baked into config. None of that is possible here:
  • internal/errtrack has zero vendor imports and its Event carries Mesh-owned identifiers only.
  • The selection switch (internal/app.errtrackProviders) knows a name and a factory. It does not know what Sentry is.
  • The vendor adapter is a leaf: nothing imports it except that switch. Deleting the directory removes the vendor from the module and the neutral core still compiles.
OTEL gets to be vendor-free by different means, not by a different standard. OTLP is a wire protocol every backend speaks, so “no vendor” costs nothing there — an endpoint change is a config change. Error tracking has no such protocol: grouping semantics, fingerprints, and stack-trace formats are per vendor. So pluggability has to be a Go interface plus one adapter per vendor. That is exactly the shape Mesh already uses where the vendor is unavoidable: internal/connector/slack and internal/model/openrouter. Slack is named in a subpackage; Slack is not named in internal/event. The properties that make it acceptable, and that a second adapter must preserve:
  • off by default; nothing runs unless MESH_ERROR_TRACKER names it
  • selected by name at runtime, never by a build tag or a compile-time swap
  • a leaf package, imported only by the provider table
  • the vendor never sees an unredacted value
  • switching or removing providers is an environment change plus a directory, never a change to the neutral core

Adding a second provider

Rollbar, Bugsnag, GlitchTip, and Honeybadger all fit the same seam:
  1. Add internal/errtrack/<provider>/ implementing errtrack.Reporter (CaptureError, Flush, Close) and exporting Name and Factory. The vendor SDK may be imported only there.
  2. Add one case to errtrackProviders in internal/app/errtrack.go.
  3. Add a row to the table above.
  4. Test against a stub transport. No adapter test may require network.

🚨 Registry versus switch, and why this is a switch

errtrackProviders is an explicit switch, not an init()-populated registry that adapters add themselves to. That is a deliberate choice against the more “extensible” option, for three reasons:
  1. It is the local convention. admin.Server carries one verifier field per model provider and its handler switches on the request’s provider name — “a second provider adds a second field here and a case in the handler’s provider switch, not a registry”. Telemetry’s protocol selection is the same shape. Mesh is not registry-happy.
  2. A registry makes “unsupported provider” ambiguous. With one, MESH_ERROR_TRACKER=sentry failing means either the name is wrong or someone dropped the blank import that pulls the adapter in. Two very different bugs behind one message, and the second is invisible in review because the missing line is the absence of a line.
  3. Import-time side effects make the selection untraceable. With a switch, “which providers can this binary build” is answered by reading one function; with init(), by auditing every import in the module.
The registry’s real advantage — adding a provider without touching a central file — is worth less to us than a selection an operator can grep for. If that list ever grows long enough that the switch is the annoying part, a registry becomes the right call; the trigger should be real providers, not anticipated ones.

Why install-level configuration, not per-agent

Error tracking is configured per install — the environment variables above, or one install_settings row plus one install-owned secret. It is never in runtime_agents or in the per-agent secrets rows. That is not an exception to internal/config’s rule that agent-level configuration never belongs in the process environment. It follows from it: an error tracker reports on the binary. A panic in the admin mux belongs to no agent at all, and neither does a failure to load the agent registry. There is no agent whose row could own this credential. Dashboard settings made it settable at runtime, which changes the delivery mechanism and not the altitude: it is still one error tracker per binary. What it stops requiring is that a process-shaped fact can only arrive as an environment variable — a redeploy to turn on the feature whose entire value is being already-configured when something breaks. The mechanical argument is stronger still: per-agent routing would require the panic path and the boot path to resolve an agent before they may report, which is exactly the ordering that loses the report. The reporter has to exist before the first thing capable of failing does, which is why App.Run installs it before the database pool opens.

What gets captured

Delivery failures are classified separately from execution failures (app.ErrDeliveryFailed) because they are different incidents with different fixes — a revoked bot token is not a model outage — and grouping them would average two unrelated stories into one useless one. No call site branches on whether error tracking is enabled. With it off, errtrack.Provider.Reporter() returns a no-op reporter, so instrumentation is inert rather than conditional — the same trick telemetry.Tracer() plays with OpenTelemetry’s global no-op tracer. A capture site that grows an if errorTrackingEnabled branch is a bug: it will be wrong in exactly the deployment that has reporting on. Successful turns, dropped bot echoes, Slack retries, and observed messages report nothing. An error tracker that fires on success is an error tracker an operator mutes, which costs exactly the failures it exists for.

The browser gets no vendor SDK

The dashboard’s ErrorBoundary posts to POST /api/client-errors and Mesh forwards the report through the same Reporter a server panic goes to. No JavaScript error SDK is bundled. Why, given that shipping one is the more common industry choice:
  • Bundle cost. A hundred-odd kilobytes on the critical path of an operator dashboard, for a code path that ideally never runs.
  • It would put the vendor in the frontend too. A hardcoded browser SDK is a second vendor coupling that no MESH_ERROR_TRACKER value can switch off, and it would make “change providers” a frontend rebuild.
  • A browser DSN is a public credential. Shipping one means a token in a public bundle that anyone can post arbitrary events to. Forwarding through an authenticated endpoint means only a logged-in operator can create an issue in the operator’s project.
  • Correlation. A dashboard error and the API failure that caused it land in one project, with one release and one environment, instead of two that have to be joined by hand.
The cost, stated honestly: no source maps, no automatic breadcrumbs, no offline queue, and an error thrown before the session cookie exists is not reportable. For a single-page operator dashboard that is an acceptable trade — the component stack the boundary already has is the part that identifies the bug. The endpoint is authenticated (session cookie plus the API’s Origin check — see internal/admin/session.go), size-limited well below the API’s default, and rate-limited per session. It accepts exactly {message, name, componentStack, route, userAgent}; DisallowUnknownFields makes anything else a 400 at the door, because that field list is the list of things forwarded to a third party.

🚨 An error report leaves the process

Same trust boundary docs/telemetry.md draws for span attributes and docs/runtime/observability.md draws when it says bodies live in Postgres only. Reports carry identifiers only: errtrack.Event has no Body, Payload, or Detail field, and that absence is the design. But the shape of a struct is not enough, because the realistic leak is not a new field — review catches that — it is a debugging line:
Both arrive at the boundary as a string, so the defence lives where strings become reports (internal/errtrack/redact.go), in three layers:
  1. Exact-literal scrubbing of credentials the process knows — MESH_AGENT_MASTER_KEY and the tracker’s own DSN. Precise rather than heuristic, which is the only technique that works on a value with no recognizable shape.
  2. Credential-shape scrubbing for xox*/xapp- Slack tokens, sk- provider keys, Bearer headers, URL userinfo, and token=/secret: pairs. Per-agent credentials are loaded from the database long after boot and can never be on the literal list, so shape is the honest tool. Deliberately aggressive: a mangled error message costs five minutes, a leaked bot token costs a workspace.
  3. Hard length caps — 1 KiB per message, 256 bytes per tag, 8 KiB per stack. This is the backstop against an interpolated payload: a thread or a prompt does not fit, so it arrives truncated and the truncation marker is visible in the issue, which is how the mistake gets noticed and fixed.
The original error object never reaches the adapter, only a redacted wrapper carrying a message and the innermost cause’s type name. A vendor SDK walks error chains and may serialize wrapped values; it is not allowed to be the component deciding what leaves the process. The type name is preserved because without it every issue in the project would be titled identically and redaction would have cost exactly the grouping it was protecting. internal/errtrack/redact_test.go enforces all of this, and it is written to fail for a future change rather than to describe the current one. If it fails, the fix is almost never to relax the assertion.

Shutdown

Reports drain on shutdown, after in-flight turns finish, inside a small bounded slice of the shutdown grace period — the same treatment telemetry gets and for the same reason (the platform SIGKILLs well before the nominal grace expires, so a wedged flush is a lost shutdown, not just lost reports). Both optional signals share one flush sub-budget rather than getting one each. Two budgets would let an operator who enables telemetry and error tracking spend twice as long flushing, and that time comes out of reply drain — while the ceiling that matters is the platform’s, which does not care how many signals are configured. A delivered reply matters more than a complete issue list. See shutdownBudgets in internal/app/app.go. A failed flush is logged, never reported through the reporter that just failed, and never fatal. Error tracking is not the product.