Telemetry
Mesh emits OpenTelemetry traces and logs. Both are opt-in, off by default, and configured entirely through standardOTEL_* environment variables.
That last clause is the contract, not a convenience. No endpoint, vendor name,
API key, or backend-specific field appears anywhere in this repository, and none
ever should. Mesh is open source and the observability backend is a deployment
choice: changing where telemetry goes must be an environment change, never a code
change. A PR that adds a vendor’s SDK, a vendor’s attribute name, or a
if backend == "..." branch is the failure mode this document exists to prevent.
Turning it on
A signal is enabled when an endpoint is named for it:
Supported protocols are
grpc (the default when unset) and http/protobuf. An
unrecognized value is rejected by name rather than silently defaulting, because
http is a plausible thing to type and it is not a value the exporter accepts.
Headers, TLS, compression, timeouts, and sampling are read from OTEL_* by the
SDK and exporters directly. Mesh does not re-expose them as its own
configuration, which would only create a second, staler surface.
Two consequences worth stating outright:
- The general endpoint enables every signal. Setting only
OTEL_EXPORTER_OTLP_ENDPOINTturns on traces and logs. That is the specified meaning of the general variable. - Signals install independently. A logs endpoint with no traces endpoint exports logs only, and a failure to construct one exporter never disables the other. A broken log exporter must not cost you the spans that would explain it.
Failure is always advisory
telemetry.Setup returns a usable *Provider and an advisory error. There is
no input to it — no malformed endpoint, unreachable collector, unsupported
protocol, or typo’d OTEL_RESOURCE_ATTRIBUTES — that makes Mesh fail to boot.
The caller logs the complaint and keeps starting.
This is deliberate and it is the inverse of how the schema step behaves (see
internal/app/app.go). A process that cannot reach its collector still answers
Slack correctly; a process that cannot reach its schema fails every request. Only
the second is worth refusing to start over.
A malformed OTEL_RESOURCE_ATTRIBUTES degrades to “the attributes that parsed”
rather than disabling telemetry, because letting one env-var typo silently turn
off all observability is the opposite of what the operator who set it wanted.
Logs
Mesh writes logs throughslog. At boot, internal/app.installLogging installs
a JSON handler on stderr and — when log export is on — fans every record out to
both stderr and the OTLP pipeline.
Exporting logs adds a destination; it never replaces one. kubectl logs and
a developer’s terminal are the diagnostics of last resort, and they are exactly
what you need when the reason nothing reaches the backend is the exporter. The
local handler therefore stays attached unconditionally.
The same boot step points the standard library’s log package at slog, which
is what makes existing log.Printf call sites export without being rewritten.
The trade-off, so nobody has to rediscover it: a record that arrives via the
stdlib logger carries no context, so it has no trace id and cannot be
correlated to a span in the backend. Records from call sites that already use
slog with a context do. Converting the hot paths to context-aware logging is
follow-on work; having the logs exist at all comes first.
🚨 A log record leaves the process
Everything below is a third-party disclosure decision, not a debugging convenience. Slack message bodies, bot tokens, the signing secret, andMESH_AGENT_MASTER_KEY must never reach a log or span attribute.
Log the routing identifiers — conversation key, event id, agent slug — and use
them to find the record in Postgres, where it is supposed to live. A log line is
the easiest place in the codebase to accidentally interpolate an entire payload,
which is why this rule is written down rather than assumed.
The same rule governs error reports, which leave the process to a third party the
operator chose. error-tracking.md states it for that
layer, along with why a named vendor adapter in its own leaf subpackage — off by
default, selected by name at runtime, imported by nothing but a provider switch —
is the internal/connector/slack and internal/model/openrouter pattern rather
than a violation of the no-vendor rule above. The rule’s subject is the neutral
core: OTLP lets telemetry stay vendor-free at no cost because every backend
speaks it, and error tracking has no equivalent protocol.
One division of labour worth stating here so the two documents cannot drift: this
layer is a latency lens, error tracking is a grouping-and-alerting lens, and
neither substitutes for the other. Neither is the system of record — see
runtime/observability.md.
Traces
Spans carry Mesh-owned attributes for correlation only:mesh.agent_slug is what makes a trace legible once one process serves many
agents: without it every span looks like it came from the same bot, and “why did
this agent not reply” has no query behind it.
Instrumentation call sites do not branch on whether tracing is enabled. With it
off, telemetry.Tracer() is OpenTelemetry’s global no-op tracer, so the
instrumentation is inert rather than conditional.
Span inventory
Each connector emits three top-level spans:<connector>.ingress (the
request, until the ack), <connector>.engagement (the gate, the inbound
write, and Tier-2 classification, on the async side of the ack), and
<connector>.turn (the whole turn). The engagement and turn spans are their
own roots, linked to the ingress span (see spanOptsLinkedTo).
Inside the engagement and the turn, child spans split the time so latency can
be attributed rather than guessed at:
Names and
gen_ai.* keys follow the OpenTelemetry GenAI semantic conventions
where one exists (chat <model>, execute_tool <tool>); the rest are
Mesh-owned. gen_ai.provider.name is the provider descriptor’s name the agent
resolved to, and mesh.workspace.backend is the backend kind label stored on
the workspace row: both are runtime data an operator chose, not a vendor the
code names.
A failed child span has error status and a bounded error.type
(cancelled, deadline_exceeded, or _OTHER), and never the error’s
text (telemetry.End). The errors that reach these spans can wrap a
provider response body, a tool’s upstream reply, or a sandbox’s stderr; the
full error stays in the local log and the run ledger. An uncertain decision is
an outcome, not an error.
Child spans take their tracer provider from the parent span in the context
(telemetry.StartChild), not from the global. That is what lets the engine,
tool, workspace, and model packages instrument themselves without a tracer
parameter on every constructor, and it keeps tests isolated (a test’s private
provider flows to every child through the context). The consequence is
deliberate: a call made with no span in its context — tracing off, or a
background sweep outside any engagement or turn — opens the no-op span and
costs one interface dispatch.
None of these spans carries message text, a prompt, a tool argument or
result, a display name, an audience’s members, a repository or path, or a
credential. The tests assert that absence directly (telemetrytest.Recording.Leak).
Shutdown
Both signals are drained on shutdown, after in-flight turns finish so their spans and logs are included, inside a small bounded slice of the shutdown grace period. The flush never starves reply drain: a delivered reply matters more than a complete trace. SeeshutdownBudgets in internal/app/app.go.