> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Telemetry

# Telemetry

Mesh emits OpenTelemetry traces and logs. Both are opt-in, off by default, and
configured entirely through standard `OTEL_*` environment variables.

That last clause is the contract, not a convenience. No endpoint, vendor name,
API key, or backend-specific field appears anywhere in this repository, and none
ever should. Mesh is open source and the observability backend is a deployment
choice: changing where telemetry goes must be an environment change, never a code
change. A PR that adds a vendor's SDK, a vendor's attribute name, or a
`if backend == "..."` branch is the failure mode this document exists to prevent.

## Turning it on

A signal is enabled when an endpoint is named for it:

| Signal | Enabled by | Protocol |
| - | - | - |
| Traces | `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` or `OTEL_EXPORTER_OTLP_ENDPOINT` | `OTEL_EXPORTER_OTLP_TRACES_PROTOCOL` or `OTEL_EXPORTER_OTLP_PROTOCOL` |
| Logs | `OTEL_EXPORTER_OTLP_LOGS_ENDPOINT` or `OTEL_EXPORTER_OTLP_ENDPOINT` | `OTEL_EXPORTER_OTLP_LOGS_PROTOCOL` or `OTEL_EXPORTER_OTLP_PROTOCOL` |

Supported protocols are `grpc` (the default when unset) and `http/protobuf`. An
unrecognized value is rejected by name rather than silently defaulting, because
`http` is a plausible thing to type and it is not a value the exporter accepts.

Headers, TLS, compression, timeouts, and sampling are read from `OTEL_*` by the
SDK and exporters directly. Mesh does not re-expose them as its own
configuration, which would only create a second, staler surface.

Two consequences worth stating outright:

* **The general endpoint enables every signal.** Setting only
  `OTEL_EXPORTER_OTLP_ENDPOINT` turns on traces *and* logs. That is the
  specified meaning of the general variable.
* **Signals install independently.** A logs endpoint with no traces endpoint
  exports logs only, and a failure to construct one exporter never disables the
  other. A broken log exporter must not cost you the spans that would explain it.

## Failure is always advisory

`telemetry.Setup` returns a usable `*Provider` and an *advisory* error. There is
no input to it — no malformed endpoint, unreachable collector, unsupported
protocol, or typo'd `OTEL_RESOURCE_ATTRIBUTES` — that makes Mesh fail to boot.
The caller logs the complaint and keeps starting.

This is deliberate and it is the inverse of how the schema step behaves (see
`internal/app/app.go`). A process that cannot reach its collector still answers
Slack correctly; a process that cannot reach its schema fails every request. Only
the second is worth refusing to start over.

A malformed `OTEL_RESOURCE_ATTRIBUTES` degrades to "the attributes that parsed"
rather than disabling telemetry, because letting one env-var typo silently turn
off all observability is the opposite of what the operator who set it wanted.

## Logs

Mesh writes logs through `slog`. At boot, `internal/app.installLogging` installs
a JSON handler on stderr and — when log export is on — fans every record out to
both stderr and the OTLP pipeline.

**Exporting logs adds a destination; it never replaces one.** `kubectl logs` and
a developer's terminal are the diagnostics of last resort, and they are exactly
what you need when the reason nothing reaches the backend *is* the exporter. The
local handler therefore stays attached unconditionally.

The same boot step points the standard library's `log` package at `slog`, which
is what makes existing `log.Printf` call sites export without being rewritten.
The trade-off, so nobody has to rediscover it: a record that arrives via the
stdlib logger carries no `context`, so it has no trace id and cannot be
correlated to a span in the backend. Records from call sites that already use
`slog` with a context do. Converting the hot paths to context-aware logging is
follow-on work; having the logs exist at all comes first.

### 🚨 A log record leaves the process

Everything below is a third-party disclosure decision, not a debugging
convenience. Slack message bodies, bot tokens, the signing secret, and
`MESH_AGENT_MASTER_KEY` must never reach a log or span attribute.

Log the routing identifiers — conversation key, event id, agent slug — and use
them to find the record in Postgres, where it is supposed to live. A log line is
the easiest place in the codebase to accidentally interpolate an entire payload,
which is why this rule is written down rather than assumed.

The same rule governs error reports, which leave the process to a third party the
operator chose. [`error-tracking.md`](/error-tracking) states it for that
layer, along with why a *named vendor adapter* in its own leaf subpackage — off by
default, selected by name at runtime, imported by nothing but a provider switch —
is the `internal/connector/slack` and `internal/model/openrouter` pattern rather
than a violation of the no-vendor rule above. The rule's subject is the neutral
core: OTLP lets telemetry stay vendor-free at no cost because every backend
speaks it, and error tracking has no equivalent protocol.

One division of labour worth stating here so the two documents cannot drift: this
layer is a latency lens, error tracking is a grouping-and-alerting lens, and
neither substitutes for the other. Neither is the system of record — see
[`runtime/observability.md`](/runtime/observability).

## Traces

Spans carry Mesh-owned attributes for correlation only:

| Attribute | What it identifies |
| - | - |
| `mesh.conversation_key` | The conversation the turn belongs to |
| `mesh.connector_type` | Which connector delivered the event |
| `mesh.event_id` | The connector's event id, for dedupe correlation |
| `mesh.agent_slug` | Which agent served the span, by stable human handle |

`mesh.agent_slug` is what makes a trace legible once one process serves many
agents: without it every span looks like it came from the same bot, and "why did
this agent not reply" has no query behind it.

Instrumentation call sites do not branch on whether tracing is enabled. With it
off, `telemetry.Tracer()` is OpenTelemetry's global no-op tracer, so the
instrumentation is inert rather than conditional.

### Span inventory

Each connector emits three top-level spans: `<connector>.ingress` (the
request, until the ack), `<connector>.engagement` (the gate, the inbound
write, and Tier-2 classification, on the async side of the ack), and
`<connector>.turn` (the whole turn). The engagement and turn spans are their
own roots, linked to the ingress span (see `spanOptsLinkedTo`).

Inside the engagement and the turn, child spans split the time so latency can
be attributed rather than guessed at:

| Span | Parent | Where | Attributes |
| - | - | - | - |
| `mesh.engagement.facts` | engagement | `addressing.Gate` graph read | outcome only |
| `mesh.inbound.record` | engagement | `app.Runtime.RecordInbound` | outcome only |
| `mesh.inbound.resolve_name` | `mesh.inbound.record` | connector display-name lookup (only when one is made) | `mesh.actor.name_resolved` |
| `mesh.inbound.resolve_audience` | `mesh.inbound.record` (also under the turn, at the pre-run audience re-check) | connector audience lookup | `mesh.audience.kind` |
| `mesh.inbound.persist` | `mesh.inbound.record` | graph write of the inbound event | outcome only |
| `mesh.decision <task>` | engagement or turn | `decision.Client.Choose` / `ChooseEach` (native decisions: engagement, commitment assent, tool preselection, ...) | `mesh.decision.task`, `gen_ai.request.model`, `gen_ai.response.model`, `gen_ai.provider.name`, `gen_ai.usage.input_tokens`/`output_tokens`, `mesh.decision.outcome` (accepted/uncertain/error, with uncertain judged as the audit judges it), `mesh.decision.questions`/`accepted` (ChooseEach) |
| `chat <model>` | turn (or engagement) | every turn-engine model call (`turn.Engine.invokeModel`), every out-of-band text call through `decision.ObservedConnector` (commitment classifier, steering, memory advisor), and the compatibility-mode Tier-2 classifier | `gen_ai.operation.name=chat`, `gen_ai.request.model`, `gen_ai.request.max_tokens`, `gen_ai.response.model`, `gen_ai.provider.name`, `gen_ai.usage.input_tokens`/`output_tokens`, `gen_ai.response.finish_reasons`, `mesh.decision.task` (out-of-band calls), `mesh.model.retryable` (on failure) |
| `mesh.model.attempt` | `chat <model>` | one attempt of the retry layer (`model.RetryConnector`), per route | `mesh.model.route_index`, `mesh.model.attempt`, `gen_ai.request.model`, `gen_ai.response.model`, `mesh.model.retryable` (on failure) |
| `execute_tool <tool>` | turn | `tool.Executor.Execute`, for registered tools only | `gen_ai.operation.name=execute_tool`, `gen_ai.tool.name`, `mesh.tool.class` (query/workspace/browser/effect) |
| `mesh.workspace.acquire` | usually `execute_tool <tool>` | `workspace.Manager.Acquire` | `mesh.workspace.backend`, `mesh.workspace.reused`, `mesh.workspace.restore` (none/baseline/checkpoint) |
| `mesh.workspace.snapshot` | turn or `execute_tool <tool>` | `workspace.Manager.PublishBaseline` / `Checkpoint` | `mesh.workspace.snapshot_kind`, `mesh.workspace.backend` |
| `mesh.workspace.release` | turn | `workspace.Manager.ReleaseForRun` | `mesh.workspace.backend` |

Names and `gen_ai.*` keys follow the OpenTelemetry GenAI semantic conventions
where one exists (`chat <model>`, `execute_tool <tool>`); the rest are
Mesh-owned. `gen_ai.provider.name` is the provider descriptor's name the agent
resolved to, and `mesh.workspace.backend` is the backend kind label stored on
the workspace row: both are runtime data an operator chose, not a vendor the
code names.

A failed child span has error status and a bounded `error.type`
(`cancelled`, `deadline_exceeded`, or `_OTHER`), and **never the error's
text** (`telemetry.End`). The errors that reach these spans can wrap a
provider response body, a tool's upstream reply, or a sandbox's stderr; the
full error stays in the local log and the run ledger. An uncertain decision is
an outcome, not an error.

**Child spans take their tracer provider from the parent span in the context**
(`telemetry.StartChild`), not from the global. That is what lets the engine,
tool, workspace, and model packages instrument themselves without a tracer
parameter on every constructor, and it keeps tests isolated (a test's private
provider flows to every child through the context). The consequence is
deliberate: a call made with no span in its context -- tracing off, or a
background sweep outside any engagement or turn -- opens the no-op span and
costs one interface dispatch.

None of these spans carries message text, a prompt, a tool argument or
result, a display name, an audience's members, a repository or path, or a
credential. The tests assert that absence directly (`telemetrytest.Recording.Leak`).

## Shutdown

Both signals are drained on shutdown, after in-flight turns finish so their
spans and logs are included, inside a small bounded slice of the shutdown grace
period. The flush never starves reply drain: a delivered reply matters more than
a complete trace. See `shutdownBudgets` in `internal/app/app.go`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.