> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Observability

# Observability

[`philosophy.md`](/philosophy) commits to "every agent turn is explainable"
and "traceable agent behavior." This document is where those commitments become
an operator-facing contract: what the runtime must capture about its own
behavior, how an operator reads it, and how it stays current without refresh.

The audience is the person asking "what did this agent just do, and should I
trust it?" That question must be answerable from the dashboard — not from a
debugger, not from `psql`, not from a third-party trace backend that may not be
configured.

## The one rule: capture before display

An audit trail that was not captured at write time is gone forever. A UI can
always be built over data that exists; no UI can be built over data that was
discarded in a stack frame. So this contract is asymmetric on purpose:

> **Capture is a runtime obligation that ships with the behavior it observes.
> Display can lag capture. Capture must never lag behavior.**

Concretely: no tool executes before its calls are persisted, no model role goes
live before its invocations are recorded, and no new execution capability lands
with a "we'll add the audit trail later" note. Later is the retrofit this
document exists to prevent. [`agent-harness.md`](/runtime/agent-harness) already
states the harness half of this ("every tool call and every observation is a
persisted event"); this document extends the same rule to model calls and makes
the read side a contract too.

## What one turn must leave behind

Today a turn leaves two `message_events` rows — the inbound message (with its
engagement verdict) and the outbound reply. Everything between them is
ephemeral. The engagement verdict already proves the pattern this document
generalizes: "why didn't you answer?" became durably answerable the day the
verdict started riding the persisted row. The same must become true of "why
*did* you answer, and what did it cost?"

When the turn ledger lands ([`agent-harness.md`](/runtime/agent-harness), roadmap
Phase 3), one completed turn must be reconstructable, in order, from durable
rows alone:

* **The turn itself.** That it ran, for which agent, in which conversation,
  triggered by which message event, with what status, error, and wall-clock
  duration. A failed turn must be distinguishable from an observed message
  without consulting logs or traces — today both leave one inbound row and
  silence, which is the single worst gap in the current schema.
* **Every model call.** Which role (`model_primary` / `model_fast`), which
  model actually served it, the request that was sent, the response that came
  back, token usage, latency, and error. Token usage is not optional
  telemetry: it is cost attribution, and per-agent billing questions are audit
  questions. (Today the OpenRouter response decoder does not even parse the
  `usage` field; capturing it is a decoder change, not just a write path.)
* **The prompt snapshot.** The assembled system prompt for each primary model
  call, stored once and referenced by digest — the
  [`../versioned-state.md`](/versioned-state) audit record of what this
  RuntimeAgent actually knew, proving which perspective supplied the call. **This
  one is wired.** The turn engine writes the snapshot *before* provider
  execution and fails the turn if the write fails; each durable model-call row
  references that snapshot. "What did it see?", what it returned, its cost, and
  how the run ended are now one traversable audit record.
* **Every tool call.** Name, the arguments exactly as the model supplied them,
  the validated form actually executed, the typed observation returned, status,
  and timing — per [`tools.md`](/runtime/tools). The distinction between
  supplied and validated arguments is deliberate: "what did the model ask for"
  and "what did the runtime do" are different audit questions, and
  normalization must not erase the first one.
* **Every memory write**, already durable via `memory_entries` and its
  revisions, made *reachable* from the conversation: `source_id` provenance
  exists but nothing indexes or traverses it. The timeline below needs
  "memories created during this conversation" to be one query, not ad-hoc SQL.
* **Progress, supersession, and budget outcomes.** The ack, `superseded_by`
  pointers, and budget-exhaustion reasons specified in
  [`agent-harness.md`](/runtime/agent-harness) and
  [`interrupt-model.md`](/runtime/interrupt-model) are all part of the same trail:
  why a run stopped is as auditable as why it started.

All of it keys back to the same spine: RuntimeAgent perspective → conversation
→ turn → steps. Linked, not siloed — a memory row that cannot find its turn,
or a tool call that cannot find its conversation, is capture without
explainability.

## Sensitivity: the database is the trust boundary

Prompt and response bodies contain conversation content. Storing them is the
right call — an audit trail of redacted stubs audits nothing — but it is a
data-sensitivity decision and it gets stated, not assumed:

* Bodies live in **Postgres only**, behind the same boundary as the
  conversation graph they quote. This is the same posture
  [`../persistence.md`](/persistence) already takes for message bodies.
* OTEL spans keep their existing policy (`internal/telemetry`): identifiers
  and verdicts only, never bodies, because spans leave the process. The trace
  backend is a latency lens, **not the system of record** — it is opt-in,
  retention-limited, and absent on most installs. Nothing in this contract may
  exist only as a span.
* Read access rides operator authentication (below). There is no anonymous or
  agent-context read path.
* Retention is a per-install knob on model-call bodies (`GET`/`PUT
  /api/settings/retention`, migration `0080`): terminal calls older than the
  window lose `request_body` and `response_body` (a null request body is what
  "pruned" means; the trace reports it as `bodies_pruned`); the turn skeleton — status, models, prompt snapshot
  reference, reply text, usage, timings — is retained. The default is keep;
  the knob exists so a privacy-sensitive install can shorten it without
  forking the schema. Prompt snapshots are the named next body class and are
  not yet bounded (they are referenced from `model_calls`, so theirs is a
  last-reference retention, not a row-age one).
* Control-plane views never feed model context. An operator can see that two
  agents are co-resident; an agent must not. This is
  [`../perspective.md`](/perspective)'s boundary restated for the read
  path: observability output is for humans, and quietly recycling it into
  context assembly would launder privileged correlations into cognition.

## The read contract: one API, two consumers

Everything captured above is served by the operator API — the JSON API under
`/api` that the dashboard already uses for agent configuration. The web UI is
a **client** of that API, never a privileged sibling with its own query path:
if the UI can render it, the API serves it, so the same facts are scriptable.

This is also a forward note: a **public, versioned operator API** is expected
later — for programmatic audit export, compliance tooling, and operators who
want their own dashboards. It is not designed here. What this contract
requires now is only that the read models be built as API resources first and
UI pages second, so that versioning them later is a naming exercise rather
than an excavation. Nothing about the dashboard may depend on a query the API
does not expose.

Read surfaces, by RuntimeAgent perspective:

* conversations: list, filterable by actor, connector, channel/shape, and
  recency — a DM is a conversation *with an actor*, a channel thread is a
  conversation *in a place*, and the filter vocabulary must respect that
* one conversation: the timeline (below)
* one turn: the full drill-down — trigger, verdict, prompt snapshot, model
  calls with bodies and usage, tool calls with observations, memory writes,
  outcome
* actors: per-actor timeline inside one perspective (roadmap "should have")
* shed turns: rows. `turn_disposals` (migration 0070) records every trigger the
  governor dropped after durable capture — reason (`queue_full`,
  `shutting_down`, `stale`, `leased_elsewhere`, `lease_unavailable`), the
  trigger's age, the staleness bound that age was judged against, the queue
  depth in force, and how many triggers had collapsed into it. Keyed to the
  message event, so a conversation timeline can show that a reply was owed and
  never came.
* delivery failures and collapsed turns: still log lines; they become queryable
  rows with everything else

## The conversation timeline

The centerpiece view, and the reason capture is keyed to the conversation
spine. One conversation renders as a single interleaved, chronological flow:

* messages, attributed, inbound and outbound
* each inbound message's engagement verdict — including `observe`, so silence
  is visible as a decision rather than an absence
* turn boundaries: where a turn started, what triggered it, how it ended
* **tool calls inline**, at the point in the flow where they happened, with
  arguments and observations expandable
* **model calls inline**, with role, model, token usage, latency, and the
  request/response bodies expandable
* memory writes inline, linking to the memory entry and its revisions
* acks, progress emissions, supersessions, and budget terminations
* when graph mode exists: the WorkGraph's node states surface through the
  same timeline, linking into the topology view (roadmap Phase 4)

The test for this view is the operator's trust question, asked at every row:
*what did the agent see, what did it do, what did it cost, and why?* Every row
answers from durable data, and every row links to its neighbors — message to
verdict to turn to calls to effects. A timeline that renders replies but hides
the tool call that produced them is a chat log, not an audit trail.

## Live updates

An operator watching a conversation must see it move — new messages, verdicts,
tool calls, and turn state — **without refreshing**. Polling per page is the
bolt-on version; the requirement is push.

The transport is an implementation choice (SSE is the likely first fit: the
dashboard needs server→client only, it rides ordinary HTTP and the existing
session auth, and reconnection is built in; a WebSocket clears the same bar).
The contract is transport-agnostic:

* **Authenticated.** The stream is an operator API surface and rides the same
  session authentication as every other `/api` route. No unauthenticated
  socket, ever — the stream carries conversation content.
* **Scoped.** A subscription names what it watches (a conversation, an agent,
  an install-wide firehose for a status page). Scope is enforced server-side
  with the same perspective discipline as the query API.
* **Notification, not source of truth.** Stream events carry identifiers and
  small deltas; the durable row remains canonical. A dropped connection is
  recovered by re-querying, so the stream can be lossy without the audit
  trail being lossy.
* **Fed at the chokepoints.** The graph writers and (once it lands) the turn
  ledger writers are the small set of places durable rows are born; they
  publish to an in-process broadcast as they commit. Multi-replica
  deployments upgrade the broadcast to Postgres `LISTEN/NOTIFY` or an
  equivalent — the subscriber contract does not change.

This is cheap now precisely because writes are already funneled through a few
seams. The expensive version of this feature is the one built after writes
have scattered.

## Ordering

1. **Spec holds the seat (this document).** The turn ledger, model-call and
   tool-call records, and the read/stream contracts are named before the
   harness hardens, so Phase 3 implements them as first-slice obligations
   rather than follow-ups.
2. **Cheap capture wins may land before Phase 3:** wire the existing
   prompt-snapshot writer into the turn engine; parse and store token usage;
   stamp the outbound event's `metadata` with model, duration, prompt digest,
   and the completion's `finish_reason` (so a `length`-truncated reply is
   detectable from the record rather than by eyeballing a reply that stops
   mid-word). Partial capture beats none, and none of these
   prejudge the ledger schema.
3. **The ledger lands with Phase 3**, capture-complete: turns, steps, model
   calls with usage, tool calls with observations.
4. **Read API and timeline land with Phase 4**, over data that already
   exists; live updates ride the same phase.
5. **The public versioned API comes later**, as a versioning pass over read
   models that were API-shaped from the start.

## Guardrails

* do not ship a behavior before the capture that explains it
* do not make an opt-in trace backend the only witness to anything
* do not put message or prompt bodies on spans or logs; bodies live behind
  the database trust boundary
* do not discard what the model actually said or asked for — validated forms
  are additional records, not replacements
* do not treat token usage as optional telemetry; it is cost attribution
* do not give the web UI a read path the API does not expose
* do not feed operator-facing views into agent context assembly
* do not let the stream become the source of truth; durable rows are canonical
* do not sever the spine: every captured record links to its turn and its
  conversation
* do not display what was never captured — a UI stitched from log scrapes is
  the retrofit, arriving late and lying about coverage


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.