> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Addressing

# Addressing

An agent in a channel is a participant, not a handler.

That distinction is the whole document. A handler runs on every event it can see.
A participant reads everything and speaks when spoken to, and the gap between
those two behaviors is the difference between a multiplayer runtime and a denial
of service pointed at your own hardware.

## The failure this exists to make impossible

An agent was a member of a busy shared Slack channel. Every message posted to
that channel — by anyone, to anyone, about anything — was treated as a request
directed at the agent. Each one opened its own session and started its own tool
loop. Ten ran concurrently against one channel, each one churning inference and
subprocesses, until the host ran out of CPU. The runtime died. Every *other*
agent co-tenanted on that host died with it.

Nothing about that outage required a bug. Each individual component did exactly
what it was written to do. The outage is the design, and it decomposes into three
independent defects that happen to compound:

1. **No addressing model.** Membership in a conversation was read as
   addressability. "I can see this message" became "this message is for me."
2. **No concurrency invariant.** Ten turns for one conversation were permitted to
   exist at once, because turn identity was derived from message arrival rather
   than from conversation identity.
3. **No fate isolation.** A turn's resource consumption could take down the
   runtime that supervised it, and therefore every other tenant of that runtime.

Each of the three is individually sufficient to cause an incident. This document
fixes the first, states the invariant for the second, and points at the doc that
owns the third. Fixing only one of them and declaring victory is the trap.

## Why not `require_mention`

There is an obvious fix: require an explicit `@` mention for every message in
every channel. We are not doing that as the general answer, and the reason is not
aesthetic.

Requiring `@`-mention markup on every message means the agent cannot participate
in a conversation, only field tickets in it. Ask a colleague "can you look at
this?" one message after saying their name, and they answer. Say "hey Lyra, what
do you think — " and then, in the next message, "specifically about the retry
budget," and a human handles it without being re-addressed. An agent that
requires re-mention on every message is not a participant. It is a form.

`require_mention` is a legitimate **per-conversation override** — a big noisy
`#general` may well want it, and an operator should be able to set it. It is the
wrong **default semantic**, because it papers over the actual missing concept:
the agent has no idea how many actors are in the room.

## Conversation shape is a first-class property

A single-user harness never needs this, which is exactly why every single-user
harness gets it wrong when pointed at a channel. Mesh already models
conversations as rows with participants (see [`../data-model.md`](/data-model)), so it can
know something OpenClaw structurally cannot: **how many actors are in here, and
am I one of two or one of fifteen?**

That gives a derived property on every conversation:

| Shape | Definition | Default engagement |
| - | - | - |
| `direct` | exactly two actors, one of which is this agent | everything is addressed to me |
| `small_group` | 3–N actors (default N=4) | addressed by default, mention-aware |
| `channel` | more than N actors, or a connector-typed channel | **addressed only when the evidence says so** |

Shape is computed from the participant set, not configured. A DM that a human
adds someone to stops being `direct`; a channel that empties out to two actors
becomes conversational. The agent's behavior changes with the room, because the
room is what changed.

This is the differentiator worth being loud about: engagement policy is a
function of conversation shape, and conversation shape is a first-class,
observable, queryable property of the graph rather than a per-integration
setting someone has to remember to configure.

## The engagement gate

Every inbound message resolves to exactly one of three dispositions:

```
  observe   — persisted to the graph, becomes context, does NOT trigger a turn
  engage    — persisted, becomes a trigger, enters the mailbox as a trigger
  ignore    — dropped pre-persistence (self-echo, retries, unsupported subtypes)
```

`observe` is the load-bearing one and the one that does not exist today. It is
what lets the agent stay contextually aware of a busy channel — "what have you
been discussing?" is answerable — while spending zero inference on it. Reading is
not replying. The graph is where ambient awareness lives; the mailbox is where
work lives.

### Signals, cheapest and most certain first

The gate is a ladder. It stops at the first tier that produces a verdict, and
that ordering is a cost-control mechanism, not a stylistic preference.

**Tier 0 — deterministic, no inference.**

| Signal | Verdict |
| - | - |
| Conversation shape is `direct` | engage |
| Message text contains this agent's own mention markup (`<@U…>`) | engage |
| Message is a threaded reply under a message this agent authored | engage |
| Message is in a thread where this agent has an open `awaiting_input` from this actor | engage |
| Author is this agent (self-echo), or a retry of a seen `event_id` | ignore |
| Author is a bot other than this agent | observe |

The mention check must resolve *this agent's own* user id from its own
credentials, not "any mention." A mention check that can never match fails open —
an agent that answers everything because it cannot find itself is the worst
possible failure mode, and it looks exactly like having no gate at all. Slack
renders mentions with the `U`-prefixed **user** id, not the `B`-prefixed bot id
that arrives as `bot_id` on echoed events; conflating them yields a check that
silently never fires.

The self-echo check is scoped the same way: *this agent's own* identity, never
"any bot." A blanket bot drop makes every agent blind to every other automated
participant — a co-resident agent's replies, a CI bot's failure reports — and
none of it ever becomes context, which forfeits the "what have you been
discussing?" property this gate exists to protect. Non-self bot authors land on
`observe`, not `engage`, and that verdict outranks every engage signal above:
the signals were designed for human authors, and letting a bot trip
"agent has replied in this thread" would have two co-resident agents answering
each other in an unbounded loop. Bot authorship is connector-visible metadata
(everyone in the room can see the bot badge), so consuming it here does not
breach the peer-recognition boundary — the gate learns *that* the author is
automated, never *which RuntimeAgent* it is. Agent-to-agent addressing arrives
with its own damping in a later phase; until then, agents read each other but
do not trigger each other. An install that does not know its own user id falls
back to the blanket drop, because an unscoped echo check that cannot recognize
itself must fail toward silence rather than toward a self-reply loop.

**Tier 1 — deterministic candidate detection, still no inference.**

A name-string match promotes the message to *candidate*, not to *addressed*.
Matching covers the agent's display name, slug, and configured aliases,
case-insensitively, at word boundaries, tolerating possessives and light typos.

Candidate is not a verdict, because "Lyra" appears in "I already asked Lyra
about that yesterday" — third-person reference, not address. Treating a name
mention as an address is a smaller version of the same mistake: it engages on
being *discussed* rather than on being *spoken to*.

**Tier 2 — bounded classification, one cheap inference call.**

Only candidates reach this tier. The question — "is this message addressed to
me, does it continue a request I'm already handling, or is it about me but not
for me?" — uses the [decision model (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/decision-models.md), with `model_fast`
compatibility for providers without native decisions or explicit rollback,
with up to the same 30 recent messages loaded for the responding agent, in
chronological order with human/agent attribution. The current message is
excluded from history and rendered separately, exactly once. The classifier
uses this context to distinguish ongoing requests from discussion between other
participants or a handoff away from the agent.

Input remains bounded: each body keeps at most 2,000 runes plus a truncation
marker, retaining both its opening and ending. The complete conversation input
is capped at 32,000 runes by dropping the oldest history first. The answer remains
a fixed-choice decision with probabilities (or a single boolean JSON object with
a 32-token ceiling in fast-model compatibility mode), and no tools. Only
conversation history already captured in this agent's graph and permitted for
the current audience is available; this does not fetch missing Slack history.
A failed history read records a degraded classification and observes the message
instead of making a confident decision without context. Database-less operation
still supports current-message-only classification.

The rule that makes this safe:

> **Deciding whether to answer must be orders of magnitude cheaper than
> answering.**

OpenClaw's real defect is that the decision *was* the answer: engaging cost a
full agent loop with tools, so an ambiguous channel message cost the same as a
real request. A classification call uses a bounded conversation input, a tiny answer, and no
tools. If the classifier ever needs a tool call, the design is wrong.

The classifier is rate-limited per conversation. Under a burst, exceeding the
budget degrades to Tier 1's verdict, which is `observe`. A chatty channel makes
the agent *quieter*, never busier. That property is the entire point.

**Default: `observe`.**

Absence of evidence is not addressed. In a channel, silence is the correct
failure mode: it is recoverable in one message ("hey, I meant you") and visibly
attributable, whereas over-engagement is neither.

### Every decision is recorded

The gate writes its verdict, the tier that produced it, and the matched signal
onto the message event. This is not telemetry for its own sake — it is the only
way to answer the two questions that will otherwise be unanswerable:

* "Why didn't you respond to me?" → because Tier 1 matched and Tier 2 classified
  it as third-person reference, at 14:31:02.
* "Why did you respond to that?" → because it was a threaded reply under your own
  message.

[`../philosophy.md`](/philosophy) says a turn we cannot explain means the system is wrong. A
turn that *did not happen* needs the same standard, or the gate becomes a black
box that people work around by @-mentioning everything — which lands us back at
`require_mention` by social convention instead of by config.

### The escape hatch

A gate this conservative will occasionally be wrong, so the recovery path is part
of the design and not a support burden. An explicit mention always engages
(Tier 0), so the human's natural repair — "@agent, this one" — always works. An
actor who addresses the agent immediately after being missed gets the preceding
unaddressed messages folded in as trigger context, so they do not have to repeat
themselves.

## The concurrency invariant

The addressing gate would have prevented the incident in the screenshot. It is
still not sufficient, because ten legitimately-addressed messages in a busy
channel would reproduce it exactly.

So, stated as an invariant rather than a knob:

> **For a given (agent, conversation) pair there is at most one turn in flight.
> Always. Not configurable.**

A second addressed message does not open a second turn. It enters the mailbox of
the existing one and is resolved by the interrupt policy — coalesce, preempt,
amend, or enqueue (see [`interrupt-model.md`](/runtime/interrupt-model)). That document assumed
this invariant; this one names it, because the failure we are designing against
is precisely its absence.

Turn identity derives from **conversation identity**, never from message
arrival. Ten rows in a session list for one channel is not a busy agent, it is a
broken key.

Above that, two ceilings that must exist before anything ships to a real
workspace:

* **Per-agent concurrency ceiling.** Max simultaneous turns across all of one
  agent's conversations. Excess turns queue; the queue has a depth; a full queue
  sheds with a visible state, not silently.
* **Per-install ceiling.** Max simultaneous turns across all agents in the
  process, because the incident's real damage was cross-tenant: one agent's fan-
  out killed every other agent on the host.

Unbounded goroutine spawn per inbound event is the mechanism by which all of this
goes wrong, and it is the current shape of the async turn path. Admission control
belongs at the point of spawn, and rejection has to be an observable state
("queued, depth 7") rather than an unlogged absence of a reply.

### Recovery must not re-fan

The nastiest version of this failure is a restart loop. If a backlog of
unaccounted messages exists when a worker comes up — because it crashed, or was
deployed, or the connector redelivered — then a naive recovery fans out one turn
per backlogged message and reproduces the outage on every boot, forever, with the
crash as its own trigger.

Two rules:

* **Staleness bound.** A trigger older than a threshold (order of minutes) is
  demoted to context. Answering a question the room resolved twenty minutes ago
  is noise even when the question was genuinely addressed to us. This is the
  answer to "what if significant time has passed": time does not license a
  *separate* turn, it *retires* the trigger.
* **Backlog collapse.** On recovery, all outstanding triggers for one
  conversation collapse into a single turn with the full set as context. Per
  conversation: one turn. Recovery obeys the same invariant as steady state; it
  is not a special case with its own concurrency model.

## Fate isolation is the third defect

The gate fixes over-engagement. The invariant fixes fan-out. Neither prevents one
legitimately busy turn from starving the runtime that supervises it, and in the
incident, the runtime dying is what escalated "an agent is being silly" into
"every agent on this host is down."

That is owned by [`sandboxed-execution.md`](/runtime/sandboxed-execution) and by the resource
accounting in it: turn workloads run in an execution backend with its own
resource envelope, and the supervisor's own budget is not drawn from the same
pool as the work it supervises. The requirement to state here is the boundary
itself: **a turn must not be able to kill the process that scheduled it.** An
addressing gate that is 99% accurate still leaves 1% of a very busy channel, and
that 1% must not be able to take the install down.

## Actor materialization (deliberately not V0)

Dropping an agent into a 10,000-member `#general` should not create 10,000 actor
rows. Flagging it now not because it needs solving now, but because the schema
decisions being made *today* determine whether it is solvable later without a
painful migration.

The distinction that keeps it open:

* **Connector identity** — "Slack user `U123` in workspace `T456`" — is cheap,
  local to the connector, and is what a message event should reference. Created
  on first sighting, or not at all.
* **Actor** — a stable cross-platform participant with a graph presence and
  history — is expensive and meaningful. It should be **promoted** on first real
  interaction, not on first sighting.

Which means the schema constraint to respect now is: a `MessageEvent` must be
persistable with a connector identity and **no** actor row, and actor promotion
must be a later linkage rather than a backfill of every row. Get that wrong and
the V1 fix is a migration over the largest table in the system.

What is genuinely undecided: what counts as "interaction" for promotion
(addressed us / we addressed them / same-thread participation), whether
promotion is reversible, and whether channel membership is materialized at all or
only ever sampled. Those are V1 questions. The foreign-key nullability is a today
question.

## What this document does not decide

* The exact small-group threshold. Default N=4 is a guess with a knob attached,
  and real channels should move it.
* Whether the Tier 2 classifier verdict should be cached per (actor, thread) to
  absorb multi-message thoughts, or recomputed per message. Caching is probably
  right and is an optimization, not a semantic.
* Per-connector addressing signals beyond Slack. GitHub review requests, email
  `To:` vs `Cc:` — the tier ladder is meant to be connector-agnostic with
  connector-supplied signals, and that boundary needs one more connector to be
  proven rather than asserted.
* Whether `observe`-class messages should be summarized into conversation state
  on a schedule rather than replayed as raw history. A busy channel produces a
  lot of context nobody has budgeted tokens for.

## Status

Steps 1–5 below are implemented. Tier 2 now uses the independent decision role
when available, retaining the fast-model compatibility path.

Implementation order, cheapest-and-most-protective first:

1. **Implemented.** Conversation shape as a derived property, and the Tier 0
   gate (`internal/addressing`). Deterministic, no inference. Shape derives from
   the connector's room classification (Slack `channel_type`) plus the
   participant count from `conversation_participants`; it is computed per
   message, never stored, so it cannot go stale. One deliberate widening of the
   table above: "threaded reply under a message this agent authored" is
   implemented as "this agent has an outbound message in this conversation",
   resolved through the agent's own connector identity. Since perspective
   ownership landed, the conversation being read already belongs to this agent, so
   a co-resident agent cannot donate participation regardless; the self-identity
   resolution stays because "I replied" is the fact actually being asked for, and
   it is the one privileged self-recognition
   [`../perspective.md`](/perspective) allows. A
   Mesh conversation *is* a thread, so this subsumes the root-author case and
   keeps the agent responsive to un-re-mentioned follow-ups in threads it
   joined mid-way. Without it, answering a mention would deafen the agent to
   the very thread it just spoke in, which is `require_mention` by another
   name.
2. **Implemented.** The per-(agent, conversation) single-flight invariant plus
   per-agent and per-install ceilings with admission control at the spawn point
   (`internal/app`'s turn governor). A second trigger collapses into the
   in-flight conversation's pending slot (newest wins; the served turn reads
   the full history, so collapsed messages are context, not casualties).
   Ceilings default to 4 turns per agent and 16 per install, tunable via
   `MESH_MAX_TURNS_PER_AGENT` / `MESH_MAX_TURNS`; the per-agent queue is
   bounded and sheds loudly when full. The staleness bound is anchored to when
   a trigger entered the governor — the queue is the thing whose age the
   runtime controls — and is the SERVING AGENT's admitted wall clock, floored at
   5 minutes. It was a flat 5-minute constant, which was a reliability defect
   once the wall clock became a per-agent dial (migration 0065): an agent tuned
   for long autonomous work holds its slots for tens of minutes, and every
   trigger queued behind it expired against a bound nobody raised. The floor is
   the load-bearing half — an agent may lengthen the window in which triggers
   behind it survive, never shorten it.
   Every drop is durable: see `turn_disposals` below.
3. **Implemented.** Decision recording on the message event
   (`message_events.engagement_disposition/tier/signal`), so silence is
   explainable before anyone has to ask. The verdict is written in the same
   transaction as the message and survives connector replays unrewritten.
   Its counterpart for the turn that never ran is `turn_disposals` (migration
   0070\): the governor's five terminal drop reasons, the trigger's age, the
   bound it was measured against, the queue depth, and the collapsed count,
   appended against the message event that went unanswered. `observe` silence
   and shed silence are now held to one standard — a dropped reply leaves a row,
   not a log line. What it deliberately does NOT do is retry: making a shed
   trigger deferrable re-opens the recovery-must-not-re-fan hazard named below,
   and that decision waits on being able to count disposals first.
4. **Implemented.** Tier 1 name-candidate detection.
5. **Implemented.** Tier 2 decision-model classification, with its budget and its degrade-to-observe
   path.

Steps 1 and 2 are the outage. Everything after them is quality. The
`require_mention` per-conversation override and backlog collapse across process
restarts (there is no durable trigger store yet — see
[`agent-harness.md`](/runtime/agent-harness)) are tracked separately.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.