> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Model

# Model connector

Mesh separates *message connectors* from *model connectors*.

* message connectors handle Slack, Telegram, GitHub, and other human-facing surfaces
* model connectors handle inference providers like OpenRouter, Anthropic, and OpenAI

That split keeps the runtime honest:

* Slack does not know about model APIs
* OpenRouter does not know about threads or actors
* the turn router can swap one provider for another without rewiring the ingress layer
* WorkGraph planning does not depend on one provider's private response format

## Why OpenRouter first

OpenRouter is the best first model connector because it gives us one vendor-neutral entry point to many models.

That lets Mesh start with one integration while keeping the provider boundary clean enough to add more later.

## The second provider: Ramp Router

Ramp Router (router.com) is the second model connector, and it earned the slot
by being the proof the boundary works: commercially it is OpenRouter-shaped —
one key, one endpoint, many models, routing on the provider's side — but it is
**not wire-compatible**. Router speaks the OpenAI *Responses* API
(`POST /v1/responses`, an `input` item list, flattened function tools,
`function_call`/`function_call_output` items) and returns 404 on the Chat
Completions path. `internal/model/router` is therefore a real protocol
adapter behind the same provider-neutral contract, and nothing above
`internal/model` knows which protocol served a turn.

The differences that shaped the adapter, recorded so nobody rediscovers them:

* **Responses, not Chat Completions.** Requests translate `Messages` into
  input items; replies come back as an `output` item list whose text and
  `function_call` items are folded into the neutral `Response`. The
  Responses-status vocabulary (`completed`/`incomplete` + reason) is translated
  into the finish reasons the run ledger already records, because a ledger
  whose stop-reason vocabulary varies by provider cannot be queried.
* **Native tools only.** Router's catalog is a curated set of frontier models
  that all support function calling, and a genuinely unsupported capability is
  a crisp `501 not_implemented_error`. The openrouter connector's text-protocol
  fallback is deliberately not duplicated here; its triggering condition is
  designed out by the provider.
* **No key-metadata endpoint.** Key validation is an authenticated
  `GET /v1/models`; a valid key confirms with no label or spend because Router
  reports none, and Mesh does not invent either.
* **The catalog is a property of the key.** `GET /v1/models` answers relative
  to the caller's key, so the admin API serves Router's catalog per agent
  (decrypting that agent's stored key) where OpenRouter's is one public,
  install-wide proxy.
* **No auto-routing model id.** OpenRouter has `openrouter/auto`; Router has
  nothing equivalent, so the process-level default model stays an OpenRouter
  id and a Router agent is expected to have `model_primary` set (the wizard
  requires it). A Router agent whose models were cleared fails loudly at the
  provider, the same visible-misconfiguration stance as a missing key.

The provider is chosen per agent at runtime (`runtime_agents.model_provider`,
an explicit switch in `internal/app`), never per deployment: two agents in one
install can bill to different providers.

## First contract

The MVP model connector may start small. It should accept:

* model name
* prompt messages
* optional temperature
* optional token budget
* optional provider-neutral tool definitions

And return:

* content
* resolved model name
* raw response for audit/debugging
* requested tool calls and finish reason
* provider-reported token usage and request id

That is enough for the first Slack call-and-response slice and its bounded
`remember` round. It is not the complete harness contract: structured output,
cache-token detail, and retry classification remain.

## Structured output is a runtime primitive

The durable harness and WorkGraph planner need typed model output rather than prose that control-plane code has to parse heuristically.

The model connector should therefore grow a provider-neutral structured-output request before WorkGraph execution ships. Conceptually:

```go theme={null}
type Request struct {
    Model        string
    Messages     []Message
    MaxTokens    int
    Temperature  *float64
    OutputSchema json.RawMessage // optional JSON Schema or equivalent
    Tools        []ToolSpec
}

type Response struct {
    Content       string
    Structured    json.RawMessage
    ToolCalls     []ToolCall
    ResolvedModel string
    FinishReason  string
    Usage         Usage
    ProviderID    string
    Raw           json.RawMessage
}
```

Exact types can change. The invariant should not: provider-specific JSON does not become the WorkGraph plan format.

A graph planner returns a typed **proposal** matching Mesh's WorkGraph proposal schema. That proposal still passes deterministic admission in Mesh. Successful schema generation is not permission to execute.

The same structured-output primitive serves other bounded runtime questions:

* `model_fast` execution-shape classification (`loop` vs `graph_candidate`)
* addressed-or-not classification
* follow-up relation classification
* verifier verdicts
* complexity/ack triage

Typed model output fails closed to the caller's declared fallback. Invalid graph-planning output means "no admitted graph," not "best-effort parse and start workers."

## Tool calls and usage

The bounded memory loop now carries provider-neutral tool definitions, tool
requests/results, finish reason, provider request id, and prompt/completion/total
usage. The durable ledger accumulates that usage into the Run's token budget.
Cache-token detail and normalized retry classification remain connector work.

Mesh owns tool grants, validation, execution, persistence, and budgets. A model
provider does not decide whether an agent has memory: when an agent has a memory
store, every model can request `remember`.

OpenRouter metadata selects only the wire protocol. Models with native tool-call
support receive provider-native definitions and return native tool requests;
Mesh asks OpenRouter to route those requests only to endpoints that accept the
required parameters. Other models receive the same granted definitions in a
strict JSON decision protocol. Mesh parses that decision, assigns call ids,
validates it against the grant and schema, executes the tool itself, and sends
the result back as ordinary model context. A metadata lookup failure uses the
same validated fallback rather than removing tools or failing the turn.

The decision must be exactly one JSON object with no unknown fields; the only
tolerated wrapping is a single Markdown ` ```json ` fence and surrounding
whitespace. A reply that is not such an object (prose, a truncated object, JSON
`null`, or an object followed by more text) gets one repair: the connector re-asks once,
with the model's own reply as the assistant turn and a short Mesh correction
as the last message, and parses the answer with the same strict decoder. Prose
is never read as an answer or as a tool having run. The repair runs inside the
same Layer A attempt (its context, deadlines, and an ownership recheck), at
most once per prepared completion across Layer A's retries, and a failed
repair returns the error the unrepaired path always returned (transient for an
unparseable reply, terminal for trailing content). That holds when the repair
request itself fails, even terminally (a 400 context overflow from echoing a
long reply): the repair is optional and must not stop Layer A from
regenerating or falling back. Only the caller's context ending or an ownership
refusal propagates unchanged. A parsed but invalid
decision (unknown type, an empty or invalid tool call) is not repaired. For a
repaired call the ledger keeps the first request body; the response body is a
`mesh_fallback_repair` envelope with both responses' exact bytes, the
messages the repair appended, and whether the repair request omitted
`temperature` after an earlier rejection, so the repair request is
reconstructable exactly from the ledger; usage is the sum of both calls
(`internal/model/openrouter/fallback_repair.go`). The router connector and the
OpenAI connector are native-tools only and have no text protocol to repair.

The catalog also gates `temperature`, with two exceptions the catalog cannot
express. Anthropic models (`anthropic/…`) never receive it: Anthropic is
deprecating the parameter on newer models, and OpenRouter's Azure endpoint for
them already answers `400 "temperature is deprecated for this model"` while
Bedrock and Anthropic direct still accept it, so which endpoint OpenRouter
picked decided whether a turn failed. And for any model, a 400 whose
*provider* error message (unwrapped from OpenRouter's `error.metadata.raw`
envelope, never the echoed request) rejects `temperature` is not terminal: the
connector remembers that the model rejects the parameter and returns the
failure re-typed as retryable, so Layer A's bounded loop — with its ownership
recheck, backoff, attempt budget and recovery deadline — runs the next attempt,
which omits `temperature`. The connector itself never sends a second request.
Later turns skip the parameter up front for the life of the process. Every
other 400 stays terminal. The run ledger records the first attempt's body, as
it does for every Layer A retry; per-attempt records are a Layer A follow-up.

A failed provider call reports **labels, never the provider's message**. Every
HTTP connector (openai, openrouter, router) parses the non-2xx body — read to a
64 KiB bound — for the OpenAI-compatible `error.type` / `error.code` /
`error.param` (preferring the upstream's when OpenRouter wraps one in
`metadata.raw`), keeps each only if it has a label's shape, and renders them
with the status: `openai: status 400: invalid_request_error
(param=tools[3].function.parameters, code=invalid_function_parameters)`. That
string is what the turn-failure record carries, and the `model attempt failed`
log line carries the same fields as `status` / `error_type` / `error_code` /
`error_param` attributes. A 2xx whose body is still a failure (Router's
`status: "failed"`, an OpenAI-compatible `{"error":…}` on 200) follows the same
rule: the connector's reason plus labels, e.g. `router: response failed:
(code=server_error)`. The provider's `message` is never rendered, logged or
reported, because providers quote the request back in it (a prompt fragment, a
tool argument, an image URL); the body is retained only privately on the error
for a connector's own decision, like the temperature rejection above.
Classification is unchanged: it reads the status, never the labels.

The current loop permits up to eight calls per round across three bounded rounds.
Requests for tools that were not offered, invalid arguments, too many calls, and
a request beyond the final round are rejected without execution. Successful
protocol-capability lookups are cached for the serving client's lifetime; that
cache is an adapter optimization, not a runtime permission boundary.

At minimum expose:

* requested tool name and typed arguments
* stable provider/request identifiers where available
* finish/stop reason
* input/output/cache token usage where available
* retryable versus terminal provider failure classification

WorkGraph and Run budgets cannot be enforced honestly if token/tool usage exists only inside a provider's raw payload.

## Model roles per agent

An agent names a model per role, not one model:

* `model_primary` — reasoning, tool use, final replies, and initial WorkGraph planning
* `model_fast` — cheap, low-latency work like complexity triage, execution-shape
  candidacy, conversation titles, context compaction, and routing
* `model_heartbeat` — periodic self-directed activity

All three resolve through the same model connector contract. The split lives in
the agent's own row, so a connector never needs to know which role it is serving,
and two agents in the same install can be on entirely different models.

Each role is optional and falls back: `model_heartbeat` to `model_fast`,
`model_fast` to `model_primary`, and `model_primary` to the process default. That
is the one thing the environment contributes — a default model for an agent that
did not name one, not the agent's configuration itself.

Do not add a `model_planner` role merely because WorkGraphs exist. Initial graph planning is `model_primary` work. A dedicated role is justified later only if measurements show planning has a materially different capability/cost profile.

Provider selection and API keys are agent-owned runtime state. The setup flow
validates the key, stores it encrypted, and records the provider on the
RuntimeAgent, so agents in one deployment may use separate provider accounts.
`OPENROUTER_API_KEY` remains only as a legacy deployment-wide fallback for an
agent created before provider provisioning existed; a newly provisioned agent
does not inherit it, and a missing or unreadable agent-owned key must not fall
back silently. See [`../agent-provisioning.md`](/agent-provisioning).

See [`../runtime/agent-harness.md`](/runtime/agent-harness) for model-role usage and [`../runtime/work-graph.md`](/runtime/work-graph) for the graph proposal/admission boundary.

## Goal for the first slice

Given a normalized Slack event, the runtime should be able to ask OpenRouter for a completion and route the result back to the Slack connector.

The next connector milestone, before WorkGraph execution, is provider-neutral
structured output and retry classification. Tool calls are no longer trapped in
OpenRouter's private JSON: the current memory loop persists them as durable
RunSteps with observations. The scheduler is still synchronous and interrupted
runs are failed rather than resumed, so this does not yet pretend to be the
crash-resumable harness.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.