Skip to main content

Model connector

Mesh separates message connectors from model connectors.
  • message connectors handle Slack, Telegram, GitHub, and other human-facing surfaces
  • model connectors handle inference providers like OpenRouter, Anthropic, and OpenAI
That split keeps the runtime honest:
  • Slack does not know about model APIs
  • OpenRouter does not know about threads or actors
  • the turn router can swap one provider for another without rewiring the ingress layer
  • WorkGraph planning does not depend on one provider’s private response format

Why OpenRouter first

OpenRouter is the best first model connector because it gives us one vendor-neutral entry point to many models. That lets Mesh start with one integration while keeping the provider boundary clean enough to add more later.

The second provider: Ramp Router

Ramp Router (router.com) is the second model connector, and it earned the slot by being the proof the boundary works: commercially it is OpenRouter-shaped — one key, one endpoint, many models, routing on the provider’s side — but it is not wire-compatible. Router speaks the OpenAI Responses API (POST /v1/responses, an input item list, flattened function tools, function_call/function_call_output items) and returns 404 on the Chat Completions path. internal/model/router is therefore a real protocol adapter behind the same provider-neutral contract, and nothing above internal/model knows which protocol served a turn. The differences that shaped the adapter, recorded so nobody rediscovers them:
  • Responses, not Chat Completions. Requests translate Messages into input items; replies come back as an output item list whose text and function_call items are folded into the neutral Response. The Responses-status vocabulary (completed/incomplete + reason) is translated into the finish reasons the run ledger already records, because a ledger whose stop-reason vocabulary varies by provider cannot be queried.
  • Native tools only. Router’s catalog is a curated set of frontier models that all support function calling, and a genuinely unsupported capability is a crisp 501 not_implemented_error. The openrouter connector’s text-protocol fallback is deliberately not duplicated here; its triggering condition is designed out by the provider.
  • No key-metadata endpoint. Key validation is an authenticated GET /v1/models; a valid key confirms with no label or spend because Router reports none, and Mesh does not invent either.
  • The catalog is a property of the key. GET /v1/models answers relative to the caller’s key, so the admin API serves Router’s catalog per agent (decrypting that agent’s stored key) where OpenRouter’s is one public, install-wide proxy.
  • No auto-routing model id. OpenRouter has openrouter/auto; Router has nothing equivalent, so the process-level default model stays an OpenRouter id and a Router agent is expected to have model_primary set (the wizard requires it). A Router agent whose models were cleared fails loudly at the provider, the same visible-misconfiguration stance as a missing key.
The provider is chosen per agent at runtime (runtime_agents.model_provider, an explicit switch in internal/app), never per deployment: two agents in one install can bill to different providers.

First contract

The MVP model connector may start small. It should accept:
  • model name
  • prompt messages
  • optional temperature
  • optional token budget
  • optional provider-neutral tool definitions
And return:
  • content
  • resolved model name
  • raw response for audit/debugging
  • requested tool calls and finish reason
  • provider-reported token usage and request id
That is enough for the first Slack call-and-response slice and its bounded remember round. It is not the complete harness contract: structured output, cache-token detail, and retry classification remain.

Structured output is a runtime primitive

The durable harness and WorkGraph planner need typed model output rather than prose that control-plane code has to parse heuristically. The model connector should therefore grow a provider-neutral structured-output request before WorkGraph execution ships. Conceptually:
Exact types can change. The invariant should not: provider-specific JSON does not become the WorkGraph plan format. A graph planner returns a typed proposal matching Mesh’s WorkGraph proposal schema. That proposal still passes deterministic admission in Mesh. Successful schema generation is not permission to execute. The same structured-output primitive serves other bounded runtime questions:
  • model_fast execution-shape classification (loop vs graph_candidate)
  • addressed-or-not classification
  • follow-up relation classification
  • verifier verdicts
  • complexity/ack triage
Typed model output fails closed to the caller’s declared fallback. Invalid graph-planning output means “no admitted graph,” not “best-effort parse and start workers.”

Tool calls and usage

The bounded memory loop now carries provider-neutral tool definitions, tool requests/results, finish reason, provider request id, and prompt/completion/total usage. The durable ledger accumulates that usage into the Run’s token budget. Cache-token detail and normalized retry classification remain connector work. Mesh owns tool grants, validation, execution, persistence, and budgets. A model provider does not decide whether an agent has memory: when an agent has a memory store, every model can request remember. OpenRouter metadata selects only the wire protocol. Models with native tool-call support receive provider-native definitions and return native tool requests; Mesh asks OpenRouter to route those requests only to endpoints that accept the required parameters. Other models receive the same granted definitions in a strict JSON decision protocol. Mesh parses that decision, assigns call ids, validates it against the grant and schema, executes the tool itself, and sends the result back as ordinary model context. A metadata lookup failure uses the same validated fallback rather than removing tools or failing the turn. The decision must be exactly one JSON object with no unknown fields; the only tolerated wrapping is a single Markdown ```json fence and surrounding whitespace. A reply that is not such an object (prose, a truncated object, JSON null, or an object followed by more text) gets one repair: the connector re-asks once, with the model’s own reply as the assistant turn and a short Mesh correction as the last message, and parses the answer with the same strict decoder. Prose is never read as an answer or as a tool having run. The repair runs inside the same Layer A attempt (its context, deadlines, and an ownership recheck), at most once per prepared completion across Layer A’s retries, and a failed repair returns the error the unrepaired path always returned (transient for an unparseable reply, terminal for trailing content). That holds when the repair request itself fails, even terminally (a 400 context overflow from echoing a long reply): the repair is optional and must not stop Layer A from regenerating or falling back. Only the caller’s context ending or an ownership refusal propagates unchanged. A parsed but invalid decision (unknown type, an empty or invalid tool call) is not repaired. For a repaired call the ledger keeps the first request body; the response body is a mesh_fallback_repair envelope with both responses’ exact bytes, the messages the repair appended, and whether the repair request omitted temperature after an earlier rejection, so the repair request is reconstructable exactly from the ledger; usage is the sum of both calls (internal/model/openrouter/fallback_repair.go). The router connector and the OpenAI connector are native-tools only and have no text protocol to repair. The catalog also gates temperature, with two exceptions the catalog cannot express. Anthropic models (anthropic/…) never receive it: Anthropic is deprecating the parameter on newer models, and OpenRouter’s Azure endpoint for them already answers 400 "temperature is deprecated for this model" while Bedrock and Anthropic direct still accept it, so which endpoint OpenRouter picked decided whether a turn failed. And for any model, a 400 whose provider error message (unwrapped from OpenRouter’s error.metadata.raw envelope, never the echoed request) rejects temperature is not terminal: the connector remembers that the model rejects the parameter and returns the failure re-typed as retryable, so Layer A’s bounded loop — with its ownership recheck, backoff, attempt budget and recovery deadline — runs the next attempt, which omits temperature. The connector itself never sends a second request. Later turns skip the parameter up front for the life of the process. Every other 400 stays terminal. The run ledger records the first attempt’s body, as it does for every Layer A retry; per-attempt records are a Layer A follow-up. A failed provider call reports labels, never the provider’s message. Every HTTP connector (openai, openrouter, router) parses the non-2xx body — read to a 64 KiB bound — for the OpenAI-compatible error.type / error.code / error.param (preferring the upstream’s when OpenRouter wraps one in metadata.raw), keeps each only if it has a label’s shape, and renders them with the status: openai: status 400: invalid_request_error (param=tools[3].function.parameters, code=invalid_function_parameters). That string is what the turn-failure record carries, and the model attempt failed log line carries the same fields as status / error_type / error_code / error_param attributes. A 2xx whose body is still a failure (Router’s status: "failed", an OpenAI-compatible {"error":…} on 200) follows the same rule: the connector’s reason plus labels, e.g. router: response failed: (code=server_error). The provider’s message is never rendered, logged or reported, because providers quote the request back in it (a prompt fragment, a tool argument, an image URL); the body is retained only privately on the error for a connector’s own decision, like the temperature rejection above. Classification is unchanged: it reads the status, never the labels. The current loop permits up to eight calls per round across three bounded rounds. Requests for tools that were not offered, invalid arguments, too many calls, and a request beyond the final round are rejected without execution. Successful protocol-capability lookups are cached for the serving client’s lifetime; that cache is an adapter optimization, not a runtime permission boundary. At minimum expose:
  • requested tool name and typed arguments
  • stable provider/request identifiers where available
  • finish/stop reason
  • input/output/cache token usage where available
  • retryable versus terminal provider failure classification
WorkGraph and Run budgets cannot be enforced honestly if token/tool usage exists only inside a provider’s raw payload.

Model roles per agent

An agent names a model per role, not one model:
  • model_primary — reasoning, tool use, final replies, and initial WorkGraph planning
  • model_fast — cheap, low-latency work like complexity triage, execution-shape candidacy, conversation titles, context compaction, and routing
  • model_heartbeat — periodic self-directed activity
All three resolve through the same model connector contract. The split lives in the agent’s own row, so a connector never needs to know which role it is serving, and two agents in the same install can be on entirely different models. Each role is optional and falls back: model_heartbeat to model_fast, model_fast to model_primary, and model_primary to the process default. That is the one thing the environment contributes — a default model for an agent that did not name one, not the agent’s configuration itself. Do not add a model_planner role merely because WorkGraphs exist. Initial graph planning is model_primary work. A dedicated role is justified later only if measurements show planning has a materially different capability/cost profile. Provider selection and API keys are agent-owned runtime state. The setup flow validates the key, stores it encrypted, and records the provider on the RuntimeAgent, so agents in one deployment may use separate provider accounts. OPENROUTER_API_KEY remains only as a legacy deployment-wide fallback for an agent created before provider provisioning existed; a newly provisioned agent does not inherit it, and a missing or unreadable agent-owned key must not fall back silently. See ../agent-provisioning.md. See ../runtime/agent-harness.md for model-role usage and ../runtime/work-graph.md for the graph proposal/admission boundary.

Goal for the first slice

Given a normalized Slack event, the runtime should be able to ask OpenRouter for a completion and route the result back to the Slack connector. The next connector milestone, before WorkGraph execution, is provider-neutral structured output and retry classification. Tool calls are no longer trapped in OpenRouter’s private JSON: the current memory loop persists them as durable RunSteps with observations. The scheduler is still synchronous and interrupted runs are failed rather than resumed, so this does not yet pretend to be the crash-resumable harness.