Skip to main content
Mesh’s openai provider uses non-streaming OpenAI Chat Completions over HTTP. The same connector serves OpenAI cloud, Ollama, vLLM, and llama.cpp-compatible servers. Mesh does not start a model process or assume it runs on the same host.

Configure Mesh

  1. For OpenAI cloud, leave MESH_OPENAI_API_BASE_URL unset. The API root is https://api.openai.com/v1. Create an OpenAI API key.
  2. For a self-hosted server, set the variable on the Mesh process, then restart Mesh:
    Include the complete API root, normally /v1, not /chat/completions. Private DNS, IPv4, IPv6, localhost, HTTPS, and reverse-proxy path prefixes work. URLs containing credentials, query strings, or fragments are rejected at startup. There is one OpenAI destination per Mesh deployment, not one per agent. Changing it requires restarting Mesh; it is not editable in the UI.
  3. In Instance model defaults, choose OpenAI / compatible server, enter the API key and an exact served primary model ID, and save. OpenAI cloud requires a key; a custom server may accept no authentication. Omit the key when first setting up an unauthenticated server. A blank key preserves any existing key: use Remove instance key to remove a previously stored key.
  4. In agent setup choose Use instance defaults, or on an existing agent’s Models page choose Use instance defaults / Use selected provider’s instance key. Alternatively, select OpenAI directly in agent setup. The connection step shows the configured compatible server and whether its key is optional. Leave the key blank and choose Connect without API key to verify anonymous access, or enter this agent’s own server key. Selecting a key-optional provider on Manage Models also verifies and activates it. A previously stored agent key is retained; entering a new key rotates it.
  5. Choose the primary model from the catalog (or enter the exact ID manually). Leave fast and heartbeat blank to inherit the primary, or name other models served by the same endpoint. Do not use OpenRouter-prefixed IDs unless your server actually serves those IDs. No model ID is guessed by the connector.
MESH_AGENT_MASTER_KEY is required whenever credentials are stored/read and for the instance-settings write path. Keyless agent activation, catalog lookup and persona generation need no encrypted model secret. Other configured features may still require the master key. Provider and model selection are runtime settings, not environment bootstrap. No database migration is required: openai was already part of the provider vocabulary. Keyless activation uses POST /api/agents/{slug}/model-provider/activate with {"provider":"openai"}. The key-save route still requires a nonempty key. Agent detail exposes key_optional and configured_endpoint; has_key remains true only for an actual agent credential. A failed connection probe leaves the current provider unchanged. Unreadable stored credentials are errors, never a reason to retry anonymously. When a key-optional agent’s model credential cannot be decrypted at reload, that agent is withheld from routing until repaired. Switching providers preserves a legacy model key under its original provider so it cannot follow the agent to an anonymous local server.

Credential and network boundaries

The URL override applies consistently to normal turns, fast-model classifiers, setup-time generation, model discovery, and key verification. It does not change OpenRouter or Ramp Router. Keys are encrypted in the existing secret store and never included in URLs. A key is sent as Authorization: Bearer ...; no header is sent when absent. Redirects are refused, including same-origin redirects, so a server cannot redirect prompts or credentials to another destination. Before changing the URL, audit/remove every OpenAI agent and instance key. Changing the deployment URL deliberately retargets all OpenAI calls and their stored credentials; Mesh cannot know whether a new hostname is the same trusted service. Do not forward a real OpenAI key to an untrusted local server. Use TLS and a trusted CA across network boundaries; TLS verification is never disabled. Keep unauthenticated servers behind a firewall or authenticated proxy. Private HTTP is supported for trusted networks but is not encrypted. Respect the Mesh process’s proxy configuration (HTTP_PROXY, HTTPS_PROXY, NO_PROXY). Container localhost refers to the container, not automatically to the host or another pod. The infrastructure API key is a usage-payer credential, not an agent’s external acting identity. Inheriting it does not share connector accounts, conversations, memory, or tool credentials between agents.

Run an inference server

No new inference wrapper needs building. Pick and pin an existing server version, preload its model artifacts, and expose the API on an address Mesh can reach:
  • Ollama: simplest pilot path. Run ollama serve, preload a tool-capable model with ollama pull <model> while online, then use http://<host>:11434/v1 and that exact model tag. Ollama’s default local API does not enforce an OpenAI bearer key. Restrict its listener/firewall or add an authenticating reverse proxy before remote access.
  • vLLM: suitable for GPU-backed concurrent serving. Serve a staged model directory, select a stable served model name, and configure bearer authentication. Native automatic tool calls need the server’s tool-calling options and a parser/chat template compatible with the specific model; these are not universal flags.
  • llama.cpp server: suitable for GGUF/quantized deployments. Stage the GGUF file, configure a served alias and API key as needed, and enable the tool-capable Jinja/chat template appropriate to the model. Use its /v1 root.
Bind remotely only on the intended network interface. Model size, quantization, GPU/CPU RAM, context size, concurrency, and tool reliability must be measured on your hardware; the HTTP adapter does not settle those trade-offs.

API contract and limitations

  • GET {root}/models for catalog/key verification, and POST {root}/chat/completions for inference. A successful catalog probe is not proof that the selected model supports tools or that completion permission is granted. A server that omits /models can use keyless instance defaults and manual model IDs, but keyed setup and keyless agent activation expect this route.
  • Native function tool definitions, assistant tool_calls, matching tool-result messages, roles (including developer), inline image input, finish reasons, request IDs, token usage, and exact prepared-request capture. Select a model and template supporting the roles and modalities Mesh sends; unsupported requests fail at the server, not through a hidden text-tool fallback.
  • Cloud requests use max_completion_tokens; compatible custom endpoints use max_tokens. A proxy serving newer reasoning-only OpenAI models must support that compatibility parameter. Responses-only models/features, audio, streaming, image generation, and automatic embeddings are outside this connector.
  • /models does not reliably describe context size, tool support, vision, or pricing. Mesh does not invent those values. Set the server’s context window and Mesh input/output budgets appropriately; test long conversations explicitly.
  • The connector’s completion timeout is 120 seconds. Mesh’s existing retry policy also applies (default 60-second attempt, 180-second total). Review MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MS and MESH_MODEL_RETRY_TOTAL_DEADLINE_MS for slower hardware. Responses are bounded to 16 MiB. Authentication and malformed requests fail without retry; transient HTTP failures follow normal retry rules.
  • No implicit fallback to OpenAI cloud or OpenRouter on a custom-server failure. Explicit fallback model IDs still use the same connector and destination.
  • Automatic native decision-provider selection falls back to the fast text model for OpenAI. An explicitly configured separate decision provider is independent and must be reviewed for offline deployments.

Deployment acceptance checklist

Automated tests use in-process HTTP servers, not a real GPU/model installation. Before deployment:
  1. Stage Mesh, frontend assets, Postgres, server binaries/containers, tokenizers, chat templates, and model weights before isolating the network. Confirm their licenses, hashes, and reproducible versions. Start the server without downloads.
  2. From Mesh’s network namespace, fetch /v1/models; verify the expected served ID. Configure Mesh, run a normal text exchange, a native tool call/result/final answer round trip, and setup-time generation. Inspect captured provider calls.
  3. Test keyed and (if permitted) keyless access, invalid credentials, an unknown model, server unavailability, cancellation, context overflow, concurrent agents, and restart persistence. Verify per-agent credentials remain distinct.
  4. Verify a custom-server outage cannot send traffic to a cloud destination. Check DNS, proxy, and firewall logs, not just the returned answer.
  5. For air-gapped operation, select lexical-only memory or explicitly provision local embeddings. Audit separate decision providers, connectors, MCP servers, search/browser tools, workspace execution, Sentry/OTLP telemetry, remote image inputs, update checks, and any cloud management dependencies. Existing agents retain their existing settings; choosing OpenAI does not rewrite them.
  6. Enforce default-deny egress and repeat the end-to-end workload, including tool execution. Inference endpoint configuration alone is not an air-gap certification. Qualify the deployment against these checks; adapter tests alone do not establish an air-gapped installation.