openai provider uses non-streaming OpenAI Chat Completions over HTTP.
The same connector serves OpenAI cloud, Ollama, vLLM, and llama.cpp-compatible
servers. Mesh does not start a model process or assume it runs on the same host.
Configure Mesh
-
For OpenAI cloud, leave
MESH_OPENAI_API_BASE_URLunset. The API root ishttps://api.openai.com/v1. Create an OpenAI API key. -
For a self-hosted server, set the variable on the Mesh process, then restart Mesh:
Include the complete API root, normally
/v1, not/chat/completions. Private DNS, IPv4, IPv6, localhost, HTTPS, and reverse-proxy path prefixes work. URLs containing credentials, query strings, or fragments are rejected at startup. There is one OpenAI destination per Mesh deployment, not one per agent. Changing it requires restarting Mesh; it is not editable in the UI. - In Instance model defaults, choose OpenAI / compatible server, enter the API key and an exact served primary model ID, and save. OpenAI cloud requires a key; a custom server may accept no authentication. Omit the key when first setting up an unauthenticated server. A blank key preserves any existing key: use Remove instance key to remove a previously stored key.
- In agent setup choose Use instance defaults, or on an existing agent’s Models page choose Use instance defaults / Use selected provider’s instance key. Alternatively, select OpenAI directly in agent setup. The connection step shows the configured compatible server and whether its key is optional. Leave the key blank and choose Connect without API key to verify anonymous access, or enter this agent’s own server key. Selecting a key-optional provider on Manage Models also verifies and activates it. A previously stored agent key is retained; entering a new key rotates it.
- Choose the primary model from the catalog (or enter the exact ID manually). Leave fast and heartbeat blank to inherit the primary, or name other models served by the same endpoint. Do not use OpenRouter-prefixed IDs unless your server actually serves those IDs. No model ID is guessed by the connector.
MESH_AGENT_MASTER_KEY is required whenever credentials are stored/read and for
the instance-settings write path. Keyless agent activation, catalog lookup and
persona generation need no encrypted model secret. Other configured features
may still require the master key. Provider and model selection are runtime
settings, not environment bootstrap. No database migration is required:
openai was already part of the provider vocabulary.
Keyless activation uses POST /api/agents/{slug}/model-provider/activate with
{"provider":"openai"}. The key-save route still requires a nonempty key.
Agent detail exposes key_optional and configured_endpoint; has_key remains
true only for an actual agent credential. A failed connection probe leaves the
current provider unchanged. Unreadable stored credentials are errors, never a
reason to retry anonymously. When a key-optional agent’s model credential cannot
be decrypted at reload, that agent is withheld from routing until repaired.
Switching providers preserves a legacy model key under its original provider
so it cannot follow the agent to an anonymous local server.
Credential and network boundaries
The URL override applies consistently to normal turns, fast-model classifiers, setup-time generation, model discovery, and key verification. It does not change OpenRouter or Ramp Router. Keys are encrypted in the existing secret store and never included in URLs. A key is sent asAuthorization: Bearer ...; no header is
sent when absent. Redirects are refused, including same-origin redirects, so a
server cannot redirect prompts or credentials to another destination.
Before changing the URL, audit/remove every OpenAI agent and instance key.
Changing the deployment URL deliberately retargets all OpenAI calls and their
stored credentials; Mesh cannot know whether a new hostname is the same trusted
service. Do not forward a real OpenAI key to an untrusted local server. Use TLS
and a trusted CA across network boundaries; TLS verification is never disabled.
Keep unauthenticated servers behind a firewall or authenticated proxy. Private
HTTP is supported for trusted networks but is not encrypted. Respect the Mesh
process’s proxy configuration (HTTP_PROXY, HTTPS_PROXY, NO_PROXY). Container
localhost refers to the container, not automatically to the host or another pod.
The infrastructure API key is a usage-payer credential, not an agent’s external
acting identity. Inheriting it does not share connector accounts, conversations,
memory, or tool credentials between agents.
Run an inference server
No new inference wrapper needs building. Pick and pin an existing server version, preload its model artifacts, and expose the API on an address Mesh can reach:- Ollama: simplest pilot path.
Run
ollama serve, preload a tool-capable model withollama pull <model>while online, then usehttp://<host>:11434/v1and that exact model tag. Ollama’s default local API does not enforce an OpenAI bearer key. Restrict its listener/firewall or add an authenticating reverse proxy before remote access. - vLLM: suitable for GPU-backed concurrent serving. Serve a staged model directory, select a stable served model name, and configure bearer authentication. Native automatic tool calls need the server’s tool-calling options and a parser/chat template compatible with the specific model; these are not universal flags.
- llama.cpp server:
suitable for GGUF/quantized deployments. Stage the GGUF file, configure a served
alias and API key as needed, and enable the tool-capable Jinja/chat template
appropriate to the model. Use its
/v1root.
API contract and limitations
GET {root}/modelsfor catalog/key verification, andPOST {root}/chat/completionsfor inference. A successful catalog probe is not proof that the selected model supports tools or that completion permission is granted. A server that omits/modelscan use keyless instance defaults and manual model IDs, but keyed setup and keyless agent activation expect this route.- Native function tool definitions, assistant
tool_calls, matching tool-result messages, roles (including developer), inline image input, finish reasons, request IDs, token usage, and exact prepared-request capture. Select a model and template supporting the roles and modalities Mesh sends; unsupported requests fail at the server, not through a hidden text-tool fallback. - Cloud requests use
max_completion_tokens; compatible custom endpoints usemax_tokens. A proxy serving newer reasoning-only OpenAI models must support that compatibility parameter. Responses-only models/features, audio, streaming, image generation, and automatic embeddings are outside this connector. /modelsdoes not reliably describe context size, tool support, vision, or pricing. Mesh does not invent those values. Set the server’s context window and Mesh input/output budgets appropriately; test long conversations explicitly.- The connector’s completion timeout is 120 seconds. Mesh’s existing retry policy
also applies (default 60-second attempt, 180-second total). Review
MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MSandMESH_MODEL_RETRY_TOTAL_DEADLINE_MSfor slower hardware. Responses are bounded to 16 MiB. Authentication and malformed requests fail without retry; transient HTTP failures follow normal retry rules. - No implicit fallback to OpenAI cloud or OpenRouter on a custom-server failure. Explicit fallback model IDs still use the same connector and destination.
- Automatic native decision-provider selection falls back to the fast text model for OpenAI. An explicitly configured separate decision provider is independent and must be reviewed for offline deployments.
Deployment acceptance checklist
Automated tests use in-process HTTP servers, not a real GPU/model installation. Before deployment:- Stage Mesh, frontend assets, Postgres, server binaries/containers, tokenizers, chat templates, and model weights before isolating the network. Confirm their licenses, hashes, and reproducible versions. Start the server without downloads.
- From Mesh’s network namespace, fetch
/v1/models; verify the expected served ID. Configure Mesh, run a normal text exchange, a native tool call/result/final answer round trip, and setup-time generation. Inspect captured provider calls. - Test keyed and (if permitted) keyless access, invalid credentials, an unknown model, server unavailability, cancellation, context overflow, concurrent agents, and restart persistence. Verify per-agent credentials remain distinct.
- Verify a custom-server outage cannot send traffic to a cloud destination. Check DNS, proxy, and firewall logs, not just the returned answer.
- For air-gapped operation, select lexical-only memory or explicitly provision local embeddings. Audit separate decision providers, connectors, MCP servers, search/browser tools, workspace execution, Sentry/OTLP telemetry, remote image inputs, update checks, and any cloud management dependencies. Existing agents retain their existing settings; choosing OpenAI does not rewrite them.
- Enforce default-deny egress and repeat the end-to-end workload, including tool execution. Inference endpoint configuration alone is not an air-gap certification. Qualify the deployment against these checks; adapter tests alone do not establish an air-gapped installation.