> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI and self-hosted models

> Connect Mesh to OpenAI or an operator-hosted compatible inference endpoint.

Mesh's `openai` provider uses non-streaming OpenAI Chat Completions over HTTP.
The same connector serves OpenAI cloud, Ollama, vLLM, and llama.cpp-compatible
servers. Mesh does not start a model process or assume it runs on the same host.

## Configure Mesh

1. For OpenAI cloud, leave `MESH_OPENAI_API_BASE_URL` unset. The API root is
   `https://api.openai.com/v1`. Create an OpenAI API key.
2. For a self-hosted server, set the variable on the **Mesh process**, then restart Mesh:

   ```sh theme={null}
   MESH_OPENAI_API_BASE_URL=http://model.internal:8000/v1
   # Other examples (choose one):
   # http://127.0.0.1:11434/v1   — Ollama in the same network namespace
   # http://192.168.10.20:8080/v1 — llama.cpp on another server
   ```

   Include the complete API root, normally `/v1`, not `/chat/completions`.
   Private DNS, IPv4, IPv6, localhost, HTTPS, and reverse-proxy path prefixes
   work. URLs containing credentials, query strings, or fragments are rejected
   at startup. There is one OpenAI destination per Mesh deployment, not one per
   agent. Changing it requires restarting Mesh; it is not editable in the UI.
3. In **Instance model defaults**, choose **OpenAI / compatible server**, enter
   the API key and an exact served primary model ID, and save. OpenAI cloud
   requires a key; a custom server may accept no authentication. Omit the key
   when first setting up an unauthenticated server. A blank key preserves any
   existing key: use **Remove instance key** to remove a previously stored key.
4. In agent setup choose **Use instance defaults**, or on an existing agent's
   Models page choose **Use instance defaults** / **Use selected provider's
   instance key**. Alternatively, select **OpenAI** directly in agent setup.
   The connection step shows the configured compatible server and whether its
   key is optional. Leave the key blank and choose **Connect without API key**
   to verify anonymous access, or enter this agent's own server key. Selecting
   a key-optional provider on Manage Models also verifies and activates it.
   A previously stored agent key is retained; entering a new key rotates it.
5. Choose the primary model from the catalog (or enter the exact ID manually).
   Leave fast and heartbeat blank to inherit the primary, or name other models
   served by the same endpoint. Do not use OpenRouter-prefixed IDs unless your
   server actually serves those IDs. No model ID is guessed by the connector.

`MESH_AGENT_MASTER_KEY` is required whenever credentials are stored/read and for
the instance-settings write path. Keyless agent activation, catalog lookup and
persona generation need no encrypted model secret. Other configured features
may still require the master key. Provider and model selection are runtime
settings, not environment bootstrap. No database migration is required:
`openai` was already part of the provider vocabulary.

Keyless activation uses `POST /api/agents/{slug}/model-provider/activate` with
`{"provider":"openai"}`. The key-save route still requires a nonempty key.
Agent detail exposes `key_optional` and `configured_endpoint`; `has_key` remains
true only for an actual agent credential. A failed connection probe leaves the
current provider unchanged. Unreadable stored credentials are errors, never a
reason to retry anonymously. When a key-optional agent's model credential cannot
be decrypted at reload, that agent is withheld from routing until repaired.
Switching providers preserves a legacy model key under its original provider
so it cannot follow the agent to an anonymous local server.

### Credential and network boundaries

The URL override applies consistently to normal turns, fast-model classifiers,
setup-time generation, model discovery, and key verification. It does not change
OpenRouter or Ramp Router. Keys are encrypted in the existing secret store and
never included in URLs. A key is sent as `Authorization: Bearer ...`; no header is
sent when absent. Redirects are refused, including same-origin redirects, so a
server cannot redirect prompts or credentials to another destination.

**Before changing the URL, audit/remove every OpenAI agent and instance key.**
Changing the deployment URL deliberately retargets all OpenAI calls and their
stored credentials; Mesh cannot know whether a new hostname is the same trusted
service. Do not forward a real OpenAI key to an untrusted local server. Use TLS
and a trusted CA across network boundaries; TLS verification is never disabled.
Keep unauthenticated servers behind a firewall or authenticated proxy. Private
HTTP is supported for trusted networks but is not encrypted. Respect the Mesh
process's proxy configuration (`HTTP_PROXY`, `HTTPS_PROXY`, `NO_PROXY`). Container
localhost refers to the container, not automatically to the host or another pod.

The infrastructure API key is a usage-payer credential, not an agent's external
acting identity. Inheriting it does not share connector accounts, conversations,
memory, or tool credentials between agents.

## Run an inference server

No new inference wrapper needs building. Pick and pin an existing server version,
preload its model artifacts, and expose the API on an address Mesh can reach:

* [Ollama](https://docs.ollama.com/api/openai-compatibility): simplest pilot path.
  Run `ollama serve`, preload a tool-capable model with `ollama pull <model>`
  while online, then use `http://<host>:11434/v1` and that exact model tag.
  Ollama's default local API does not enforce an OpenAI bearer key. Restrict its
  listener/firewall or add an authenticating reverse proxy before remote access.
* [vLLM](https://docs.vllm.ai/en/latest/serving/online_serving/openai_compatible_server/):
  suitable for GPU-backed concurrent serving. Serve a staged model directory,
  select a stable served model name, and configure bearer authentication. Native
  automatic tool calls need the server's tool-calling options and a parser/chat
  template compatible with the specific model; these are not universal flags.
* [llama.cpp server](https://github.com/ggml-org/llama.cpp/tree/master/tools/server):
  suitable for GGUF/quantized deployments. Stage the GGUF file, configure a served
  alias and API key as needed, and enable the tool-capable Jinja/chat template
  appropriate to the model. Use its `/v1` root.

Bind remotely only on the intended network interface. Model size, quantization,
GPU/CPU RAM, context size, concurrency, and tool reliability must be measured on
your hardware; the HTTP adapter does not settle those trade-offs.

## API contract and limitations

* `GET {root}/models` for catalog/key verification, and
  `POST {root}/chat/completions` for inference. A successful catalog probe is
  not proof that the selected model supports tools or that completion permission
  is granted. A server that omits `/models` can use keyless instance defaults and
  manual model IDs, but keyed setup and keyless agent activation expect this route.
* Native function tool definitions, assistant `tool_calls`, matching tool-result
  messages, roles (including developer), inline image input, finish reasons,
  request IDs, token usage, and exact prepared-request capture. Select a model
  and template supporting the roles and modalities Mesh sends; unsupported
  requests fail at the server, not through a hidden text-tool fallback.
* Cloud requests use `max_completion_tokens`; compatible custom endpoints use
  `max_tokens`. A proxy serving newer reasoning-only OpenAI models must support
  that compatibility parameter. Responses-only models/features, audio, streaming,
  image generation, and automatic embeddings are outside this connector.
* `/models` does not reliably describe context size, tool support, vision, or
  pricing. Mesh does not invent those values. Set the server's context window and
  Mesh input/output budgets appropriately; test long conversations explicitly.
* The connector's completion timeout is 120 seconds. Mesh's existing retry policy
  also applies (default 60-second attempt, 180-second total). Review
  `MESH_MODEL_RETRY_ATTEMPT_TIMEOUT_MS` and `MESH_MODEL_RETRY_TOTAL_DEADLINE_MS` for
  slower hardware. Responses are bounded to 16 MiB. Authentication and malformed
  requests fail without retry; transient HTTP failures follow normal retry rules.
* No implicit fallback to OpenAI cloud or OpenRouter on a custom-server failure.
  Explicit fallback model IDs still use the same connector and destination.
* Automatic native decision-provider selection falls back to the fast text
  model for OpenAI. An explicitly configured separate decision provider is
  independent and must be reviewed for offline deployments.

## Deployment acceptance checklist

Automated tests use in-process HTTP servers, not a real GPU/model installation.
Before deployment:

1. Stage Mesh, frontend assets, Postgres, server binaries/containers, tokenizers,
   chat templates, and model weights before isolating the network. Confirm their
   licenses, hashes, and reproducible versions. Start the server without downloads.
2. From Mesh's network namespace, fetch `/v1/models`; verify the expected served
   ID. Configure Mesh, run a normal text exchange, a native tool call/result/final
   answer round trip, and setup-time generation. Inspect captured provider calls.
3. Test keyed and (if permitted) keyless access, invalid credentials, an unknown
   model, server unavailability, cancellation, context overflow, concurrent agents,
   and restart persistence. Verify per-agent credentials remain distinct.
4. Verify a custom-server outage cannot send traffic to a cloud destination.
   Check DNS, proxy, and firewall logs, not just the returned answer.
5. For air-gapped operation, select lexical-only memory or explicitly provision
   local embeddings. Audit separate decision providers, connectors, MCP servers,
   search/browser tools, workspace execution, Sentry/OTLP telemetry, remote image
   inputs, update checks, and any cloud management dependencies. Existing agents
   retain their existing settings; choosing OpenAI does not rewrite them.
6. Enforce default-deny egress and repeat the end-to-end workload, including tool
   execution. Inference endpoint configuration alone is **not** an air-gap
   certification. Qualify the deployment against these checks; adapter tests alone do not
   establish an air-gapped installation.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.