Skip to main content
Every model request has a bounded input allowance. Memory retrieval alone cannot bound a request: history, images, tool schemas, and tool results also occupy the model’s context window. Mesh applies admission before the first completion and before each later call, including the final response without tools.

Request structure

The initial assembled prompt separates three messages:
  1. System: runtime-authored instructions, the agent’s persona and settings, and tool policy.
  2. Attributed context, in a user message: eligible memory, prior turns, conversation labels, and current-message attribution.
  3. Current user message: the inbound body and images.
The context message is transient and regenerated as admission changes. External messages and memories are not inserted into the system role. Tools and later assistant messages join the request through the ordinary model/tool loop.

Limits and accounting

The default input ceiling is 131,072 estimated tokens. If the provider reports a model context window, Mesh further limits input to that window minus the requested output reserve. The output default is 8,192 tokens; an engine configuration may override it. The engine’s WithPromptBudget API can set a total input ceiling and memory/history section ceilings. It is an internal configuration seam, not a dedicated operator setting or environment variable. Accounting includes text, tool schemas, call arguments, results, and framing. The deterministic estimator uses ceil(UTF-8 bytes / 3) plus reserves, including 1,024 tokens for provider protocol overhead and 8,192 per image. These estimates are admission controls, not exact tokenization or provider billing. Missing or failed context-window metadata leaves the configured ceiling in place; it does not establish the unknown model’s capacity. The metadata lookup has a two-second deadline. Provider retries check each fallback’s advertised window, so a smaller fallback cannot bypass admission.

Selecting context

Required content is reserved first: persona and policy, identity, current input, attachment descriptions, and the other ongoing request messages and tool schemas. Optional context is reduced in this order:
  1. Enforce any memory and history section ceilings.
  2. Remove the lowest-ranked memories.
  3. Remove the oldest history records.
  4. If required content still cannot fit, fail with a context-budget error before sending the request.
Records are removed whole, preserving attribution and meaning. Retained memory is a relevance-ordered prefix; retained history is a chronological suffix. Older tool results have a separate compaction policy. As tool results accumulate, optional memory and history may shrink again. Omitted optional context is not reintroduced later in that turn. These limits do not reset the run’s model-call, tool-call, token-usage, or time allowances. See the harness.

Permissions and audit

Audience eligibility precedes budget selection. Once a model has seen evidence, later pruning does not erase that exposure: memory writes and the final reply retain its lineage. Budgeting cannot turn private facts into public ones. Changed assembled prompts are persisted before their model calls when snapshot storage is configured; a failed snapshot write stops the call. Model-call rows reference the snapshot actually used. Successful reply metadata includes context_assembly reports with estimates, ceilings, section costs, and included/omitted counts. The reports contain no evidence bodies or actor identities. Provider-reported usage remains the source for usage accounting; these estimates do not replace it. Implementation: internal/prompt (repository) and internal/turn/context_budget.go (repository).