Request structure
The initial assembled prompt separates three messages:- System: runtime-authored instructions, the agent’s persona and settings, and tool policy.
- Attributed context, in a user message: eligible memory, prior turns, conversation labels, and current-message attribution.
- Current user message: the inbound body and images.
Limits and accounting
The default input ceiling is 131,072 estimated tokens. If the provider reports a model context window, Mesh further limits input to that window minus the requested output reserve. The output default is 8,192 tokens; an engine configuration may override it. The engine’sWithPromptBudget API can set a total input ceiling and memory/history
section ceilings. It is an internal configuration seam, not a dedicated operator
setting or environment variable.
Accounting includes text, tool schemas, call arguments, results, and framing.
The deterministic estimator uses ceil(UTF-8 bytes / 3) plus reserves, including
1,024 tokens for provider protocol overhead and 8,192 per image. These estimates
are admission controls, not exact tokenization or provider billing.
Missing or failed context-window metadata leaves the configured ceiling in place;
it does not establish the unknown model’s capacity. The metadata lookup has a
two-second deadline. Provider retries check each fallback’s advertised window,
so a smaller fallback cannot bypass admission.
Selecting context
Required content is reserved first: persona and policy, identity, current input, attachment descriptions, and the other ongoing request messages and tool schemas. Optional context is reduced in this order:- Enforce any memory and history section ceilings.
- Remove the lowest-ranked memories.
- Remove the oldest history records.
- If required content still cannot fit, fail with a context-budget error before sending the request.
Permissions and audit
Audience eligibility precedes budget selection. Once a model has seen evidence, later pruning does not erase that exposure: memory writes and the final reply retain its lineage. Budgeting cannot turn private facts into public ones. Changed assembled prompts are persisted before their model calls when snapshot storage is configured; a failed snapshot write stops the call. Model-call rows reference the snapshot actually used. Successful reply metadata includescontext_assembly reports with estimates,
ceilings, section costs, and included/omitted counts. The reports contain no
evidence bodies or actor identities. Provider-reported usage remains the source
for usage accounting; these estimates do not replace it.
Implementation: internal/prompt (repository)
and internal/turn/context_budget.go (repository).