Skip to main content

Philosophy

Mesh is being built around a simple critique of existing agent harnesses:
  • they assume one conversation partner at a time
  • they blur participant identity into a generic “user” bucket
  • they treat threads as context blobs rather than structured relationships
  • they make it too easy to lose attribution when multiple people talk in the same place
  • they inherit tools and credentials from one developer’s machine
  • they let model output become runtime policy without an admission boundary
  • they add persistence and explanation after behavior instead of with it
  • they often hide the topology of complex work inside one model loop
  • they treat a promise as a sentence, so work an agent commits to evaporates when the turn that mentioned it ends
That works for a private DM bot doing a bounded job. It does not work for a shared workspace where multiple participants share a thread and agents take on long-running objectives. This document is the reasoning. The normative rules it implies live in exactly one place — roadmap.md §Design guardrails (repository) — because that is the list a PR is actually checked against. Where this document states a principle and that list states a do not, the list is the one to add to.

The failure mode this comes from

The critique is not abstract. It is what happens in any real shared thread. Victor says something. Sergey says something else. An agent replies. Another participant joins later. The system has to know who each message belongs to — and if the harness only sees a generic user bucket, attribution breaks immediately. What follows is predictable:
  • wrong names in replies
  • confused memory
  • broken turn ownership
  • impossible multi-participant collaboration
  • a conversation history that cannot be reconstructed cleanly
The same class of hidden-state failure appears once one agent takes on complicated work. A single loop may internally decide to research three things, verify two of them, retry one branch, and then merge the results. If that topology exists only in transient model context, the runtime cannot reliably explain what remains, what is blocked, why something ran, or how much work the plan can still create. None of those are primarily prompt problems. They are consequences of system models that cannot express the structure they need to govern.

What Mesh is instead

Mesh is an identity-first, multiplayer conversation substrate and durable agent runtime. It treats conversation as a graph of actors, connectors, threads, messages, and turns; treats sufficiently complex work as a separate graph of bounded jobs and dependencies; and gives agents cloud capabilities through explicit contracts rather than through the machine hosting the control plane. That lets Mesh answer two different families of questions: Conversation / perspective
  • who said this?
  • who was it for?
  • what have I discussed with this person today?
  • which thread does this belong to?
  • which runtime owns this turn?
  • how do I reconstruct the conversation across platforms?
Execution / work
  • what is this run trying to do?
  • which jobs are independent?
  • what is blocked on what?
  • which result was verified?
  • which artifacts fed this conclusion?
  • why did this node run now?
  • where is human approval required?
Mesh is the shared runtime underneath agents, not another bot. Slack is the first connector because it exercises the hardest version of the conversation problem; long-running WorkGraphs exercise the hardest version of the execution problem. The short version: Mesh turns messy multiplayer conversation into structured, persistent, explainable perspectives, and turns agent work into bounded, portable, explainable execution.

The identity-first principle

There is no generic user. Each RuntimeAgent also has its own identity, personality, and external accounts. When two agents work on the same GitHub repository, their actions need to appear under their respective identities. A shared bot account erases that distinction at the service, even if Mesh records which agent invoked it internally. Hosting agents together can share infrastructure and billing without pooling who they are. The normative boundary is the agent-owned identity guardrail (repository). An utterance belongs to an actor as observed through a connector inside one RuntimeAgent’s perspective. Those scopes are part of the identity, not metadata attached after a prompt has already collapsed several participants together. If two people write in one thread, the system must preserve who said each thing and who each reply targets. If one person is known through several connectors, linking those identities is an explicit, perspective-owned claim rather than a string match or a deployment-global fact. Actor identity is deliberately opaque. Mesh does not need to classify another participant as human, agent, or service in order to converse with them. It may privately recognize a RuntimeAgent’s own connector identity for routing and loop prevention, but it does not expose peer topology as conversational truth. This keeps attribution stable when a participant moves between implementations or when two agents happen to share one Mesh deployment. Identity resolution therefore precedes prompt assembly, addressing, memory, and execution. A downstream component never receives an unattributed “user” message and attempts to reconstruct the speaker later. See data-model.md and perspective.md.

The conversation-graph principle

Conversation is durable structure, not a prompt transcript. Actors, connector identities, conversations, participants, messages, replies, and turns are related objects. Message events are append-only observations with an author, audience, time, direction, and owning perspective. An outbound reply is an event in the same graph as its trigger, not an unrecorded side effect whose only surviving copy happens to be in Slack. Model context is a bounded projection of that graph for one turn. It is not the canonical store. That distinction lets Mesh retrieve by actor, conversation, time, or connector; preserve attribution through compaction; and rebuild a prompt without pretending the prompt itself is memory. The conversation graph and the WorkGraph are intentionally separate. The first describes who observed whom saying what. The second describes how one admitted objective executes. Work dependencies must not become conversational identity, and conversation edges must not become an accidental scheduler. “Graph” names the relationships that must remain explicit; it does not require a graph database. See persistence.md, runtime/context-resolution.md, and runtime/work-graph.md.

The explicit-state and blank-slate principle

Mesh boots with no agents. That is not a missing feature or a staging convenience. It follows directly from the critique above. The critique is about identity: a harness that assumes one conversation partner cannot represent several. The same error has a deployment-layer version, and it is easier to miss because it hides in configuration rather than in code. If a runtime has to be told at startup which single identity it is — one bot token, one signing secret, one persona, read out of the process environment — then that runtime has one privileged identity baked into it. Everything else becomes a second-class citizen of its own process. Call it the blessed agent. It shows up as:
  • one agent’s credentials living in process config while every other agent’s live somewhere else
  • code paths that are only correct when exactly one agent exists
  • “the agent” appearing in function signatures where an agent id belongs
  • a redeploy being required to add a coworker
Each of those is a small compromise. Together they mean the system can only ever be honest about one participant, which is the thing Mesh exists not to do. So the rule is inverted. Zero agents is a valid, fully supported boot state. Mesh starts empty, comes up healthy, says it has no agents configured, and waits. Agents are created at runtime, through the UI, into the database — all of them, including the first one. There is no first-among-equals because there is no bootstrap path that only the first agent can take. What this buys is the actual product: one running system hosting many agents, each appearing in Slack as its own coworker, with its own credentials, connectors, and models. To everyone else in the channel they are colleagues. Underneath, it is one deployment. That property is not reachable from a design where one agent is configured differently from the rest. The empty first boot is also the honest one. A system that starts empty has to make adding an agent a real, first-class flow, because there is no other way to get one. The broader rule is that mutable product state is explicit, owned, and durable. Agents, connectors, provider choices, credentials, personas, memory, prompt snapshots, tools, grants, and execution backends are runtime state. Their canonical representation belongs in the database, with ownership and history appropriate to the object. Markdown, process memory, and generated prompts are views of that state, never secret canonical stores. Environment variables configure the deployment boundary: database access, listen address, public URL, telemetry, and the master key or KMS path needed to decrypt stored secrets. They do not define an agent. The process-level OPENROUTER_API_KEY survives only as a migration fallback for agents created before provider provisioning existed; it is not the path for a fresh agent. Adding an agent, rotating its connector, changing its provider, or granting a tool must not require a redeploy. See agent-provisioning.md and versioned-state.md.

The cloud-first principle

The control-plane host is not the agent’s computer. Most agent harnesses inherit their capabilities from the machine on which they run: its working directory becomes the agent’s workspace, its browser becomes the agent’s browser, its network becomes the agent’s network, and credentials already present for an interactive user become the agent’s credentials. That is a convenient development shortcut. It is not a production architecture for a deployment hosting dozens of agents. Mesh assumes the opposite. The control plane may run on an interchangeable cloud instance with no project checkout, no interactive browser, no developer account, and no useful ambient credentials. Merely placing a file, executable, token, or logged-in session on that host must not expand what any RuntimeAgent can do. Every capability crosses a declared service boundary:
  • query tools read perspective-scoped Mesh state or call a registered hosted provider
  • workspace tools run in an isolated, run-scoped execution backend
  • browser tools use an explicitly provisioned cloud browser session
  • external effects go through a connector or registered provider with attributable delivery state
  • third-party extensions use registered remote MCP servers, not local subprocesses of the control plane
Credentials follow the capability. They are scoped to the agent, connector, provider, and run that need them; minted or decrypted only at the boundary; budgeted, audited, and revocable. They are never inherited wholesale from the control-plane process. This is simultaneously a scale, isolation, and reproducibility rule. A heavy build cannot consume the control plane that schedules every other agent. Two runs cannot collide in one implicit working directory. Moving the control plane to a fresh host cannot silently remove an agent’s tools or expose a different set of files. The same contracts can be backed by a managed service, an operator’s Kubernetes cluster, or another cloud execution system without changing the harness. Cloud-first does not mean public-SaaS-only. Operators may keep execution inside their own perimeter, and Mesh may provide an explicit local backend for development. It means locality is a configured backend with an honest security boundary, never ambient authority or a production default. The portability test is straightforward: given the same database, declared providers, backends, and grants, moving Mesh’s control plane to a clean host must not change an agent’s capabilities. See runtime/tools.md and runtime/sandboxed-execution.md.

The perspective principle

Co-residency is not shared cognition. A single Mesh deployment may host Lyra and Daedalus, but that is a control-plane fact. It must not give Lyra privileged conversational knowledge about Daedalus, or vice versa. From Lyra’s perspective, Daedalus is another participant observed through a connector. He may be a human, an externally hosted agent, a service, or a RuntimeAgent in the same Mesh process. Mesh does not expose that distinction to her merely because the control plane happens to know it. This produces a hard boundary:
  • the control plane may know RuntimeAgent mappings for routing, self-recognition, authorization, budgets, loop prevention, graph scheduling, and auditability
  • an agent perspective contains only the actors, conversations, messages, memory, and state that agent is entitled to observe
  • control-plane facts do not become prompt or memory context unless they were learned through the external conversation like any other fact
The portability test is simple: moving Daedalus from this Mesh deployment to a completely unrelated implementation must not change anything Lyra can observe, retrieve, remember, or infer about him, provided his external connector behavior stays the same. That means actors are deliberately opaque. Mesh does not classify an Actor as human, agent, or service. Self-recognition is privileged because an agent must be able to recognize its own connector identity; peer-recognition is not. A RuntimeAgent may know “this is me” without being told “that participant over there is another RuntimeAgent.” It also means perspectives do not share actors or memory. If Lyra and Daedalus both know Victor, “Victor as known by Lyra” and “Victor as known by Daedalus” are allowed to be distinct graph identities with distinct histories and memory. Cross-connector identity linkage is scoped to the perspective that established it. Internal WorkGraph workers do not break this rule. A fresh worker context is an execution child of one RuntimeAgent, not another persistent agent perspective. See perspective.md for the normative boundary.

The participation principle

Being in the room is not being asked a question. This is the same critique as the blank slate principle, pointed at behavior instead of at identity. A harness that assumes one conversation partner also assumes every message is for it, because in a DM that assumption is free and correct. Point that assumption at a shared channel and it becomes: every message anyone posts, about anything, to anyone, is a request directed at the agent. That is not a rate-limiting problem. It is the identity error again. The agent does not know how many actors are in the room, so it cannot know that it wasn’t the one being addressed. It has no concept of overhearing. So Mesh models conversation shape as a first-class property, and engagement is a function of it. Two actors: everything is for me. Fifteen actors: almost nothing is, and I need evidence before I speak. The agent reads the whole room and answers when addressed — which is what every human in that channel is already doing. The naive alternative is to require an explicit @-mention on every message. That is available as a per-conversation override and it is the wrong default, because it converts a colleague into a ticket form: no follow-up without re-addressing, no “hey Lyra” then “specifically the retry budget.” It also hides the real defect rather than fixing it. The agent still has no idea how many people are in the room; it has just been muzzled. Two corollaries that fall out of taking this seriously:
  • Deciding whether to answer must be far cheaper than answering. If engaging and evaluating cost the same, an ambiguous channel message costs a full agent loop, and a busy channel is a self-inflicted denial of service.
  • Not answering is explainable too. every agent turn is explainable has to cover the turn that did not happen, or people will work around the gate by @-mentioning everything and we will have shipped require_mention by social convention.
See runtime/addressing.md.

The adapter principle

Slack is not the application, OpenRouter is not the intelligence layer, and a sandbox vendor is not the execution architecture. Each external system sits behind the narrow contract appropriate to its role. A conversation connector translates wire events and delivery operations; it does not own identity, addressing, memory, or turn policy. A model connector translates prompts, structured output, tool-call syntax, and usage; it does not decide which tools the agent is allowed to use or execute them on Mesh’s behalf. An execution or browser backend provisions declared resources; it does not set their grants, budgets, or lifecycle policy. That separation is what makes plurality real. One RuntimeAgent may have several Slack installations without becoming several agents. Another connector may carry the same conversational contracts without copying Slack assumptions into the graph. A model that lacks a provider-native tool-call envelope can still participate through Mesh’s strict structured fallback because the harness—not the provider—owns the tool loop. A backend can move from a managed service to an operator’s own perimeter without leaking vendor branches into the harness. Adapters are allowed to expose capabilities and limitations. They are not allowed to redefine the semantic layer above them. Code above a boundary may branch on a declared capability, never on a vendor name. The anatomy a seam needs before that holds under growth — contract, descriptor, catalog — and the layer model this principle lives inside are in layers.md. See also connectors/slack.md, connectors/model.md, and runtime/sandboxed-execution.md.

The work topology principle

A loop is the right abstraction for one coherent job. It is not the only execution topology Mesh should be able to express. When a complex objective decomposes into independent work, verification, joins, or approval boundaries, that topology should become a first-class WorkGraph rather than remaining a private plan inside one context window. This does not mean “use more agents.” It means the runtime should be able to represent the actual dependency structure of the work. The governing distinction is:
  • loop — one bounded worker repeatedly reasons/acts toward one coherent objective
  • graph — an admitted DAG coordinates several bounded jobs whose dependencies matter
The simple case remains simple. An inherently sequential task stays one loop. A graph is earned by real decomposition, not by the number of boxes somebody can draw. For graph mode:
  • the model may propose nodes and dependencies
  • Mesh deterministically validates and admits the topology
  • only admitted ready nodes may execute
  • parallel workers get fresh bounded contexts, not new RuntimeAgent identities
  • workers hand off typed artifacts instead of broadcasting transcripts
  • verification and human gates are explicit nodes
  • loops stay inside nodes; V1 does not add arbitrary graph-level cycles
  • replanning creates immutable graph revisions instead of rewriting history
  • fan-out remains under the same budget, authority, and concurrency ceilings as the root Run
That boundary follows the same philosophical pattern as the rest of Mesh: models may exercise judgment inside explicit contracts; the runtime owns invariants. See runtime/work-graph.md.

The runtime-governance principle

Models propose; the runtime governs. The model is responsible for judgment: interpreting a request, choosing a strategy, deciding whether a granted tool would help, proposing structured work, and synthesizing a result. It is not the authority for what may execute. Mesh validates every structured proposal and owns the invariants around it:
  • addressing and single-flight admission decide whether a root Run exists
  • tool policy decides which capabilities are offered in this conversation
  • schema validation decides whether proposed arguments are executable
  • budgets bound iterations, tools, tokens, time, compute, and graph fan-out
  • graph admission decides which proposed nodes and dependencies become runnable
  • effect policy and attributed approval govern externally visible actions
  • the commit boundary decides whether interruption may cancel or must reconcile
  • persistence decides what state survives a crash and what work may resume
Provider-native tool calls are one wire encoding, not delegated execution. Mesh may use them when available or a strict structured fallback when they are not; either way, the same harness validates, invokes, records, and returns the observation. A provider metadata failure may change encoding but cannot revoke a tool that Mesh granted. Deterministic governance does not mean replacing model judgment with a giant rules engine. It means judgment happens inside typed, bounded contracts whose safety properties do not depend on the model remembering to enforce them. See runtime/agent-harness.md, runtime/tools.md, and runtime/interrupt-model.md.

The capture-before-behavior principle

Explanation is part of the feature, not an observability follow-up. Every consequential transition must leave a durable record connected to the RuntimeAgent perspective, conversation, trigger, and Run that caused it. That includes engagement verdicts, model requests and responses, prompt snapshots, token usage, raw and validated tool arguments, observations, memory writes, budget outcomes, approvals, external effects, failures, and the decision to remain silent. What the model requested and what the runtime actually executed are separate facts and neither may overwrite the other. Capture happens with the behavior. A tool call cannot execute first and acquire an audit row later; a crash between those actions would make the most important event the one Mesh cannot explain. External effects additionally require idempotent delivery state and a recorded commit boundary because replay after a crash must not post twice. The database is the system of record. Logs and traces are operational lenses, not substitute memory, and operator UI is a client of the same read API that makes the history scriptable. Display may arrive after capture. Capture may never arrive after behavior. The test is blunt: if an operator cannot reconstruct from durable data what the agent observed, why it acted or did not act, what authority it used, what it cost, and how it ended, the behavior is not ready to ship. See runtime/observability.md, persistence.md, and runtime/turn-pipeline.md.

The thing Mesh should be

Mesh should behave like a multiplayer conversation substrate and durable execution runtime:
  • every message has an actor within one agent’s perspective
  • every actor has connector-specific identity within that perspective
  • every thread can be reconstructed from that perspective
  • every reply has a target
  • every mutable capability has an explicit owner, grant, and revocation path
  • every agent turn is explainable
  • every agent knows how many actors are in the room, and behaves differently because of it
  • no agent gains hidden knowledge because another participant is co-resident in Mesh
  • no agent gains tools or credentials from the machine hosting Mesh
  • no model response becomes executable merely because the model emitted it
  • complex work can expose its dependency topology instead of hiding it inside one loop
  • every admitted WorkGraph node has a reason it became runnable and attributable inputs/outputs
  • every consequential action is captured before it can disappear into the world
  • silence is a decision with a recorded reason, not an absence

What we are optimizing for

  • attribution
  • continuity across platforms within an agent’s perspective
  • correct turn ownership
  • low-friction connector support
  • runtime provisioning without redeployment
  • traceable agent behavior
  • strong isolation between agent perspectives
  • cloud-scale execution without ambient host authority
  • explicit, portable capability boundaries
  • bounded, resumable work with attributable cost
  • durable, inspectable work topology for complex objectives
  • dependency-aware parallelism when the task genuinely supports it
  • independent verification and explicit effect gates
  • eventual contribution from many people, not just one operator

What we are not optimizing for

  • clever prompt tricks that paper over identity problems
  • a giant monolithic “assistant” abstraction
  • vendor-native model or tool behavior as runtime policy
  • pretending all participants are interchangeable
  • an agent that mistakes presence in a channel for being addressed by it
  • hiding the conversation structure inside an opaque memory blob
  • hiding a complex execution plan inside an opaque model loop
  • a privileged first agent configured through the environment
  • treating the control-plane host as an agent’s workstation
  • inheriting files, browser sessions, network reach, or user credentials from that host
  • local-only tool protocols as the production extension model
  • mutable product state hidden in prompts, files, or process memory
  • behavior whose audit trail is deferred to a later observability pass
  • hidden shared cognition between agents merely because they share a process
  • fan-out for its own sake
  • turning every temporary worker into a persistent agent identity

Preferred design style

  • contract first
  • explicit types
  • small stable primitives
  • append-only and versioned state where history matters
  • the database and event log as source of truth
  • connectors as replaceable adapters
  • capability providers and execution backends as replaceable adapters
  • least-authority credentials scoped to the capability and run
  • perspective ownership enforced in data access, not only in prompts
  • model proposals separated from runtime validation, admission, and scheduling
  • capture designed into the write path before display
  • typed hand-offs over transcript broadcast
  • UI that reflects the model instead of inventing one

Language choice

The current bias is Go for v1 because the problem is mostly about concurrency, connectors, workers, graph scheduling, and API surfaces. Rust is still attractive for performance-sensitive pieces, but it should not become an entry fee that delays the product.

Naming bias

“Mesh” is good because it implies:
  • linked nodes
  • many participants
  • structure
  • texture / weave / fabric
That is the right emotional shape for the product.