> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mesh.texturehq.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Work graph

# Work graph

Implementation note: the [finite worker-task storage foundation (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/worker-tasks.md)
provides a separate child-task ledger, not WorkGraph execution or a launchable
sub-agent. See that contract for implemented lifecycle behavior and enablement gates.

A loop is how one worker does a job. A graph is how a RuntimeAgent organizes work.

Mesh needs both.

The agent harness already treats a useful agent turn as a bounded loop: inspect state, call a model, use tools, observe the result, repeat until done. That is the right primitive for a locally coherent job. It is the wrong primitive for every large objective, because a single loop hides the shape of the work inside one context window.

When an objective naturally decomposes, Mesh should make that decomposition explicit as a **WorkGraph**: a durable, inspectable graph of bounded jobs, dependencies, verification, joins, and human gates.

> **Loops are nodes. The WorkGraph is the plan.**

This document owns that contract.

## Why this is a first-class primitive

A complex request often contains several different shapes of work at once:

* some tasks are genuinely sequential
* some evidence can be gathered independently
* some results should be checked by a fresh verifier
* some branches should fan out in parallel
* some results must join before a conclusion is possible
* some effects require a human decision before they are allowed to happen

A single self-directed loop can discover and execute all of that implicitly. The problem is that the topology then exists only inside model reasoning and transient context. Mesh cannot reliably answer:

* what is the plan?
* which work is independent?
* why did this task run now?
* what is blocked on what?
* which branch failed?
* which result did this conclusion depend on?
* where can a human safely intervene?
* how much budget remains for the unfinished graph?
* what changed when the plan was revised?

Those are exactly the questions a runtime that values auditability should be able to answer without asking the model to narrate what it thinks it did.

The WorkGraph moves topology out of hidden reasoning and into durable runtime state.

## This is not the conversation graph

Mesh now has two graph-shaped domains with very different meanings.

**The conversation graph** is epistemic. It is one RuntimeAgent's perspective on actors, messages, conversations, and remembered facts. It answers: *what has this agent observed?*

**The WorkGraph** is operational. It is the execution topology for one objective. It answers: *what work remains, what depends on what, and what may run next?*

Do not reuse `ThreadEdge`, conversation identity, or Actor relationships to represent work dependencies. The fact that both domains are graphs does not make them the same graph.

A WorkGraph belongs to exactly one RuntimeAgent and one root Run. It never creates shared cognition between RuntimeAgents. See [`../perspective.md`](/perspective).

## The portability boundary still holds

A WorkGraph is an internal execution artifact of one RuntimeAgent.

If Lyra decides that a request needs five independent research jobs, those jobs do not become five new `RuntimeAgent` rows, five Slack coworkers, or five Actors. They are bounded worker executions owned by Lyra's run and Lyra's perspective.

That distinction is load-bearing:

* a `RuntimeAgent` is a durable product identity with connectors, memory, and a perspective
* a WorkGraph worker is an ephemeral execution context for one bounded node
* workers do not own connector identities
* workers do not appear in conversations
* workers do not get private long-term memory
* workers do not learn hidden facts about co-resident RuntimeAgents
* every worker reads only context authorized from the parent RuntimeAgent's perspective

Calling a worker a "subagent" in implementation code is fine if useful. It must not accidentally turn into a second class of persistent Mesh agent identity.

## Graphs are not mandatory fan-out

Graph engineering is not "use more agents."

The default shape for a small or inherently sequential task is still one bounded loop. Conceptually that is a WorkGraph with one agent-loop node. An implementation does not need to materialize a graph row for the degenerate case until doing so is useful.

Mesh should choose graph execution when the **shape of the work** warrants it, for example:

* independent branches can run without reading one another's results
* a large objective needs explicit decomposition to fit bounded contexts
* independent verification is materially useful
* different branches require different tools, budgets, or execution environments
* a human gate belongs between evidence and an irreversible action
* a fan-in step needs to merge several independently produced artifacts

Do not graph a five-step chain merely because it can be drawn as five boxes. If every step genuinely needs the previous step's full result, one loop is cheaper, simpler, and often more accurate.

This is an admission decision, not an aesthetic one.

## Execution shape

A root Run has an execution shape:

```text theme={null}
loop   — one bounded agent loop owns the objective

graph  — a WorkGraph owns the objective and schedules bounded node executions
```

The first classifier can be conservative. `model_fast` may classify whether the task is plausibly decomposable, while `model_primary` does any actual planning. A classifier failure defaults to `loop`, not to speculative fan-out.

A graph can also emerge later. A Run that began as a loop may discover that the objective has several independent branches and propose a WorkGraph before crossing its external-effect commit boundary.

## The graph lifecycle

A WorkGraph moves through five conceptual stages:

```text theme={null}
propose -> admit -> execute -> revise -> terminal
```

### 1. Propose

The planner proposes explicit nodes, dependencies, expected outputs, and budgets.

The proposal is typed data, not prose that a scheduler tries to interpret. At minimum each node proposal names:

* stable node key
* objective
* node kind
* dependencies
* required inputs
* declared output schema/artifact type
* allowed tool subset
* budget
* side-effect classification

### 2. Admit

**The model never directly schedules its own proposed graph.**

Mesh validates the proposal deterministically before any node can run. The admission gate checks invariants such as:

* all referenced nodes exist
* graph is acyclic
* node count is below the configured cap
* depth is below the configured cap
* fan-out is below the configured cap
* the graph has at least one reachable terminal path
* node budgets fit inside the parent Run budget
* requested tools are subsets of the parent Run's grants
* no node increases the parent Run's authority
* effectful nodes satisfy the required approval policy
* declared node kinds exist in the runtime
* artifact dependencies are satisfiable
* no two nodes are configured to mutate one non-isolated working state concurrently

Rejection is structured. The planner may receive the list of failed constraints and propose a new revision, subject to a replan budget.

The invariant is simple:

> **The model proposes topology. The runtime admits topology.**

### 3. Execute

The scheduler executes only nodes whose dependencies and conditions are satisfied.

Independent ready nodes may run concurrently, subject to all existing limits:

* per-graph parallelism cap
* per-RuntimeAgent concurrency ceiling
* per-install concurrency ceiling
* workspace/compute limits
* model/provider rate limits

Graph fan-out does not bypass admission control merely because the planner drew parallel edges.

### 4. Revise

Real work discovers information the planner did not have.

A WorkGraph is therefore versioned rather than frozen. After a node succeeds, fails, or reveals a new dependency, the planner may propose a new graph revision.

Revisions are append-only and admitted by the same deterministic gate as the original proposal.

Completed history is never rewritten. A revision may:

* add future nodes
* add or replace dependencies among unstarted work
* supersede pending nodes
* add a verifier or human gate
* collapse work that is no longer necessary
* route around a failed branch

A revision may not pretend an already completed or externally committed effect never happened.

### 5. Terminal

A graph completes when its admitted terminal condition is satisfied, or ends legibly with a terminal failure such as:

* failed
* cancelled
* budget\_exhausted
* blocked
* rejected

The user-visible root Run owns the final answer and terminal status. Worker failures are graph state, not orphaned background errors.

## V1 is a DAG of loops

At the WorkGraph level, V1 is deliberately a **directed acyclic graph**.

Loops already exist where they belong: inside bounded agent-loop nodes, with iteration, token, tool-call, wall-clock, and compute budgets.

Allowing arbitrary graph cycles immediately creates a second looping mechanism with harder reasoning about termination, replay, and side effects. Mesh does not need two different places where infinity can hide.

Retries do not require cycles. A failed node gets another attempt. Rework does not require cycles. A graph revision can add a replacement node. Iterative reasoning does not require cycles. The worker loop already owns it.

If a future use case genuinely requires graph-level cycles, add them with explicit exit conditions after the DAG scheduler is proven. Do not accidentally acquire them through a permissive edge model.

## Node kinds

The graph contract should stay small. V1 needs only a few semantic node kinds.

### `agent_loop`

A bounded model/tool loop with one objective.

An agent-loop node receives a deliberately narrow context assembled from:

* the node objective
* the required upstream artifacts
* the minimum relevant slice of the owning RuntimeAgent's perspective
* the allowed tool set
* the node budget

It does **not** automatically receive every sibling node's transcript or the entire parent Run history.

For durable execution, an `agent_loop` node is implemented as a child Run or equivalent child execution unit under the root Run. It may have its own `RunStep`s and, when exec-class tools are needed, its own isolated workspace.

### `deterministic`

A runtime-defined function that does not need an LLM to decide what to do.

Examples include:

* normalize several typed artifacts
* compute a score
* run a schema validator
* execute a fixed test command
* select values using a declared deterministic rule

Do not spend model tokens on control flow the runtime can express directly.

### `verify`

An independent check of an upstream artifact.

Verification is first-class because "the worker says it checked" is weaker than an independently executed check. A verifier may be deterministic or model-backed, but it runs in a fresh context and receives the artifact plus the verification rubric, not the producing worker's private reasoning transcript.

A verifier produces structured evidence: pass, fail, uncertain, plus findings or measurements.

### `human_gate`

A node that pauses the graph until an authorized actor approves, rejects, or supplies required input.

Place human gates where the next edge crosses an expensive or irreversible boundary, not reflexively before every action.

The gate's decision is persisted and attributed to the actor who made it.

### `join`

A controlled fan-in point that owns the merge of several upstream artifacts.

A join exists because parallel workers should not silently race to mutate one shared artifact. One node owns the merge, conflict handling, and resulting canonical artifact.

The merge may itself be deterministic or model-backed, but the graph has one explicit owner for it.

## Edges are dependencies, not conversation

A WorkEdge means that execution or data dependency flows from one node to another.

At minimum:

```text theme={null}
A -> B
```

means B cannot become ready until A satisfies the dependency.

Edges may also carry a declared condition over **typed outputs**, for example:

```text theme={null}
verify.result == "pass"
```

Avoid arbitrary model-generated code or expressions as edge conditions. Conditions should be drawn from a small runtime-defined predicate language over validated artifacts.

## Typed artifacts are the hand-off contract

The graph should move **artifacts**, not giant concatenated transcripts.

A WorkArtifact is a durable output produced by a node and consumed by downstream nodes. Examples:

* structured research findings
* a list of files
* a patch
* test results
* a decision record
* a citation set
* a generated plan fragment
* a sandbox snapshot reference

Every artifact records provenance:

* graph and revision
* producing node
* producing attempt
* schema/type
* creation time
* source inputs where relevant

For structured outputs, the producer should declare a JSON Schema or equivalent typed contract and Mesh should validate it before unlocking dependent nodes.

This buys two things at once:

1. **context economy** — downstream workers receive the compact product of upstream work instead of all upstream context
2. **auditability** — an edge carries an inspectable hand-off rather than "whatever that agent happened to say"

## Fresh contexts, shared perspective

Parallel worker nodes should usually execute with fresh model contexts.

Fresh context is useful precisely because it prevents one worker's framing, mistakes, and token load from contaminating every sibling. But fresh context does **not** mean a fresh Mesh identity.

All workers remain inside the root RuntimeAgent's perspective and authority boundary.

A child worker may receive a strict subset of the parent's context and tools. It may never receive more.

> **Delegation may preserve or reduce authority; it may not increase it.**

## Verification and the diamond

The highest-value graph pattern is often a diamond:

```text theme={null}
                -> worker A -> verifier A -
root / scope ---|                         |-> join / synthesize
                -> worker B -> verifier B -
                -> worker C -> verifier C -
```

Use it when work genuinely splits.

The separation matters:

* workers own bounded discovery or production
* verifiers independently test claims/results
* one join owns reconciliation

Do not let three workers all mutate one final artifact and call the resulting race "collaboration."

## Relationship to the commit boundary

[`interrupt-model.md`](/runtime/interrupt-model) defines the transition from speculative `buffering` work to externally committed work.

The WorkGraph does not weaken that rule.

An effectful WorkNodeAttempt crosses the root Run's commit boundary when its first external side effect becomes observable. From then on:

* the root Run is committed
* an interrupt may not silently restart the entire graph
* graph revisions must preserve the committed effect in history
* compensation/reconciliation is explicit work, never a pretend rollback

Effect classification belongs on nodes and attempts so the scheduler knows before execution which nodes can still be safely discarded.

## Workspace isolation

Parallel graph nodes must not share one mutable working directory.

An `agent_loop` node that needs exec-class tools runs as a child execution unit with its own workspace under [`sandboxed-execution.md`](/runtime/sandboxed-execution). The existing workspace-per-run invariant therefore applies at the child Run level, not as one shared workspace for the entire root WorkGraph.

For coding work, parallel branches may begin from the same project snapshot, but each gets isolated mutable state. A later join/merge node owns reconciliation.

That makes graph fan-out compatible with the sandbox rule rather than an exception to it.

## Budgets compose

A WorkGraph has a root budget and each node has a child budget.

The planner proposes allocation; Mesh enforces it.

At minimum bound:

* max graph nodes
* max graph depth
* max parallel nodes
* max graph revisions
* max replans
* max aggregate tokens
* max aggregate tool calls
* max aggregate wall clock
* max aggregate compute-seconds

Node budgets are slices of the root budget, not independent credit cards.

Unused node budget may return to the root pool. A planner may request reallocation in a graph revision, but it cannot create budget by drawing more nodes.

## Failure propagation

Failure semantics must be explicit on the graph, not inferred ad hoc by workers.

A failed node may cause one of a small set of outcomes:

* retry the node within its attempt limit
* mark dependent nodes blocked
* follow a declared failure edge
* request a graph revision/replan
* fail the root graph

The scheduler owns that transition.

A worker never recursively spawns hidden work because it is unhappy with its result. If more topology is needed, it becomes a graph revision and passes admission again.

## No recursive explosion

A WorkGraph is already a concurrency multiplier. Unbounded nested delegation turns one request into a fork bomb with model tokens.

V1 rules:

* only the root graph planner may propose topology
* worker nodes are loop executors, not autonomous graph schedulers
* workers may request decomposition, but that request returns to the root planner
* every graph revision re-enters admission control
* graph depth, width, revision count, and aggregate budget are hard bounded

Nested subgraphs may be useful later, but they should compile into the same root scheduling/accounting model rather than becoming invisible mini-runtimes.

## Persistence model

The exact schema can evolve, but the conceptual entities should exist before the harness implementation hardens around a linear queue.

### WorkGraph

One root execution plan for a graph-mode Run.

Suggested fields:

* `id`
* `run_id`
* `agent_id`
* `status`
* `current_revision_id`
* aggregate budget fields
* `created_at`
* `updated_at`

### WorkGraphRevision

An immutable admitted or rejected proposal.

Suggested fields:

* `id`
* `work_graph_id`
* `revision_number`
* `proposed_by_run_step_id`
* `status` (`proposed`, `admitted`, `rejected`, `superseded`)
* `rejection_reasons`
* `created_at`

### WorkNode

A stable logical job across graph revisions.

Suggested fields:

* `id`
* `work_graph_id`
* stable `key`
* `kind`
* `objective`
* `effect_class`
* declared input/output contract
* budget
* `created_at`

### WorkEdge

A dependency in one admitted revision.

Suggested fields:

* `revision_id`
* `from_node_id`
* `to_node_id`
* condition/predicate
* artifact contract

### WorkNodeAttempt

One execution attempt for a node.

Suggested fields:

* `id`
* `work_node_id`
* `child_run_id` for `agent_loop` nodes
* `status`
* `started_at`
* `finished_at`
* failure reason
* effect/commit state

### WorkArtifact

A durable typed hand-off.

Suggested fields:

* `id`
* `work_graph_id`
* `producer_node_id`
* `producer_attempt_id`
* artifact type/schema
* inline value or external reference
* provenance metadata
* `created_at`

The tables above are execution/control-plane state owned by one RuntimeAgent. They are not part of the conversation graph and are not cross-perspective retrieval sources.

## How a complex turn should look

A conceptual example:

```text theme={null}
Victor: "Figure out why checkout latency regressed and tell me what we should do."

root Run
  |
  +-> decide execution shape = graph
  |
  +-> planner proposes:

      scope
        |
        +-------------------+-------------------+
        |                   |                   |
        v                   v                   v
     metrics             deploy diff           logs
        |                   |                   |
        v                   v                   v
     verify              verify              verify
        \                   |                  /
         +------------------+-----------------+
                            |
                            v
                         correlate
                            |
                            v
                       recommendation
                            |
                            v
                       final response
```

Every box is bounded. Every dependency is visible. Every artifact is attributable. The scheduler, not the model, decides which admitted nodes are ready to run.

If the `metrics` node reveals a database event nobody anticipated, the planner can propose a revision that adds `database changes` as another branch. The completed nodes remain completed; the plan grows around reality instead of restarting from zero.

## UI and observability

The WorkGraph should eventually become the natural progress view for a long task.

An operator should be able to inspect:

* graph revision history
* ready/running/blocked/completed nodes
* dependency edges
* node attempts
* per-node and aggregate token/compute cost
* artifact hand-offs
* verifier results
* human gates
* failure/replan reasons
* which effect crossed the commit boundary

This is not visualization for its own sake. The graph is already the runtime's execution state; the UI should display that state rather than reconstructing a plausible diagram after the fact.

## Relationship to existing Mesh contracts

* [`addressing.md`](/runtime/addressing) decides whether work should exist at all. WorkGraph planning happens only after a message legitimately creates or joins a turn.
* [`agent-harness.md`](/runtime/agent-harness) owns bounded loops, Runs, RunSteps, progress, and budgets. WorkGraph composes those loops rather than replacing them.
* [`interrupt-model.md`](/runtime/interrupt-model) owns conversation changes while work is running and the external-effect commit boundary.
* [`context-resolution.md`](/runtime/context-resolution) owns which perspective data a root or worker context may read.
* [`sandboxed-execution.md`](/runtime/sandboxed-execution) owns effect isolation and workspaces for child executions.
* [`../perspective.md`](/perspective) owns epistemic isolation. WorkGraph workers are internal execution contexts, never peer RuntimeAgents.
* [`commitments.md` (repository)](https://github.com/TextureHQ/mesh/blob/main/docs/runtime/commitments.md) owns the objective that outlives a turn. A graph-mode Commitment points at a WorkGraph and lets a reconciler drive it across many ticks; this document owns the topology, that one owns the liveness that wakes it up.

## Prior art and what Mesh takes from it

The July 2026 "graph engineering" discussion gave a new label to patterns that predate the label. Mesh should borrow the useful mechanics without coupling itself to one framework.

### LangGraph

LangGraph models workflows with explicit state, nodes, and edges and adds durable execution/checkpointing around long-running stateful agents.

Mesh takes: graph-as-runtime-state, durable node execution, explicit routing, human intervention.

Mesh does not take: a dependency on LangGraph or Python as the orchestration substrate.

* [https://docs.langchain.com/oss/python/langgraph/graph-api](https://docs.langchain.com/oss/python/langgraph/graph-api)
* [https://docs.langchain.com/oss/python/langgraph/overview](https://docs.langchain.com/oss/python/langgraph/overview)

### AutoGen GraphFlow and Microsoft Agent Framework Workflows

GraphFlow represents sequential, parallel, conditional, and looping execution with a directed graph. Microsoft's newer Workflow abstraction broadens nodes beyond agents to executors/functions/subworkflows and emphasizes typed edges and checkpointing.

Mesh takes: nodes need not all be agents; execution topology and message/context flow are separate concerns; typed hand-offs are better than broadcast context.

* [https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/graph-flow.html](https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/graph-flow.html)
* [https://learn.microsoft.com/en-us/agent-framework/workflows/workflows](https://learn.microsoft.com/en-us/agent-framework/workflows/workflows)

### Anthropic multi-agent Research

Anthropic's research system uses an orchestrator-worker architecture in which a lead agent dynamically decomposes a query and launches parallel subagents with bounded objectives. Their published lessons emphasize detailed delegation boundaries and avoiding duplicated worker effort.

Mesh takes: runtime decomposition, parallel specialist contexts, explicit task boundaries, one owner of synthesis.

* [https://www.anthropic.com/engineering/multi-agent-research-system](https://www.anthropic.com/engineering/multi-agent-research-system)

### LLMCompiler

LLMCompiler treats planning like compilation: an LLM produces a dependency graph of function calls, a task-fetching unit dispatches ready work, and an executor runs independent calls in parallel.

Mesh takes: planner versus scheduler separation, dependency-aware parallelism, graph admission before execution.

* [https://proceedings.mlr.press/v235/kim24y.html](https://proceedings.mlr.press/v235/kim24y.html)

### Scaling-agent research

Google Research / DeepMind and collaborators evaluated 180 agent-system configurations and found that topology should follow task structure: parallelizable tasks benefited from centralized multi-agent coordination, while sequential reasoning tasks degraded under multi-agent architectures. Independent systems also amplified errors more strongly than centralized coordination.

Mesh takes: **graph when decomposable; loop when sequential; centralized scheduling and verification rather than "more agents is better."**

* [https://research.google/blog/towards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work/](https://research.google/blog/towards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work/)

### Graph of Thoughts is adjacent, not the same thing

Graph of Thoughts models intermediate LLM reasoning as an arbitrary graph of thought units and dependencies. That is useful prior art for non-linear reasoning, but it is not the same abstraction as Mesh's WorkGraph.

Mesh's WorkGraph is an **execution graph** whose nodes have durable runtime semantics, budgets, permissions, artifacts, and side effects. The model's private thought topology remains private model reasoning unless it is deliberately compiled into executable work.

* [https://arxiv.org/abs/2308.09687](https://arxiv.org/abs/2308.09687)

## Design guardrails

* do not replace one opaque loop with an opaque planner
* do not let a model execute a proposed graph before deterministic admission
* do not graph work that is genuinely sequential merely to create parallelism
* do not model every node as a persistent RuntimeAgent
* do not leak worker execution contexts into the conversation graph
* do not broadcast every node transcript to every other node
* do not allow child workers to gain authority beyond the parent Run
* do not allow arbitrary graph cycles in V1
* do not let retries become hidden cycles
* do not let workers recursively fan out outside the admitted graph
* do not share one mutable workspace across parallel agent-loop nodes
* do not allow graph fan-out to bypass concurrency ceilings
* do not treat model self-review as equivalent to independent verification
* do not let several parallel workers race to own one merge
* do not mutate an admitted revision in place
* do not erase completed history when replanning
* do not pretend a graph revision can undo an external side effect
* do not let drawing more nodes create more budget

## The principle

A good agent harness makes one loop reliable.

A good WorkGraph makes the relationships between reliable loops explicit.

Mesh needs the loop because individual jobs remain uncertain and agentic. Mesh needs the graph because the organization of those jobs should not be uncertain, invisible, or unauditable once the runtime has accepted a plan.

**The model decides what work might be useful. Mesh decides what work is allowed to exist and when it may run.**


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.