Skip to main content

Organizations

Mesh knows who said something. It does not know who they work for. Actors are first-class (perspective.md). Conversations are a graph. Memory is versioned and audience-aware. But the entity every one of those things is about in practice — the company, the utility, the vendor, the counterparty — exists nowhere in the schema. Grep the repository: no table, no Go type, no column. An agent deployed inside Texture has no record that Texture is a thing, let alone that Acme Corp is a customer it first spoke to three months ago. This document proposes organizations as a distinct, versioned, perspective-owned knowledge class built on the memory machinery rather than beside it. Status: plan, and the most speculative of the current specs. The phasing section marks what should be built first and what should be deferred until the first phase has taught us something.

Why it is not just memory with a tag

The honest counter-proposal is: memory entries already have subjects, provenance, audience evidence, and revisions — add subject_kind = 'organization' and stop. That gets 60% of the value for 5% of the work, and it is the right first move (see phasing). What it does not get is the thing that makes organizations interesting:
  • Identity resolution. “Texture,” “Texture HQ,” “texturehq.com,” and “the utility platform folks” are one entity. Free-text memory has no mechanism to converge on that, so an agent accumulates four half-populated shadow records and answers inconsistently depending on which one retrieval surfaces.
  • Structure worth querying. “Which customers have we not spoken to in 90 days” is a query over rows. It is not a similarity search over prose.
  • Aggregation across actors. An organization’s knowledge is assembled from many actors’ statements across many conversations, each with its own audience. A single free-text entry has one provenance chain; an organization needs many, per attribute.
  • Relationship edges. actor → organization, organization → organization (vendor, customer, parent), organization → conversation.
So: reuse the mechanics of memory — append-only revisions, evidence capture, audience predicates, perspective ownership — and add an entity with resolution and structure on top. Do not build a second memory system. Do not pretend a graph entity is a sentence.

Non-negotiables inherited from the rest of Mesh

These are not new rules. They are the existing rules, and they bind this feature harder than they bind most.
  1. Perspective ownership. runtime_agent_id NOT NULL, composite foreign keys, no deployment-global organization table. If Daedalus learns Texture was founded in 2023, Lyra does not thereby know it. A shared organization directory would be the single largest breach of the boundary ever proposed for Mesh, and it would be convenient, which is exactly why it is called out here.
  2. Audience-aware disclosure, per attribute. An attribute learned in a private DM carries that DM’s audience. Retrieval into a public channel must omit it. This means audience lives on the attribute assertion, not on the organization — an organization’s existence may be public while its revenue figure is not.
  3. Capture precedes behavior. Every attribute records who said it, in which message, when, and under what audience. An attribute with no provenance is a rumor, and unknown provenance fails closed.
  4. Actors stay opaque. Mesh does not classify an actor as human, agent, or service (Actor.kind was dropped in migration 0014). “Works at Acme” is an edge with provenance, not a type. Do not reintroduce classification through this door.

Shape

Why every table carries the perspective

runtime_agent_id is repeated on each child table and is part of every foreign key, rather than being reached by joining back to organizations. This is not denormalization for convenience; it is the only shape in which the perspective boundary is a constraint instead of a habit. Migration 0016_perspective_ownership.sql added actors_perspective_id_key UNIQUE (runtime_agent_id, id) for exactly this reason, and 0018_actor_attributed_memory.sql spends its header explaining the consequence: attributing one agent’s memory to another agent’s Actor is rejected by the database, not caught in review. organizations needs the same composite key so organization tables can reference it the same way, and organization_attributes needs it on source_actor_id so an attribute about Acme cannot cite an Actor from another perspective as its source. conversations and message_events deliberately do not carry an ownership column (data-model.md) because conversations.connector_id is NOT NULL and makes their perspective total one join away. Organization tables are not in that position: an organization_attributes row is reachable directly from retrieval and carries its own audience decision, so the perspective must be on the row being filtered. Two ordering notes for whoever writes the migration. MATCH SIMPLE (the default) is what lets the one intentionally nullable reference work — when actor_id is NULL the constraint is satisfied without consulting actors, which is the “affiliation claimed but identity unresolved” state the Jim section below depends on. And ON DELETE RESTRICT rather than SET NULL, because SET NULL on a composite FK nulls every column in it, including the NOT NULL runtime_agent_id.

Nullability is a per-column decision, not a table-wide one

MATCH SIMPLE cuts both ways, and the cut is easy to miss because the two cases sit in the same constraint. A composite foreign key whose referencing columns are partly NULL is satisfied without checking the parent table at all. That is precisely the behavior organization_actors.actor_id wants. It is precisely the behavior every organization reference does not:
  • organization_id on organization_revisions, organization_aliases, organization_attributes, and organization_actors is NOT NULL, and so are from_organization_id and to_organization_id on organization_relations. Without it, a row carrying a perspective and a null organization is storable, and the composite FK will not object. That row is an orphan fact: it can never be retrieved (retrieval reaches attributes through an organization), never merged (the operator merge path keys on organization), and never explained. Absent data is recoverable; unattributable data that looks like knowledge is the failure mode this whole document is arranged against.
  • actor_id on organization_actors stays nullable, deliberately and alone. Nulling it is the representation of the unresolved-affiliation state, not an accident of not knowing yet.
So the rule for the migration author: an organization child row must reference a concrete organization, and the only nullable reference in these tables is the one whose null is load-bearing. A CHECK (actor_id IS NOT NULL OR claimed_name IS NOT NULL) on organization_actors is worth adding alongside, so the nullable case still has to say who it means. Attributes are append-only with supersession rather than updated in place, so “founded in 2023” asserted today and corrected next month leaves both rows and a readable history. This is the remember / amend split expressed as data: a writer that cannot overwrite cannot silently rewrite what an agent believed. Evidence columns deliberately mirror memory_revision_evidence so the existing mesh_audience_allows predicate applies unchanged, and an attribute-level mesh_organization_attribute_allowed(...) can be written as a near-copy of mesh_memory_base_allowed. Reusing the predicate is the point; a parallel hand-rolled check would drift.

The self organization

One row with is_self = true per perspective, created at provisioning, not inferred, and held to one row by the partial unique index above rather than by this sentence. The agent’s own organization is a configured fact — it comes from the operator during agent setup, the same way the persona does (agent-provisioning.md) — because guessing your own employer from chat is both unnecessary and embarrassing when wrong. It also anchors a useful default: a workspace-audience attribute about the self organization is broadly usable inside that workspace, while the same attribute about a counterparty organization usually is not.

Learning: the part that needs discipline

The tempting design is “run an extraction pass over every message and update the graph.” That produces a confidently wrong knowledge base within a week, because chat is full of hypotheticals, jokes, quoted third parties, and plans that never happened. “We should tell Acme we founded in 2023” is not an assertion that either fact is true. Constraints that keep extraction honest:
  • Extraction is a proposal, never a commit. It produces candidate assertions with confidence and a source span. The runtime governs what is retained — models propose, the runtime governs (first principle 9).
  • Only from messages the perspective legitimately holds, with the inbound message’s audience attached at extraction time. The audience of the assertion is the audience of the message it came from, never widened by aggregation.
  • Attribution is derived, never model-supplied. Exactly as the memory verbs refuse any parameter naming a source, extraction takes the speaker from the persisted graph row. A model must have no channel for “Bob told me X.”
  • Speaker authority is recorded, not assumed. Someone at Acme describing Acme is a strong source; a third party describing Acme is weaker. Store the distinction as the actor’s affiliation at assertion time and let retrieval weigh it, rather than encoding a trust hierarchy in the writer.
  • Contradiction does not auto-resolve. Two conflicting attributes coexist, both visible, newest first. Silent auto-resolution loses exactly the information a human needs to fix it.
  • Cost is real. An extraction model call per inbound message is a per-message tax on a system that deliberately keeps observe cheap (addressing.md). Follow the same cost-ordered ladder: a deterministic cheap gate (does this message mention a known alias or an organization-shaped pattern?) before any inference, and consider extracting on engaged turns only in phase one.

Explicit verbs beat ambient inference

Before any automatic extraction ships, give the agent verbs and let the existing tool audit trail carry the weight:
These are auditable, cheap, and immediately useful, and they generate the corpus that tells us whether automatic extraction is worth its cost. Shipping verbs first is not timidity; it is how the extraction prompt gets written from evidence rather than from imagination.

Jim the IT guy

Victor’s example is the sharpest test in the whole design: someone says “I have to check with Jim, he’s our IT guy.” What should Mesh retain? Do not create an Actor. An Actor in Mesh is an identity with connector identities, addressability, and a memory subject. Creating one from a third-party mention manufactures an identity from hearsay, and the damage is concrete: the next time a real Jim appears on Slack, the runtime faces a phantom to disambiguate against. Identity linking requires a live possession proof: a code echoed from one private session into another. A name claim cannot select an account or link anything. A fabricated Actor would still misrepresent hearsay as an observed, addressable person. Do record the claim, as an organization_actors row with actor_id = NULL, claimed_name = 'Jim', role = 'IT', verified = false, and full provenance. That is genuinely useful — the agent can say “you mentioned Jim handles IT” — and it is honest about what it is: something a person said once. Promotion is a separate, verified step. When someone identifiable as Jim later appears on a connector, a claim can be linked to that Actor through the existing linking protocol. The unverified claim becomes evidence for a link; it never becomes the link. That is the same rule as identity linking, applied to a weaker input, and it is why the verified flag is in the schema from day one rather than added later.

Retrieval

Organizations enter prompt context the way memory does, through context-resolution.md, and they are subject to the same bounded allowance in prompt-budget.md (repository).
  • Which organizations are relevant comes from the participating actors’ affiliations and from aliases mentioned in the trigger message.
  • Which attributes are eligible is decided in SQL by the audience predicate per attribute, before ranking. Same rule as memory and as conversation-search.md: hard eligibility precedes relevance.
  • How much is capped. An organization with 400 attributes must not evict history. Render a small, ranked digest — self organization first, then organizations of the actors present — not the record.
  • Provenance flows through. An organization attribute used in a reply is recorded like a memory revision id, so the disclosure lineage in migration 0063 can walk it later. This likely needs an organization_attribute_ids array alongside memory_revision_ids in MessageProvenance, and a corresponding clause in mesh_messages_allowed. That is a schema and predicate change, not an afterthought, and it is the main reason phase one should stay small.

Costs and risks, stated plainly

  • A new entity class is a permanent maintenance surface: migrations, retrieval, budget, operator UI, export. Memory-with-a-tag costs almost nothing by comparison.
  • Duplicate organizations are the predictable failure. Alias resolution will be wrong sometimes; there must be an operator merge path, and merge must preserve both provenance chains.
  • Wrong facts are stickier than absent ones. An agent that says “you founded Texture in 2022” with confidence is worse than one that says nothing. Hence confidence, provenance, coexisting contradictions, and no auto-resolution.
  • Perspective pressure will be constant. Every deployment will eventually ask for a shared org directory across agents. The answer is no; the compliant shape is an operator-provisioned self organization per perspective, with per-agent learning on top.
  • Privacy surface grows. Organization records aggregate what many actors said across many rooms. Per-attribute audience is what keeps aggregation from becoming leakage, and it must never be relaxed to per-organization audience for convenience.

Phasing

  1. Self organization only. Operator-provisioned at agent setup, one row, rendered in context. No extraction, no learning. This alone fixes “the agent does not know what company it works for.”
  2. Organizations, aliases, attributes, and the three explicit verbs, with per-attribute audience evidence reusing the existing predicates and full provenance into MessageProvenance. Model-initiated, auditable, no ambient inference.
  3. Affiliation edges, including unverified claimed_name rows, plus promotion through verified identity linking.
  4. Candidate extraction, gated behind the cheap deterministic tier, writing proposals with confidence — evaluated against the corpus phase 2 produced.
  5. Relationship graph and structural queries (“customers we have not contacted in 90 days”), once there is enough real data to know which queries people actually ask.
Phase 1 is small and obviously correct. Phase 2 is where the design earns its keep. Nothing past phase 3 should be committed to before phase 2 has run in a real deployment.

Open questions

  • Is is_self singular? Settled: singular, enforced by the partial unique index in the shape above. A consultancy or MSP agent expresses the rest through organization_relations and through per-client affiliation, not through several self organizations. The cost is real and worth naming: an agent that genuinely belongs to two organizations has to pick one as self, and the retrieval default (“self organization first”) will favor it. That is recoverable — dropping a partial unique index is a trivial migration, while retrofitting a uniqueness rule onto a table that already holds duplicates is not. Constrain first, relax on evidence.
  • Does an organization deserve its own retrieval audit table, or does it extend memory_retrieval_events? Leaning: extend, so one query answers “why did you say that.”
  • How does an organization record interact with versioned-state.md’s export view? An org digest is a plausible markdown export; it must stay a view, never a source of truth.
  • Should attribute keys be a controlled vocabulary or free text? Free text learns faster and queries worse. Probably: free text with a small promoted set (founded, industry, relationship, first_contact_at).