Engineering deep dive · August 11, 2026

Organizational memory as a reviewed, auditable graph

How Planet Arkency turns messy conversations into typed knowledge—without giving an LLM direct write access.

Explaining Piotr Jurewicz’s Maintaining an organizational knowledge graph with an LLM and event sourcing.

1. The important invention is the boundary around the model

A transcript is useful evidence, but it is not yet a database. “May is joining Atlas” contains a person, a project, and a relationship. A model can interpret that sentence. It cannot, by interpretation alone, establish that May is the same person as Maya Chen, that the relationship is allowed, or that nothing changed in the database while it was thinking.

Planet Arkency separates these responsibilities. The LLM interprets unstructured input and proposes a structured result. Ordinary application code resolves that result into exact changes, checks it against existing state, and applies it through a separate workflow.

Think of the LLM as a contributor submitting a patch, not a database administrator executing arbitrary commands. A valid patch must fit the vocabulary, refer to the right entities, explain its changes, and still apply to the current state.
1 · IngestOne endpoint receives source content.
→
2 · Read & extractModel searches the graph, then returns nodes and edges.
→
3 · ProposeServer computes before/after differences.
→
4 · Review & applyApply safely, or stop on conflicts.
→
5 · Audit & queryExpose graph, source history, and read history.

Inputs include meeting transcripts, flagged Slack threads, email, RSS, calendar invitations, and notes. Zapier or n8n pushes them to the same ingestion endpoint. This keeps source-specific connectors outside the core knowledge-maintenance machinery.

Why store a graph instead of only a wiki?

In a wiki, “Maya works on Atlas” is prose that a reader must interpret. In a typed graph, it becomes a machine-readable relationship:

Maya Chen [person] ── works_on ──▶ Atlas [project]

You can list everyone working on a project, follow dependencies, or count relationships without reinterpreting every sentence. This does not make a wiki unusable by software; it means a typed graph supplies an explicit contract that prose and ordinary hyperlinks do not.

The implementation uses PostgreSQL: nodes, edges, and JSONB attributes on both. Edges have a unique (source, target, relation) triple. A graph need not begin with a specialized graph database. The author notes that dedicated databases may better suit workloads such as deep multi-hop traversal; the plain storage model also makes adapters easier to replace.

2. Ontology: a controlled vocabulary with relationship rules

For this system, an ontology defines the permitted kinds of entities and the permitted relations between them. It answers not “What facts do we know?” but “What shapes may our facts take?”

Schema / ontology

node_kinds:
  - kind: person
  - kind: project
edge_relations:
  - relation: works_on
    signature: "person --works_on--> project"

Teaching example, simplified from the post’s YAML format.

Instances / graph

person: Maya Chen
aliases: [May]

project: Atlas

Maya Chen --works_on--> Atlas
edge attrs: { since: "2026-08-11" }

The ontology is the grammar; these nodes and this edge are statements in that grammar.

The ontology is closed: a model cannot freely invent another kind or relation. The author tried an open vocabulary and reports rapid chaos. With unrestricted extraction, synonymous labels can split one relationship into works_on, assigned_to, and helps_with; later queries must guess which labels mean the same thing.

The same ontology is rendered into prompt tables for explanation and into the structured-output schema as enumerated choices. These are complementary: the prompt explains semantics, while the schema constrains output shape. Neither proves that a fact is true. The application still needs checks against database state.

Different domains get different graphs

Planet Arkency is multi-tenant. The internal organizational graph uses concepts such as people, projects, and decisions. The Rails Event Store maintainers’ graph uses releases, known problems, and community content. The shared machinery does not force them to share one universal ontology.

This follows domain-driven design’s notion of a bounded context: a boundary within which a vocabulary has a consistent meaning. “Release” in a software-maintenance graph and “client” in an organizational graph need not be squeezed into one generic type system.

Identity resolution: do these names refer to the same thing?

A graph with ten spellings of one person is not ten times richer. It is fragmented. The post uses three defenses:

  1. Look before writing. The extraction model gets read-only tools such as search_nodes and get_node_edges. It must search before proposing a new entity and inspect relationships when context or ambiguity warrants it.
  2. Canonical names plus aliases. A node has one canonical name and multiple surface forms. A rename uses new_name; the old name can remain an alias. An alias denotes the same entity, never a separate person.
  3. Hybrid search. PostgreSQL trigram similarity over names and aliases catches spelling variation. Embedding search catches semantic matches with little lexical overlap. Results are merged and ranked.

The code shown uses a trigram threshold of 0.3, self-hosted bge-m3 embeddings through Ollama, and pgvector cosine nearest-neighbor search with a semantic cutoff. The post does not specify the full merge/ranking policy or provide an identity-resolution benchmark.

A retrieved match is a candidate, not proof of identity. Two people can share a name; a semantic neighbor can be related but distinct. Inspecting connections gives the model more evidence, while recorded tool results let a reviewer see what informed the choice.

3. From extraction to a safe change proposal

The model returns one structured result containing nodes and edges, using RubyLLM’s schema support. A node includes its name, kind, descriptions, attributes, optional aliases, and a new or existing status. Existing nodes must use the exact canonical name returned by the tools. A stable, identity-focused short description supports search; the longer description should synthesize old and new information without discarding prior facts.

Edges name their endpoints, choose an allowed relation, explain why it holds, and carry optional attributes. Crucially, the model does not compute the authoritative database operation.

Let the server determine what actually changed

The server loads or initializes an ActiveRecord object, checks the model’s new/existing claim, assigns the proposed values, and uses dirty tracking to calculate {field: [before, after]}. Attributes are merged with existing attributes. Edges are looked up by their unique triple and diffed similarly.

// Simplified field-level change, computed by the application:
{
  op: "update",
  node_id: "atlas",
  changes: { status: ["planning", "active"] }
}

If the model calls an existing entity “new,” or claims an unrecognized node is “existing,” validation fails. The model receives natural-language feedback and another attempt in the same conversation. This is not a guarantee against every duplicate: an unresolved alias could still be proposed as a genuinely new entity. Search quality remains important.

Review is separate from execution

Extraction produces a proposal, not a write. A Slack summary announces it. A human may inspect the diff and apply it early; otherwise it may apply after a configured delay. Thus the design provides an opportunity for review, not mandatory approval of every fact.

A delay creates a concurrency problem. Suppose a proposal expects Atlas to be planning and would change it to active. Another edit sets it to paused first. Applying the old proposal without a check would overwrite newer state.

Optimistic concurrency, in one line:

Apply only if the current relevant state still matches the proposal’s expected “before” state; otherwise stop and explain the conflict.

The post says applying stops when current state no longer matches the proposal’s basis, and affected rows are marked conflicted. It does not document all transactional, locking, or retry details. The playground below uses an atomic all-or-nothing toy application rule and exact expected-before checks to illustrate the principle.

4. Playground: a model proposes; the application decides

Try alias resolution, a clean application, and a stale proposal. This is a deterministic local simulation, not an LLM or embedding service. It has two starting nodes, one known alias, a tiny closed ontology, and scripted extraction. Nothing is sent over the network.

The relation control matters only for the assignment source. Selecting an invalid relation mimics an unconstrained model output reaching server validation; schema-constrained generation aims to prevent that output earlier.

Latest proposal / exact diff

Recorded read set

Replay the accepted graph through event history

Drag the slider backward. Proposing and reading do not alter the graph. Only accepted writes change it. The panels above always show the latest extraction; this replay affects only the graph and graph-state panel below.

Graph at selected event

Event log / provenance

Three experiments to run
  1. Alias + clean write: Reset → keep assignment and works_on → Extract & propose → Apply. The read set resolves May to Maya Chen; only one person node exists, and a sourced edge appears.
  2. Lost-update prevention: Reset → choose “Atlas is now active” → Extract & propose → competing edit → Apply. The graph remains paused; the proposal expected planning and is conflicted.
  3. Closed vocabulary: Reset → assignment → helps_with → Extract & propose. Validation rejects the relation, leaving the graph untouched. No automatic correction is simulated.

A minimal executable kernel

The live simulation uses this pattern. The excerpt omits UI code, extraction, and event storage; it shows the small consistency check at the heart of safe application. In production, checking and writing must be protected against interleaving, for example by a database transaction and suitable locks or conditional updates.

function applyPatch(state, patch) {
  const conflicts = patch.filter(p =>
    JSON.stringify(state[p.key] ?? null) !==
    JSON.stringify(p.before)
  );
  if (conflicts.length) return { ok: false, conflicts };
  for (const p of patch) state[p.key] = structuredClone(p.after);
  return { ok: true };
}
// Example:
// state = { "atlas.status": "planning" }
// patch = [{ key: "atlas.status",
//            before: "planning", after: "active" }]

The prototype stores full snapshots for cheap browser replay and uses flat keys. Planet Arkency instead uses Rails Event Store, aggregates, relational projections, and real extraction/tool calls. This is a teaching reduction, not its source code or persistence design.

5. Provenance records both writes and reads

Write provenance

Which extraction created or updated this entity? What field changed? Which ingested content supports the change?

node_extractions and edge_extractions hold the extraction/entity pair, operation, status, and field-level diff.

Read provenance

What did the model see before it decided? Which search did it run, and which nodes and edges were returned?

ExtractionToolCalled events project into tool_invocations, linked to the returned graph objects.

Write provenance explains the origin of a change. Read provenance helps diagnose a mistaken merge: perhaps a search returned two similar people and the model chose one. Neither exposes the model’s private internal reasoning; it records the observable evidence and actions around the decision.

The post describes every fact as tracing back to its source. Concretely, the described implementation tracks extraction-linked entity changes and field diffs. That is strong auditability, but it is not an independent proof that every synthesized sentence or attribute is supported. A reviewer still has to compare a claim with its underlying content.

Event sourcing connects the workflow

TranscriptIngested
  → ExtractionRequested
  → KnowledgeExtracted
  → GraphChangeProposed
  → GraphChangeApplied OR GraphChangeConflicted

An event records something that happened. A read model is a convenient view built from those events. Planet Arkency’s UI views—ingestions, extractions, diffs, and tool calls—are read models.

Two small aggregates guard the important invariants. An ingestion aggregate prevents two concurrent extractions of the same content. An extraction aggregate governs the proposal-to-application state machine. The review window becomes a state in that workflow; provenance becomes another projection of events already being recorded.

This is why event sourcing fits the author’s design: the question is not only “What is true now?” but “How did we reach this state, and what happened before an accepted change?” It does not make all engineering complexity disappear. Event versioning, projection consistency, access control, and operational recovery still need design; the post does not evaluate those trade-offs in depth.

6. Costs, research, and access use the same boundaries

Costs are attached to each extraction

The system tracks tokens and price per extraction. Long transcripts and multi-round tool use can be expensive. Prompt caching helps because the system prompt and content stay the same between rounds, allowing much of the input to be billed at a cache-read rate. The author reports that most of their extractions cost well under a dollar, while emphasizing model dependence. This is an operational observation, not a universal price estimate or a latency benchmark.

Research is new input—not privileged knowledge

A user can request research from a node. A model equipped with Anthropic server-side web_search and web_fetch tools creates a structured research brief. Each tool is configured with up to ten uses in the shown code.

The prompt provides the entity’s existing kind, description, and attributes to reduce ambiguity. It asks for factual information, source URLs beside claims, and a completed or aborted status. The model must abort when it cannot confidently identify the exact entity or cannot find substantive, verifiable information, rather than pad the result with facts about a similar topic.

A completed brief is then ingested as ordinary content with its own kind. It goes through extraction, proposal, and review like any other source. That reuse is the key design choice: research does not bypass the write boundary. It also does not eliminate the risk of a mistaken brief; evidence quality still matters.

MCP exposes the graph to assistants

The graph is available through MCP so an assistant with access can search it in natural language and answer with sources. The post describes this capability but does not give the complete MCP tool schema, authorization model, or a downstream agent evaluation.

7. What was demonstrated—and what remains unmeasured

This is an engineering account of a running system, not a controlled research study. Its evidence is architecture, code snippets, UI examples, and reported production scale.

≈2,000nodes
>5,200edges
≈300extractions
>3,600events

Numbers reported for the main production Arkency graph at the time of writing. “Almost 2,000” and “around 300” are approximate, not exact counts.

Claim or designEvidence in the postWhat that does not establish
Closed ontology improves organizationAuthor reports chaos with an open ontology; shows YAML and prompt/schema use.No controlled comparison or measured ontology quality.
Identity resolution has multiple safeguardsLook-before-write instructions, aliases, hybrid-search code, and status validation.No precision/recall, duplicate rate, or ambiguity error rate.
Changes are reviewable and auditableField diffs, provenance joins, recorded reads, and UI examples.No measured reviewer workload or proof that all extracted claims are correct.
Stale proposals do not silently overwrite stateDescribed conflict detection and conflicted rows.No detailed concurrency test suite or complete locking semantics.
Pipeline operates at organizational scaleProduction counts and per-extraction cost observations.No retrieval-quality benchmark, agent-success study, throughput, or end-to-end latency results.

Important boundaries of the design

These are limitations of the claims and description, not assertions that the production system lacks every unreported feature.

8. The lesson to keep

The graph’s value depends less on asking a model to “remember everything” than on making its contribution governable: define the vocabulary, search before creating, preserve identities, compute exact changes outside the model, separate proposal from application, reject stale writes, and retain the evidence trail.

Closing connection: For a context platform, this post offers a concrete pattern for the typed-graph layer: domain-specific ontology plus reviewed, source-linked changes. It is not a complete context-lake architecture, but it shows how ingested material can become durable, queryable knowledge without losing the path back to its sources.