Don’t just retrieve context.
Give it structure, meaning, and a lifecycle.
A first-principles guide to Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence.
Zhishang Xiang and collaborators · Survey / conceptual synthesis
Original paper listing · PDF · Full text read for this reportThe central lesson
A more capable agent is not automatically a better organized system. A large prompt can describe a plan, but it is a poor substitute for a durable record of dependencies, ownership, evidence, and state transitions. The survey argues that these relationships should be explicit graphs that affect execution—not just pictures drawn after execution.
Scope boundary: This is a survey and architectural argument, not a new benchmarked context-platform implementation. The ontology direction is partly a research agenda. The “context lake” architecture and invoice prototype below are this report’s engineering synthesis; they are not a system implemented or evaluated by the authors. The full text returned by the reader uses v2 assets; this report follows that returned text rather than assuming it is identical to the original submission.
1. Start with four different things
Imagine an assistant asked to pay an invoice. It must read the invoice, find the purchase order, check the applicable policy, reconcile quantities and prices, obtain approval, and finally request payment. “Here are all the documents” does not tell the system which facts are current, what approval means, or whether a payment already happened.
Ontology: meaning
The shared vocabulary and constraints. An Invoice is not a Payment. An amount has a currency and unit. A reviewer must be authorized.
Context lake: evidence
Durable source material: documents, observations, tool results, and events, with identity, versions, timestamps, and access rules.
Context graph: relationships
Queryable links among entities, claims, sources, tasks, agents, and state. A claim can point back to the exact evidence that supports it.
The fourth thing is the runtime: code that schedules work, checks permissions, validates proposed changes, and commits state. An ontology saying “payment requires approval” does not stop an API call unless the runtime enforces that requirement.
Graph from first principles
A graph consists of nodes and edges. The important question is not “Do we have a graph database?” but “What does each edge mean, and what behavior does it permit?”
Invoice I-17 ──supportedBy──► Source invoice-v1.pdf Task Reconcile ──requires──► Task Extract Task Review ──assignedTo──► Agent Reviewer Evidence E-3 ──validFor──► Invoice revision 1 Payment Intent P-1 ──requiresApproval──► Review R-1
The first link helps retrieve evidence. The second constrains scheduling. The third assigns responsibility. The fourth prevents stale evidence from validating a changed invoice. The fifth gates an action. Treating them all as generic “related to” edges throws away their operational meaning.
2. What the survey contributes
From a model call to an organized system
The paper distinguishes three levels. Model intelligence is capability expressed in a bounded interaction. Training provides capabilities; prompts specify behavior and context supplies relevant information. Individual intelligence adds persistent resources and feedback-controlled action. System intelligence adds organization across interdependent tasks, heterogeneous components, and shared state.
Agent = Loop(Foundation Model + Harness)
System = (Agent Team, Shared Resources, Environment,
Coordination Mechanisms, System State)The harness supplies tools, memory, skills, interfaces, and controls. The loop decides what to do next after observing results, including when to stop. Neither automatically provides a coherent division of work across several agents.
The authors identify three recurring single-loop bottlenecks: independent work becomes serial; writing and verification become entangled; and errors contaminate a long conversational history without a clean recovery boundary. These motivate an architectural change, not simply a larger context window.
Task organization
Decompose objectives into subgoals, expose dependencies, compile them into executable operators, and adapt workflows to feedback.
Agent coordination
Model capabilities, assign ownership, define team topology, and route useful communication instead of connecting everyone to everyone.
Runtime state
Record validated transitions, diagnose failures, and recover affected regions without discarding all successful work.
Three distinctions worth preserving
- Subgoals versus operators. “Check the invoice” is a semantic goal. Parsing a PDF, querying a purchase order, and running a deterministic comparison are executable operations.
- Team organization versus communication. A reviewer’s responsibility can remain stable while messages and evidence requests change each round. Sending another message need not redefine ownership.
- Adaptation versus evolution. Retrying one task changes a run. Persisting a validated new workflow for future runs changes the system. The paper emphasizes that durable cross-run structural evolution remains uncommon.
Representative works illustrate the taxonomy: LLMCompiler exposes schedulable dataflow dependencies; GPTSwarm optimizes computational graphs; LangGraph supplies explicit graph-oriented workflow and state abstractions; Graphiti represents temporal contextual knowledge. These are examples discussed in the survey, not interchangeable implementations of its entire vision.
Why the ontology section matters
Two agents can agree on a graph’s topology and disagree on its semantics. One calls an invoice “complete” when parsing ends; another interprets that as “approved for payment.” The paper proposes layered, modular ontologies: a shared core plus modules for goals, capabilities, evidence, policies, actions, states, and domain concepts.
Ontology connects graph views by standardizing what their entities and relations mean. It should be grounded in observations and governed through provenance, validation, versioning, migration, and rollback. It does not choose society’s values, replace factual verification, or solve all system-level evaluation problems.
3. Turn the framework into a context platform
Engineering synthesis: The following architecture translates the survey’s requirements into a buildable design. “Context lake” here means a governed, durable evidence substrate; the paper does not prescribe that storage layer or name this design.
Original explanatory diagram. The feedback arrow records observations; it does not grant agents permission to overwrite authoritative facts.
A minimal data contract
| Object | Fields to start with | Purpose |
|---|---|---|
| Source | ID, immutable version/hash, URI, ingestion time, observed time, ACL | Keep exact evidence addressable and governed. |
| Claim | Subject, predicate, value, units, valid interval, recorded time, source span, status | Separate an assertion from its support and its acceptance state. |
| Task | ID, requirements, dependencies, assignee, state, acceptance criteria | Make readiness and ownership explicit. |
| Event | Run ID, actor, prior version, patch, validation result, event ID | Make transitions auditable and reconstructable. |
| Action intent | Target, parameters, authorization, approved evidence version, idempotency key | Prevent an evidence query from becoming an unchecked side effect. |
Time deserves two fields: when the fact applies and when the platform learned it. An invoice correction learned today may apply to yesterday. Keeping both supports historical questions and prevents “latest ingested” from masquerading as “currently valid.”
Write path: propose → validate → commit
- Ingest original material with stable IDs and source permissions. Preserve raw evidence rather than replacing it with a summary.
- Extract candidate entities and claims. Resolve identity carefully; a vendor’s name alone is not a unique key.
- Validate types, units, source references, role permissions, freshness, and graph invariants. Use OWL reasoning for applicable logical consistency checks and SHACL or code for concrete data-shape constraints; neither proves real-world truth.
- Commit accepted patches against an expected state version. On a concurrent conflict, reject or reconcile explicitly instead of allowing silent lost updates.
- Append an event and update graph projections. Rejected proposals remain diagnostic records, not accepted facts.
Read path: build a purpose-specific context view
Start from a task or entity, traverse relevant typed links, filter for authorization and validity, rank the remaining evidence, and package a bounded view with citations. The invoice extractor needs the source invoice; the reviewer needs the invoice, purchase order, policy, reconciliation result, and evidence versions. Neither needs every conversation in the organization.
Authorize before retrieval expansion and again before delivery. Otherwise, summaries, graph paths, and even existence of hidden nodes may leak information. Permission-scoped views should apply to edges and derived claims as well as source documents.
Recovery is not time travel for the outside world
If extraction is wrong, invalidate the affected descendants and retain the independent policy lookup. If a bank transfer was already sent, resetting a task flag cannot undo it. Recovery may require reconciliation, compensation, or human intervention. Use idempotency keys and explicit settlement records at the external boundary.
4. Playground: an invoice workflow with semantic gates
This local deterministic simulation contains no LLM, network request, bank connection, or persistent storage. It illustrates readiness, parallelism, validation, evidence grounding, and selective repair. Each simulated task takes one tick; the numbers are teaching parameters, not paper results.
Ingest ──► Extract ──┐
└─► Policy ───┴─► Reconcile ──► Review ──► Payment intent
compare independent no real paymentSettings are captured at reset. Repair disables the injected error for this run and preserves unrelated completed work. Review independently compares the original invoice with the parsed claim even if extraction validation is disabled.
Scoped context view
Append-only teaching event log
- Run the defaults: unit validation rejects extraction, while Policy still completes.
- Repair and run again: Extract and its descendants resume; Policy is reused.
- Disable validation, keep the error, and reset: the claim travels farther, but independent Review rejects it.
- Disable the error; compare one versus two workers after reset. The idealized workflow finishes in six versus five ticks.
- Switch context roles: the progress observer sees task state, not amounts or documents.
A small scheduling prototype
The live code uses this basic rule. A task may start only after all prerequisite tasks have committed accepted results. A failed task is not equivalent to a completed dependency.
const ready = tasks.filter(t =>
t.state === "pending" &&
t.dependencies.every(id => state[id] === "done"));
for (const task of ready.slice(0, availableWorkers)) {
const proposal = execute(task); // not authoritative yet
const error = validate(proposal); // schema + evidence + role
if (error) recordFailure(task, error);
else commitResultAndEvent(task, proposal);
}
// Repair traverses dependency descendants, not every task.A production runtime additionally needs atomic commits, concurrency control, durable checkpoints, evidence-version binding, permission enforcement at every boundary, and safe handling of external effects. The role picker models a policy-scoped projection, not a security boundary in this client-side file.
5. Building incrementally, without a giant ontology
- Choose one bounded workflow. Start with invoice reconciliation or incident investigation, not “all enterprise context.” Define success and forbidden actions.
- Preserve sources and lineage first. If a claim cannot link to original evidence, a larger graph only scales uncertainty.
- Define a small semantic core. Source, Claim, Entity, Task, Agent, Evidence, Policy, and ActionIntent are enough to expose many mistakes. Use explicit units and evidence status.
- Add an operational task graph. Schedule deterministic operations alongside model calls. Keep verification separate from generation.
- Add state transitions and repair. Store events, state versions, approval gates, and descendant invalidation.
- Only then evolve structure. Propose a changed workflow, replay or test it under controlled conditions, evaluate transfer, and commit or roll it back.
An initial implementation can use an object store for source bytes, relational tables for claims and events, and adjacency tables for graph links. A dedicated graph store is useful when traversal needs justify it, but it is not the definition of graph engineering.
6. How to tell whether the platform helps
The survey argues that final task success is insufficient. Better results may come from a stronger model, more samples, or a larger budget rather than better organization. Compare systems with matched models, tool access, retry limits, and compute budgets, and report both effectiveness and overhead.
| Question | Practical measurement | Useful intervention |
|---|---|---|
| Does context stay grounded? | Fraction of accepted claims with resolvable evidence; source-check accuracy | Insert contradictory and stale source revisions. |
| Does structure help? | Dependency violations, completion latency, communication cost | Compare serial, fixed graph, and adaptive graph under matched budgets. |
| Can the system recover? | Valid work retained, recomputation, recovery time, repeated external effects | Fail one branch or revoke a required tool. |
| Are semantics consistent? | Unit/type violations caught; approval gates correctly enforced | Change units or overload the meaning of “done.” |
| Is access governed? | Unauthorized source and derived-claim exposure | Revoke access and test retrieval, summaries, and cached projections. |
| Does evolution transfer? | Held-out cross-run gain and regression rate per graph version | Evaluate a retained structural change on different tasks. |
What this paper does—and does not—establish
Its strength is a unifying taxonomy: work, actors, state, and shared semantics belong in one architectural conversation. It surveys benchmarks, frameworks, and applications across software, science, healthcare, enterprise automation, personal agents, and simulation.
It does not establish a universal numerical benefit for graph engineering or specify a finished interoperable platform. Its own account says existing systems are fragmented and persistent structural evolution is rare. Several connections are future directions rather than validated production recipes.
Be especially skeptical of two shortcuts: “more agents means more intelligence” and “ontology consistency means factual correctness.” Additional agents can amplify shared errors and increase coordination cost. Formal consistency only checks selected logical commitments; it cannot guarantee that evidence is accurate or an action is safe.
Questions to carry into a design review
- Which graph links change behavior rather than merely decorating retrieval?
- What exact evidence makes a task complete, and for which source version?
- Who may propose a fact, who may accept it, and who may execute an action?
- Which state can be replayed, and which external effects require compensation?
- Which structural improvement survives this run, and what test justifies retaining it?
Takeaway
Build a context platform as a governed evidence-and-execution system: the lake preserves observations, the ontology defines meaning, the graph organizes relationships, and the runtime enforces valid change.
The paper’s most useful shift is from asking “What should I put in the prompt?” to asking “What must the system know, preserve, verify, and coordinate so that the next action is justified?”
Primary source: Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence, especially §§2–6 for concepts and architecture, §§7–9 for evaluation, ecosystem, and applications, and Appendix §11.2 for the distinction from graph-enhanced agent capabilities. Authors’ resource collection.