Paper explained · Knowledge graph systems

Quipu: make the store hard to convince

A governed knowledge graph should remember not only what it knows, but who was allowed to write it, which rules applied, and why a write was refused.

Reading: Steve Brown, Quipu: A Governed Bitemporal Knowledge Graph Store · Read version dated August 24, 2026 · Original paper · PDF

01 / THE PROBLEM

Well-formed is not the same as trustworthy

A knowledge graph is a collection of facts connected by explicit relations. A tiny graph might contain household:h3 — placedIn → district:south. This is easy for software to follow. It is also easy for an agent to invent, misplace, duplicate, or write without authority.

Traditional workflows can accept a fact and review it later. But between acceptance and review, another process can use it to make a decision. Agent writers increase both the volume of writes and the speed at which plausible errors spread. Quipu moves governance into the store’s own write path rather than asking every client to remember the same instructions.

The thesis: start strict; let agents bear the cost of retrying. A refusal is structured feedback, not just a failed operation. But acceptance means “satisfies the checks we authored,” not “is true in the world.”

Four defaults the paper challenges

  1. Accept, then clean. Invalid facts are visible before anyone repairs them.
  2. One clock, or no clock. Historical data may survive while historical rules disappear.
  3. Flat trust. Combining graphs can make weak evidence look authoritative.
  4. Governance outside. Policy in a prompt or dashboard can drift away from what the store enforces.

Four corresponding inversions

  1. Gate writes against the state they would create.
  2. Make governance bitemporal along with ordinary facts.
  3. Partition authority and trust; never strengthen labels through composition.
  4. Store policies, traces, and verdicts as facts, so an audit can query them.

The novelty claim is the combination. Immutable fact logs, temporal databases, named graphs, validation, and security lattices all have predecessors. Quipu argues that these mechanisms reinforce one another when implemented together.

02 / FIRST PRINCIPLES

Validate the world after the proposed change

Suppose the current graph places household h1 in North. An agent proposes another placement in South. The current graph is valid. The proposed fact, viewed alone, is also valid. Their combination is not, if a household may have only one placement.

candidate state = current state + proposed delta
allow only if applicable constraints hold on candidate state

This is post-state validation. “Post” here means the state that would exist, not a check performed after making a bad write permanent. Quipu stages the delta inside a SQLite savepoint, evaluates it, and commits or rolls it back.

1. AttributeActor, delegation chain, target graph
→
2. Check authorityIntersect grants along the chain
→
3. Stage & checkPolicy claims and shapes see pending post-state
→
4. ResolveCommit allowed data; roll back denied delta
→
5. Preserve verdictFlush attributed, signed decision after resolution

Two kinds of checks

Shapes constrain structure: a record must have a provenance property, for example. Quipu uses SHACL, a validation language for graph data.

Policy claims constrain acceptable states. They are stored as policy entities with SPARQL ASK queries—queries that return a yes/no answer. A closed-world vocabulary policy can reject a fabricated predicate even when a shape would tolerate extra properties.

Two important boundaries

A policy’s graph scope is part of its meaning. If records live in named graphs but a query looks only at the default graph, it may evaluate an empty view rather than the intended evidence.

Policies are indexed by target type. A write that touches no governed type runs zero policy claims. It still pays authority-checking costs: “no policy work” does not mean “no governance overhead.”

Why a refusal needs its own persistence path

If the verdict were part of the same transaction as the candidate facts, rolling back a denied write would erase the reason for denying it. Quipu stages the verdict separately and flushes it after the savepoint resolves. It records policy, target, outcome, actor, principal chain, evidence hash, and an Ed25519 signature. Attribution is inside the signed evidence hash, so swapping the writer invalidates verification.

The signing identity must be registered against a human-authored root of trust. The store cannot appoint itself trustworthy. Without a signing identity, it records no verdict rather than an unsigned one. Thus the operational signing setup is a prerequisite for the intended evidence trail.

Approval is a retry protocol, not a waiting transaction

A require-approval refusal leaves a decision request with an evidence hash and expiry. A human approval is bound to that same evidence. The agent retries; the next attempt can succeed. Rejection outranks approval, and an expired or zero approval window refuses. The store does not hold a write transaction open while waiting for a human.

03 / HANDS-ON

Playground: a strict gate is only as complete as its rules

This teaching model uses household placements. The graph starts with h1 in North and h3 in South. A recorder is authorized only for the North graph. Try a conflicting placement, a missing source, an invented relation, and a false but structurally valid record.

The optional registry is an illustrative additional rule, not a claim that Quipu’s original scenario had this oracle.

Ready. No write attempted.

Committed facts

Decision trail

Simulation only: in-memory records, not cryptographic signatures, durable storage, SPARQL, or the actual Quipu engine. A refused candidate is not added to committed facts.

Try this: load the conflict and submit. Authority passes because the North graph is authorized, but h3 would have two placements. Disable the single-placement check and submit again. The false North claim can now land. Reset, enable the trusted-registry rule, and repeat: the missing semantic check now catches it.

This reproduces the kind of loophole found in the paper’s agent experiment, not its exact data fixture: one agent refiled a household into a district it could write, and the false record passed because no policy stated that household’s actual district. Formal validity is not an independent truth oracle.

A small executable core behind the playground
function evaluateWrite(facts, w, cfg) {
  const reasons = [];
  if (w.graph !== 'north') reasons.push('No authority');
  if (!w.source.trim()) reasons.push('Missing provenance');
  if (w.predicate !== 'placedIn') reasons.push('Undeclared relation');
  const duplicate = facts.some(f =>
    f.house === w.house && f.graph === w.graph &&
    f.predicate === w.predicate);
  const post = duplicate ? facts : [...facts, w];
  const places = new Set(post.filter(f =>
    f.house === w.house && f.predicate === 'placedIn')
    .map(f => f.graph));
  if (cfg.single && places.size > 1)
    reasons.push('Post-state has two placements');
  if (cfg.world && registry[w.house] !== w.graph)
    reasons.push('Trusted registry disagrees');
  return {allow: reasons.length === 0, reasons, post};
}
// Resolve data first; retain a separate decision record either way.
const r = evaluateWrite(facts, write, config);
if (r.allow) facts = r.post;
verdicts.push({actor: 'recorder', outcome: r.allow ? 'allow' : 'deny',
               reasons: r.reasons});

This strips away persistence and temporal semantics to isolate candidate-state validation. Production code also needs transactional safety, authority chains, signed evidence, temporal query semantics, and policy-definition validation.

04 / THE TWO CLOCKS

When was it true? When did we learn it?

A single timestamp cannot answer both questions. Valid time says when a fact holds in the world. Transaction time says when the store learned or recorded it.

Imagine a recorder learns on day 5 that a household moved on day 3. A query made using knowledge available on day 4 should not see that update—even if its effective date is day 3. A retrospective query using knowledge available on day 6 should see it.

fact = (entity, attribute, value, graph, transaction,
    valid_from, valid_to, operation)

Quipu’s append-only log supports assert, retract, and tombstone. Retraction logically closes a valid interval without deleting history. A tombstone hides a triple in a composed view without modifying the underlying layer.

Day 1North placement recorded, effective day 1.
Day 3Household moves to South.
Day 5Store learns the move, effective day 3.
Day 7Rule changes from “source optional” to “source required.”

A two-clock query

Teaching fixture: the move is learned on day 5 and applies from day 3. Rule v2 is recorded and effective on day 7. Production rules can have distinct valid and transaction dates too.

Historical decision

A hypothetical source-less write was allowed on day 2 under rule v1.

The important extension

Do not preserve only old data. Preserve old labels, authority grants, verdicts, policies, and shapes. Otherwise you can reconstruct what was known but not what was required.

The paper’s amendment test makes this difference concrete: all 50 accepted verdicts re-derived under their historical facts and claims; all 50 became unsatisfied under the amended claim. Evaluating yesterday’s decision against today’s rule would falsely diagnose a historical error.

Denied writes are asymmetric. Quipu keeps the denial verdict but discards the staged delta. For those cases, historical replay verifies the rules in force, not the denied outcome from scratch. Full re-derivation requires a writer-side trace containing the attempted delta.

How the layers fit together

Named graphs partition facts. Committed graphs are self-rooted; an overlay binds once to a committed parent and cannot later claim a different base. A view resolves nearer overlays first, including tombstones. A dataset, in contrast, selects an arbitrary set of graphs. The parent tree and dataset composition are different structures.

For cross-store composition, Quipu allocates distinct identifier spaces and mounts attachments read-only, verifying rather than migrating them. Knowledge packs transfer facts by re-interning terms, not copying numeric rows. Pack identity is a hash over sorted N-Triples, so local identifier assignments do not change the content identity.

05 / COMPOSITION WITHOUT TRUST LAUNDERING

A union of graphs must not upgrade its ingredients

Consider an attested district graph and a quarantined extraction graph. A query can join them, but the resulting view must not silently inherit the attested graph’s prestige. Quipu attaches labels to partitions and conservatively composes those labels.

Meet: choose no stronger than the weakest

Trust and freshness fold by meet. For a simple, shared rank chain, the minimum is the weakest admissible rank. Combining “attested” with “quarantined” cannot yield “attested.”

Trust from different named chains is not compared numerically. Quipu refuses the comparison and names the incompatible chains.

Join: retain every obligation

Policy obligations fold by join. If one member says “no export,” the composed set inherits that obligation.

Labels also carry durability. The key reported composition rules are conservative trust/freshness folds and accumulating obligations; the playground below models those, not the full four-axis algebra.

Unknown is not just the lowest trust rank

An undeclared label is missing information, not a declared “untrusted” value. Quipu therefore returns a pair: folded label and coverage. Coverage is Empty, None, Partial, or Full. Empty means no members; None means members exist but no declarations. Partial coverage fails enforcement floors even if every declared member looks strong. An expired label becomes absent, degrading coverage—it does not become a false label.

Trust composition sandbox

Graph A is fixed: attested, fresh, on the “civic” trust chain. Configure Graph B. The read floor requires full declarations and at least curated trust.

Authority is related, but different. A writer’s grants specify which partitions it may write. Along delegation, effective authority is the intersection of every principal’s grants. A wildcard is the identity; an empty intersection refuses. Relabeling a graph requires authority over the reserved meta-graph, preventing a tenant from upgrading itself.

effective authority = grant₁ ∩ grant₂ ∩ … ∩ grantₙ
{North, South} ∩ {North} = {North}

Labels are not row-level access control. A label floor can refuse a composed query, but it is not a security system that selectively hides rows.

06 / GOVERNANCE AS DATA

An audit asks what the evidence establishes

Quipu stores the specification Σ, the trace T, and verdicts in the graph being governed. The expression T ⊨ Σ means “the trace satisfies the specification.” Its checker does not consult an LLM or reverse-engineer a prompt.

Four audit passes

  1. Coverage: are relevant actions and constraints represented?
  2. Class-placement: was each constraint checked at the right enforcement point?
  3. Outcome consistency: does the observed response agree with the declared constraint?
  4. Attribution: can the action and decision be assigned to an actor and chain?

Three outcomes must not collapse

Allow, deny, unknown are governance outcomes. Separately, an audit distinguishes violation from incompleteness.

A trace that contradicts a policy is not the same as a trace that omits the evidence needed to judge it. The first needs correction; the second needs better records or an explicitly acknowledged bypass.

The checking framework describes a trace-by-constraint cost of O(|T| × |C|). However, coverage checking is only half-decidable and the implementation reports which passes are total. “Audit is a query” is not a promise that finite records prove there are no undeclared actions anywhere.

Three evidence planes, not three magic containers

Writer-side guard traceTool, target, graph, instant, actor, chain
Signed verdict factPolicy, outcome, attribution, hash, signature
Historical policy snapshotClaim and authority grants as of the decision

Having all three containers is insufficient if their contents are missing or contradictory. A signature attests to sealed evidence; it does not establish that the fact is true, that every action was logged, or that the policy covered everything relevant.

The external evidence benchmark exposed two implementation gaps: the original signature did not seal the writer binding, and denied writes did not retain their actor in the store. Quipu changed the verdict to include actor and principal chain inside its signed evidence hash. The published sufficiency results are for the improved format.

07 / WHAT WAS ACTUALLY MEASURED

Strong controlled results, narrow empirical reach

The central benchmark, Census, is a seeded deterministic lifecycle. It runs the same script twice: with gates enabled and with gates disabled. No LLM is used in this core loop. It creates three district graphs, three recorder identities, four policies, six planted defects, and 100 clean writes; then it corrects, composes, amends rules, and audits. The manifest is the ground-truth oracle.

0 vs 6Planted defects remaining, gated vs ungated, out of six
7 / 7Composition probes upheld the contract
50 / 50Accepted verdicts re-derived under historical rules
QuestionReported resultWhat it establishes / boundary
Does strictness cost only governed writes?Median 2.7 ms for compliant governed writes; 1.3 ms off-target in gated arm; 0.7 ms off-target in control.Policy claims are skipped off-target, but authority checks still cost time. These are illustrative single-run latencies.
Does gating prevent the planted defects?0/6 remain versus 6/6 without gates.Catches missing provenance, unauthorized and over-delegated writes, policy failure, a post-state conflict, and an invented predicate.
Does the audit agree externally?56 decisions agree verdict-for-verdict with SARC. Faithful export has 168 coverage discrepancies; padding non-applicable checks removes them.The mismatch is whether abstention must be explicitly recorded. It is not a disagreement over the gate’s verdicts.
Does composition remain conservative?All 7 probes succeed; four clean compositions pass with zero false refusals.Includes undeclared and expired labels, cross-chain refusal, obligations, bind-once overlays, and stable pack identity.
Does historical replay work?50 accepted outcomes re-derived; all 50 fail the amended claim. All 6 denials verify rules-in-force.Denial replay is an attestation check, not an outcome re-derivation: attempted deltas were rolled back.

The costs are measurable, not free

Control, off-target
0.7 ms
Gated, off-target
1.3 ms
Gated, governed
2.7 ms

Seed-42, release-mode Census measurements; not a throughput comparison with other databases.

Quipu is a Rust crate of roughly 67,000 lines with over 1,000 tests, backed by one SQLite file and exposed through CLI, REST, and MCP handlers. Its custom SPARQL 1.1 evaluator works directly over temporal facts, with temporal query parameters and RDFS widening.

The engineering characterization reveals a serious scaling ceiling: basic graph-pattern joins use nested loops and grow quadratically with store size. An interned-term cache gives a measured 3.9–4.9× constant-factor benefit. A separate in-memory read-model prototype changes one 133-second two-hop join to 0.15 ms, at roughly 385 resident bytes per fact. That is a prototype result, not the production evaluator’s general speed. Storage growth was approximately 8.3 KB per ingested episode from 1,000 to 20,000 episodes.

Do agents actually improve after refusals?

The agent arm supplied five actions, no view of Σ, and one revision using verbatim refusal feedback. Beyond an initial single run, it tested four Claude models with three trials each.

So the evidence supports feedback-driven repair for expressible constraints—not a general claim that strict stores make agents truthful. The experiment covers one task, one model family, and three trials per model.

A real trace did not get a clean bill of health

Five genuine decisions from Yupana’s pre-edit guard—two denies and three allows—were replayed against its shipped policy catalog. The audit found two violations and six incompleteness findings. The guard enforced a locally configured blast-radius rule not authored in Σ. Other gaps involved unexercised constraints, undeclared placement, and partial attribution. The result matters because the audit surfaced a real discrepancy rather than merely passing a hand-planted success case.

External evidence sufficiency: read the contents

DEMM-Bench asks whether records suffice to reconstruct eight properties: identity, authority, action boundary, policy basis, decision basis, resource touch, lifecycle context, and verification strength. Quipu’s self-added regime applies eight content-degradation conditions and evaluates 64 cases, yielding 512 property judgments.

Reader of the same exported evidenceCases marked sufficientOverclaim rateInterpretation
Property-level reader8/640%All 512 property cells correct; no underclaim; mean property accuracy 1.0.
Internal validity check8/640%Rejects degraded records, but does not localize each missing property.
Three-container checklist56/6475%Confuses surviving containers with intact evidence.
Trace-, ledger-, or schema-presence64/6487.5%Content deletion leaves the container present, making presence a poor sufficiency test.

This is a self-administered run using the external benchmark’s published construction oracle, with Quipu-authored degradation transforms and reader. The reader consumes evidence content, not degradation labels. Its perfect score is about this deterministic exported format and test pipeline—not an independently administered leaderboard win or proof of general audit completeness.

Census manifests were byte-identical across three repeated runs per seed and arm. Determinism makes regressions and independent re-execution easier; it does not make a planted scenario representative of every deployment.

08 / HOW TO READ THE CLAIMS

Six guarantees—and the things they do not guarantee

Design principleMechanismDo not infer
GS1 · Gated writesCandidate-state policy and shape checksPassing policies means factual truth.
GS2 · Permanent verdictsSeparate persistence, signed attributionA denial’s attempted facts survive, or signatures prove exhaustive evidence.
GS3 · Partitioned authorityDelegation intersection; protected relabelingA writer can expand a delegator’s grant or upgrade its own trust.
GS4 · Non-widening compositionConservative folds plus explicit coverageLabels are row-level access control, or every pair of trust chains is comparable.
GS5 · In-store auditStored specification, traces, verdicts; deterministic passesA complete specification can be discovered automatically from a finite trace.
GS6 · As-of replayHistorical data, labels, grants, policies, and shapesLatest-only rules can faithfully judge old decisions, or all denials can be re-derived.

The evaluation’s open questions

How well does this design work across many real ingestion tasks? How much policy-authoring work is required? Can the store remain practical under large joins and concurrent production workloads? The paper does not answer those with broad comparative experiments. Its future directions are a richer multi-task agent evaluation with competency-question oracles, and a portable governed-store contract that other substrates can implement and Census can score.

A useful mental model: Quipu is not a truth machine. It is an evidence-preserving contract boundary: writes must convince explicit checks, mixed graphs must carry honest labels, and later auditors must be able to distinguish “violated the rule” from “we cannot tell.”

Closing platform note: For a context platform, this paper contributes a governed graph-storage layer; it does not define a full context lake or an ontology-construction pipeline.