Paper explained · Ontology & agent execution

An ontology that agents can execute

OaK turns domain concepts, evidence, and recurring reasoning steps into a task-specific interface—then repairs that interface using execution failures.

Toward Effective and Reliable LLM Agents via Dynamic Ontology
Xiaohui Zhang, Zequn Sun, Chengyuan Yang, Yuanning Cui, Lingbing Guo, and Wei Hu

Original paper · PDF

1. The problem: a fact is not yet a valid decision

Suppose a travel agent finds a cheap hotel with the requested room type. The answer sounds reasonable. But the hotel requires a five-night stay and accommodates only two guests; the trip is for three guests over two nights. The problem is not necessarily a hallucinated hotel. It is a missing connection between known facts and the procedure that selects a hotel.

Putting more documents in the prompt may expose those facts without ensuring that the agent uses them. Giving the agent a graph may make the relationships searchable without ensuring that it checks all the relevant constraints. OaK asks a different question: can we make both the domain vocabulary and the permitted computations explicit?

Kernel K = (schema S, function catalog F)
Evidence graph G = Build(S, task corpus)

Schema S: what can be represented

Entity types, properties, identity fields, and typed relationships. For example, Hotel, City, and Hotel → located_in → City.

Graph G: what the data says

Concrete instances extracted from the reference corpus: this hotel costs $80 per night, has capacity two, and is in a particular city.

Functions F: what can be computed

Typed executable procedures: find hotels in a city, apply price and occupancy constraints, and return eligible candidates.

The graph is the evidence substrate; it is not one of the two components in the paper’s definition of the kernel. The schema and functions together mediate access to that evidence. At inference, the LLM chooses a declared function and binds its arguments rather than inventing a fresh computation from scratch.

The distinctive move: OaK does not stop at “retrieve graph facts for the model.” It compiles repeated graph reasoning into tools, then improves both the data model and those tools using downstream task feedback.

This is a lightweight, operational use of “ontology”: a YAML or JSON schema plus executable functions, not simply a large axiomatic description of a domain. Data requirements live in schema declarations; computational and workflow requirements live in function control flow.

2. Two stages—not an ontology that rewrites itself during every answer

OaK construction and inferenceDuring training, schema, graph, functions and agent execution feed a judge that sends repairs back. At inference the schema and functions are frozen but the graph is built from the test corpus.Construction · training examplesInference · unseen queryDraft & verifyschema SInstantiategraph GComposefunctions FRun agent& scoreJudge traces + scoresPropose artifact repairsFrozen S* and F*Build Gq from query corpusAgent calls F* over Gq
Conceptual redraw of the paper’s workflow. The judge repairs schema and functions between construction rounds; the test-time agent does not edit them.

“Dynamic” means task-conditioned construction and refinement. Once construction ends, the schema and function catalog are frozen. For each unseen query, OaK still extracts a graph from the accompanying test corpus under that fixed schema. A new graph is not a new kernel.

This separation matters. Training failures can improve the interface, but test answers are produced without another round of kernel repair. The paper describes the frozen kernel as the only channel through which the agent reaches the data.

3. How the construction loop works

1 Analyze requirements, draft a schema, check consistency

Each round samples fresh training examples, where each example contains a query and the corpus needed to answer it. An LLM analyzes task scope, entities, relations, and constraints. It drafts a schema using that specification and the previous round’s repair feedback.

Every entity type has a primary key, and each relationship declares source and target types. OaK translates the lightweight schema to OWL and uses the HermiT reasoner to check disjointness, restrictions, property axioms, global consistency, and unsatisfiable classes. A failing draft is returned for revision with reasoner feedback.

Consistency is not completeness. A schema that omits hotel capacity can be logically consistent and still useless for safe accommodation selection. HermiT checks the logical encoding, not whether the schema captures every business requirement or whether extracted facts are true.

2 Extract a graph using chunk–map–merge

The corpus is split into token-bounded chunks. An LLM extracts typed entities and relations from each chunk. Because chunks overlap in what they describe, the same entity can appear several times. OaK reconciles duplicates using a key signature:

key(entity) = (entity type, primary-key value)

Candidates with matching signatures collapse into one canonical node. Relationships are reattached to canonical endpoints, and duplicate relations collapse too. For example, two extracted Hotel records with primary key H-17 become one hotel; a Restaurant with the same key value is still a different typed entity.

The equality rule is exact, not a general solution to entity resolution. It assumes the chosen primary key identifies the real object correctly. Names alone can collide; inconsistent identifiers can split one object into several. The paper explains duplicate identity merging but does not specify a general policy for conflicting property values.

3 Compile recurring reasoning into domain tools

A composer reads the current schema, graph, training questions, generic operator library, and prior function feedback. It identifies repeated patterns and realizes each as a function with typed inputs, typed outputs, and an executable implementation. Functions are tested for executability on the current graph.

Composition
Chain lookup → filter → projection → aggregation.
Specialization
Fix an operator’s parameters for a domain-specific use.
Adaptation
Add lightweight preprocessing or postprocessing.

Appendix B lists nine public operators: runtime-slot extraction, entity lookup, relation traversal, property projection, categorical filtering, relation-connectivity filtering, numeric filtering, set-overlap filtering, and aggregation. Aggregation supports counts, sums, minima, maxima, and averages, optionally grouped by a field.

Most of these are familiar data operations. The value is in bundling them into reusable task-level procedures. Instead of repeatedly asking an LLM to remember how occupancy and minimum stay interact, one function encodes the checks.

Not everything becomes deterministic: the runtime-slot extraction operator uses an LLM to map a natural-language query to declared typed slots. Entity lookup also supports fuzzy matching and optional truncation. Declared interfaces reduce freedom; they do not remove all ambiguity.

4 Score the answers, inspect traces, repair artifacts

A ReAct agent runs the training questions using those functions. The trajectory records operator traces, function outputs, and final answers. Answers receive official dataset scores; a judge then examines scores and trajectories alongside the schema, graph, and functions.

repair = (target artifact, add/delete/modify, patch, diagnosed reason)

Schema-level suggestions return to schema drafting; function-level suggestions return to the composer. New examples are sampled for the next round. The loop ends when no blocking fault is identified or the iteration budget is exhausted.

A useful diagnosis distinguishes “the constraint cannot be represented” from “the constraint exists but no function enforces it.” The hotel case in the paper requires both repairs: add minimum_nights and maximum_occupancy, then modify the accommodation function to filter on them.

4. Playground: repair a hotel-selection kernel

This small, deterministic teaching simulation follows the paper’s accommodation case. It is not the authors’ code or benchmark data. Three synthetic hotels are connected to the same city. The function selects the cheapest eligible hotel; the nightly budget is a ceiling, not a whole-trip budget.

Synthetic reference data. Room type is “entire unit” for every hotel. Constraint fields exist in the source, but stage 0 does not represent them.
Hotel / source IDPrice per nightCapacityMinimum nightsFunction decision

Declared schema

Executed operator trace

Try this: advance from stage 0 to stage 1. The graph now contains capacity and minimum-stay fields, but the answer remains wrong. Advance to stage 2: the procedure finally uses the facts. Then lower the budget below $120; a grounded “no eligible hotel” is better than an invented solution.

The separate teaching audit checks the full synthetic source constraints even when the simulated kernel cannot see them. This makes the defect visible; it is not an extra data channel available to the simulated agent. Changing the stage manually is a way to compare versions—not test-time self-modification.

The small executable idea behind the simulation

function accommodation(graph, args, repaired) {
  assertTypedArguments(args); // nonnegative budget; positive integer guests/nights
  let rows = lookupHotelsConnectedTo(graph, args.city);
  rows = rows.filter(h => h.price <= args.budget);
  rows = rows.filter(h => h.room_type === args.room_type);
  if (repaired) {
    rows = rows.filter(h => h.maximum_occupancy >= args.guests);
    rows = rows.filter(h => h.minimum_nights <= args.nights);
  }
  return rows.sort((a,b) => a.price - b.price);
}

The important detail is the direction of the comparisons: capacity ≥ guests, but minimum_nights ≤ stay_nights. Adding the fields without adding these predicates does not change the outcome.

This browser prototype implements typed input checks, a small city–hotel relation, versioned schema projection, numeric and categorical filters, cheapest-first selection, and a readable trace. It does not implement LLM extraction, HermiT verification, automated function composition, or a judge-driven construction loop.

5. What the evaluation actually shows

The paper evaluates three settings: TravelPlanner (multi-day itineraries), CRMArenaPro (synthetic CRM workflows, policy, text, and database tasks), and ToolQA (compositional questions over heterogeneous corpora). All methods use two backbones: DeepSeek-v4-flash and GPT-4o-mini. Construction and inference use the chosen backbone; the judge is fixed to Claude Sonnet 4.6.

Each construction round samples 20% of the training split; construction has at most five rounds. The ReAct agent has at most 20 steps per query. Comparators are ReAct, AFlow, MemP, ReCode, and AgentSquare.

Selected primary aggregate results from Tables 1–3. Scores are percentages; gains are percentage points versus the strongest listed baseline for that row, not relative-percent improvements.
Benchmark / metricBackboneBest baselineOaKGain
TravelPlanner · FinalDeepSeekMemP · 51.5055.90+4.40
TravelPlanner · FinalGPT-4o-miniMemP · 16.3019.70+3.40
CRM B2B · Avg.DeepSeekMemP · 66.4478.38+11.94
CRM B2C · Avg.DeepSeekMemP · 67.2775.20+7.93
CRM B2B · Avg.GPT-4o-miniMemP · 49.2463.93+14.69
CRM B2C · Avg.GPT-4o-miniMemP · 49.8967.62+17.73
ToolQA · weighted Avg.DeepSeekMemP · 55.0264.58+9.56
ToolQA · weighted Avg.GPT-4o-miniReAct · 50.4559.32+8.87

TravelPlanner: checking pieces is not checking the whole

Micro scores measure individual satisfied constraints. Macro scores measure plans satisfying all applicable constraints within one family. Final pass requires satisfying both commonsense and hard-constraint families completely.

A plan can have excellent local scores and still fail because one dependency is broken. If a hotel choice violates capacity, many otherwise correct itinerary decisions do not rescue the plan. OaK leads TravelPlanner’s macro and final metrics on both backbones, but not every micro metric.

With DeepSeek, MemP has better hard-constraint micro performance (63.91 vs. 59.31), yet OaK has better final pass (55.90 vs. 51.50). This is consistent with better coordination, not simply more individually satisfied checks.

Final pass · DeepSeek

ReAct
15.30
ReCode
48.20
MemP
51.50
OaK
55.90

Bars share a 0–100% scale. These are reported task outcomes, not playground measurements.

CRM: large database gains, not uniform dominance

The CRM average is the equal-weight mean of four categories. Workflow, Policy, and Database use exact match; Text uses exact match for discrete outputs or token-level F1 for free-form outputs. These are therefore category scores with different underlying scoring rules, not one universal accuracy measurement.

OaK is particularly strong on database and policy tasks. GPT-4o-mini reaches 84.69 on B2B Database versus MemP’s 49.38. But on DeepSeek B2C Text, MemP scores 43.77 and ReAct 43.47, both above OaK’s 39.87. Structured execution is not automatically the best route for every free-form task.

ToolQA: best aggregates still hide regressions

ToolQA uses normalized exact match, with a weighted average over Flight, Coffee, Airbnb, DBLP, and Yelp. OaK leads all five DeepSeek subsets. With GPT-4o-mini, however, ReAct beats it on Flight (37.38 vs. 31.10) and Airbnb (72.28 vs. 54.40). Strong improvements in other subsets yield the higher overall score.

Interpretation: the benchmark results support improved task performance and the usefulness of executable, schema-grounded procedures. They do not by themselves establish universal reliability, formal answer correctness, or robustness to prompt injection.

6. What the ablations, loop analysis, and costs add

Without composition

The agent gets generic operators directly. It must reconstruct recurring multi-step computations for every query. Performance falls; the main-text CRM analysis identifies this as the largest degradation there.

Without functions

The LLM reasons over the graph itself, with minimal retrieval when the graph is too large for context. Graph access alone does not recover the full model’s performance.

Without refinement

Only the first construction round is used. Missing constraints and incomplete procedures remain uncorrected, reducing performance.

These experiments help isolate the contribution from making computations reusable, rather than merely supplying structured facts. For GPT-4o-mini on TravelPlanner, Appendix D reports final pass dropping from 19.70 to 9.90 without the function module. The paper’s remaining ablation plots are summarized qualitatively here; numerical values not stated in the text are not reconstructed.

Construction-round curves rise rapidly early and largely plateau by rounds four to five. This suggests early feedback repairs important gaps, but it does not prove the loop converges or that five rounds is optimal for other domains.

Appendix C reports three DeepSeek TravelPlanner runs: final pass is 54.00, 58.20, and 55.50, summarized as 55.90 ± 2.13. That is useful evidence against a single lucky run in this setting. It is not a multi-seed comparison of every baseline across all benchmarks.

Cost is a trade-off, not a free improvement

The TravelPlanner cost analysis uses runtime and token counts as resource proxies. OaK uses fewer input tokens than ReAct and MemP, but more runtime and output tokens than several baselines. Query-specific graph construction is a major source of the extra work.

Packaging reasoning into a function can shrink what the LLM must repeatedly read and decide. Extracting the graph still costs something. The authors suggest graph reuse could amortize that cost; reusable and incrementally updated graphs remain future work rather than a demonstrated deployment result here.

7. Boundaries of the claim

The strongest defensible takeaway is not “ontologies eliminate hallucinations.” It is: jointly engineering the representation and the executable reasoning interface can improve benchmark outcomes and make failures easier to localize.

8. A practical mental model

Think of OaK as building a small domain API from examples. The schema says which objects and facts the API understands. The graph supplies its records. Functions encode repeated computations. A task evaluator reveals failures; a judge routes repairs to the schema or to the code. The resulting interface is then fixed for unseen queries.

Three questions to retain:
Can the relevant constraint be represented?
Does the graph faithfully contain it?
Does the selected procedure actually enforce it?

Closing connection: For a context platform, this paper offers a concrete pattern for placing typed, testable procedures above an ontology-governed evidence graph, rather than relying on graph retrieval alone.