Explained · enterprise semantics

Business meaning before business answers

How Genie Ontology connects governed definitions, learned context, live analysis, and repeatable product workflows.

A close reading of How Genie Ontology powers product development at Databricks
Nicole Esibov · Databricks · October 5, 2026

The central idea: an agent needs more than access to data. It needs to know what the data means, which sources deserve trust, and which information the current user may access.

1. Why a correct query can still produce a wrong answer

Imagine a product manager asking, “How did weekly active users change?” A database can count rows perfectly while answering the wrong business question. Does “active” mean opening a page, submitting a question, or completing an analysis? Are mobile users included? Are internal test accounts excluded? Is one person using three surfaces counted once or three times?

These are not SQL syntax questions. They are questions about business semantics: the meanings, rules, and relationships an organization has agreed to use. An agent that guesses them can produce a polished report with misleading numbers.

Data tells us what happened

An event records a user, time, surface, and action. A table stores those records. A query computes a result.

Context tells us how to interpret it

A governed definition says which events count, a source policy says which asset is authoritative, and provenance tells us where a rule came from.

The article describes Genie Ontology as a continuously updated understanding of institutional knowledge. It begins with human-curated, governed concepts in Unity Catalog, then supplements them with knowledge learned from dashboards, documents, queries, notebooks, and connected applications.

Here, “ontology” should be read in the product’s practical sense: business definitions, rules, relationships, and trusted sources. The post does not specify a formal ontology language, an inference calculus, or a particular graph-storage design.

Four terms worth separating

Term in the postIts roleWhat it does not guarantee by itself
Metric viewA governed KPI with one shared definition; the example finds a certified weekly-active-users view.That the underlying data is complete or that every query uses the intended time window.
PageA governed wiki in Unity Catalog that supplies curated business context.That every organizational concept has been documented.
Ontology snippetA business fact automatically learned and evaluated from existing assets; its definition, origin, author, and authority score can be inspected.That a learned fact is equivalent to a human-certified rule.
OntoRankRanks context by authority using signals such as certification, usage, and authorship, to support trusted and relevant retrieval.A disclosed scoring formula or a proof that the selected source is correct.

2. The mechanism: choose meaning, retrieve evidence, compute

The demonstrated task is a weekly review of Genie One adoption. Users interact across web, desktop, mobile, Slack, and Teams. Some dashboards are authoritative; others are stale. Plans are in Google Docs, bugs in Jira, and feedback in Slack. The report must combine these different kinds of information rather than treat them as interchangeable text.

1 · InterpretFind business definitions and the review template.
2 · SelectIdentify trusted assets and relevant context.
3 · ExecuteSearch live systems and query usage data.
4 · InspectReview the report and its citations.

Conceptual reconstruction of the demonstrated workflow, not a disclosed internal architecture.

1

Start from the team’s actual reporting task

The manager requests an adoption review using a standardized template in a Google Doc. The template supplies the output structure; the ontology supplies the business meaning needed to fill it.

2

Locate data assets and interpretive context

The visible trace first examines relevant dashboards, agents, and tables, then curated Pages, then potentially relevant ontology snippets. It finds the certified weekly-active-users metric view. The post says Genie favors human-certified assets before exploring on its own.

3

Combine search with fresh computation

Genie uses MCP to search live Google Drive documents and Jira tickets, then runs SQL against usage tables. MCP is the tool-connection mechanism here; it is not the business definition. Documents can explain plans and incidents, while SQL measures current usage. Neither alone supplies the complete review.

4

Expose the sources behind the report

The manager opens citations and sees metric views, Pages, and learned snippets. For a snippet, the interface exposes its definition, origin, author, and OntoRank authority score. The author’s profile provides further domain context. This makes a result inspectable rather than merely fluent.

Two different checks: authority asks “Should I trust this source for this question?” Permission asks “May this user access it?” The post says context is permission-aware through Unity Catalog and connected sources. Permission must not be traded off against relevance or popularity.

Why supplying context can reduce exploration

The post contrasts this workflow with a general-purpose agent that may crawl schemas, sample tables, read documents, and run repeated queries to discover the right definitions. Pre-existing business context can narrow that search before and during reasoning. This is a plausible mechanism for reducing effort and misinterpretation, but the article provides no measured latency or token-cost comparison.

3. From one answer to an ongoing workflow

The article continues beyond the initial report. These later steps matter because the ontology supports a recurring product-management activity, not just a one-off lookup.

Next stepWhat the manager doesWhat a careful reader should distinguish
Investigate a spikeAsks why new users increased rather than manually breaking usage down by surface and account and searching roadmap, Jira, and Slack context.A breakdown establishes where growth occurred. A nearby launch or incident suggests an explanation, but does not establish causality.
ForecastAsks Genie One to forecast the trend and flag accounts to watch.The capability is demonstrated, but model choice, assumptions, uncertainty, and forecast accuracy are not reported.
SchedulePlaces the adoption review on a recurring schedule.Automation repeats the workflow; it does not remove the need to monitor changed definitions, permissions, or data availability.
Share an agentTurns the conversation into a Genie Agent using the sources from this and prior topic-related conversations.Reusing context is useful, but a shared agent still needs access checks for each user.
Take actionThe closing account says the workflow also kicked off a follow-up ticket.The post gives much less detail about the write operation than about retrieval and reporting.

The summary also mentions custom skills and MCP writes to external tools. Those are stated capabilities, not mechanisms evaluated in detail in this walkthrough.

Hands-on · synthetic data

4. Playground: popular is not the same as authoritative

Two assets both match the words “weekly active users.” One defines a governed metric across surfaces; the other is a popular legacy dashboard that counts web page-open events. Change the ranking policy and see how the apparent adoption trend changes.

This is an educational simulation, not Genie, OntoRank, or a recreation of Databricks data. All records, scores, definitions, and permissions below are invented.

The remaining weight goes to relevance, usage, and authorship.

Only the incident reviewer may read the simulated restricted Jira note. The role does not change the metric.

Metric candidateRelevanceCertificationUsageAuthorshipToy score

Computed adoption

Each bar is the selected definition applied to the available event feed.

Authorized explanatory evidence

A note is evidence about an event, not proof that the event caused the observed change.

Inspect the selected definition and provenance
Inspect the complete synthetic event ledger

Each row represents one event, not necessarily one active user. “Internal” marks a test account. Disabling the mobile feed hides its events from computation, not from this teaching ledger.

WeekUserSurfaceActionInternal?

Three experiments to try

  1. Move certification weight to zero. The popular legacy asset wins. Its exact match to the words does not make its event-count definition match the governed KPI.
  2. Reset, then disable the mobile feed. Even the right definition undercounts when evidence is incomplete. Authority is not data completeness.
  3. Change the user role. Restricted explanatory evidence becomes visible only to the authorized role. A higher source score could not make it visible to the product manager.

5. Work through the example by hand

In the complete synthetic feed, Week A has questions from external users u1, u2, u3, u4. Week B has questions from u1, u2, u3, u4, u5, u6. A repeated question by u1 is still one active user. An internal account is excluded. The governed result is therefore 4 → 6 users, or 50% growth.

The legacy dashboard instead counts web page-open events: two in Week A and two in Week B. It reports 2 → 2 events, or 0% growth. This is not a computational error. It answers a different question.

With the mobile feed unavailable, the governed query sees only three active users in each week, producing 0% observed growth. The definition remains correct, but the evidence is incomplete. The report should flag that condition rather than present the result as complete cross-surface adoption.

Important lesson: ranking is only one layer. A defensible answer also needs definition compatibility, adequate data coverage, correct execution, and a clear statement of what is known versus inferred. A source winning the ranking is not a reason to silently redefine a KPI.

6. A small, inspectable prototype

The playground uses a deliberately simple ranking formula:

score = w × certification
      + (1 − w) × (0.7 × relevance + 0.2 × usage + 0.1 × authorship)

This formula is invented for teaching. The article names certification, usage, and authorship as authority signals, but does not disclose OntoRank’s algorithm or its weights. Here, the invented relevance feature makes the contrast easy to explore.

A minimal context-assisted reporting loop can be written as follows. The stricter compatibility check below is an implementation recommendation; the playground intentionally allows a mismatched candidate to show the failure mode.

// All policies and schemas below are illustrative.
function review(user, request, sources, events) {
  const visible = sources.filter(s => mayRead(user, s));
  const compatible = visible.filter(s =>
    s.concept === request.concept &&
    s.definitionVersion === request.definitionVersion);
  if (!compatible.length) return {status: 'needs clarification'};

  const source = compatible.sort((a,b) => score(b)-score(a))[0];
  const data = events.filter(e => mayRead(user, e));
  const counts = computeUsing(source.definition, data);
  const notes = searchLiveTools(user, request); // access enforced here too
  return {
    counts, coverage: checkCoverage(data, request),
    citations: [source.provenance, ...notes.map(n => n.provenance)],
    explanations: notes, // associations, not automatic causal claims
    definitionVersion: source.definitionVersion
  };
}

The essential separation is between selection and execution. Context chooses the rule and the source; deterministic analysis computes the measure. The answer carries the definition and provenance so a reviewer can inspect both.

For a production system, access control belongs at retrieval, tool access, and query execution—not just in this illustrative client-side filter. Learned snippets also need source links and handling for outdated or conflicting definitions. The blog describes automatic maintenance, but does not explain those maintenance procedures.

7. What the article establishes—and what remains untested

This is a product-use walkthrough, not a controlled evaluation. Its evidence is a narrated internal workflow with screenshots and interface examples. Its reported outcome is an executive-ready review, follow-up investigation and forecasting, recurring scheduling, a follow-up ticket, and a shareable agent.

Supported by the walkthrough

  • Curated metric views and Pages can be used alongside learned ontology snippets.
  • The manager can inspect source citations, snippet provenance, authorship, and authority scores.
  • Live document/ticket search and SQL analysis are combined in one task.
  • The conversation can lead to follow-ups, scheduling, and a reusable agent.

Not established by this post

  • A numerical improvement in answer accuracy, latency, cost, or adoption.
  • A benchmark against a general-purpose agent.
  • Calibration or ranking quality of OntoRank.
  • Forecast quality or causal validity of spike explanations.
  • Technical details of graph schemas, snippet extraction, conflict resolution, or freshness guarantees.

Practical failure modes to keep in mind

Authority can lag reality. Certification is valuable, but a formerly correct definition may become outdated. High usage can reflect familiarity rather than truth. Authorship is a trust signal, not a universal endorsement. These are general risks of the design pattern, not failures measured in the article.

Permission-aware answers may be partial. An agent should not disclose restricted context, but it should avoid turning missing evidence into a confident explanation. A shared agent cannot simply inherit the creator’s effective access for everyone.

Citations make checking possible, not automatic. A citation can identify a source without demonstrating that the claim follows from it. For a KPI, the reviewer still needs the chosen definition, filters, time window, and data coverage. For a forecast, the reviewer needs assumptions and uncertainty.

Recurring workflows amplify both good and bad choices. Scheduling saves repeated effort, but repeated semantic mistakes also scale. A robust review process would monitor definition changes and missing feeds and require appropriate approval for external writes. The post does not describe these controls in detail.

Bottom line: the article’s strongest contribution is a concrete illustration of enterprise context in action: curated meaning plus learned context, authority-aware source selection, live computation, and inspectable provenance. It motivates the architecture; it does not quantify its superiority.

8. One closing connection

For a context platform, the transferable lesson is to preserve the distinction between governed meaning, learned context, and live evidence—and attach permissions and provenance to each. The article supports that pattern, but does not provide a blueprint for a context lake or a specific context-graph implementation.