Ground the familiar.
Preserve the novel.
How a hybrid pipeline turns noisy skill declarations into shared concepts—without flattening every specialization into its nearest dictionary entry.
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
Emma Jouffroy, Warren Jouanneau & Marc Palyart · Malt
A string is not a concept
A freelancer writes “gestion de projet,” another writes “Project Management,” and a third writes “Web PM.” A search system that compares only text can miss the first pair’s equivalence. A system that merges everything similar can make the opposite mistake: losing the web specialization.
The paper tackles this at Malt, a multilingual freelance marketplace. Its inputs are free-text expertise declarations: translations, misspellings, abbreviations, compound skills, job titles, and new jargon. The intended output is a skills knowledge graph and a simpler taxonomy derived from it for downstream matching.
Entity identity
A stable identifier answers which concept? Different labels can name one entity. A Wikidata QID is a language-independent identity, not an English word.
Canonical label
A preferred name answers what should the interface display? One concept can have preferred labels in French, English, German, Dutch, and Spanish.
Relation
A connection answers how do concepts differ or relate? “Web Project Management” can be narrower than “Project Management” without being its synonym.
A taxonomy emphasizes broader–narrower organization. A knowledge graph can also carry identities, multilingual names, raw members, source metadata, and other relations. The paper describes a metadata-rich graph from which a streamlined taxonomy is extracted; it does not specify a complete formal ontology or executable relation schema.
Why combine top-down and bottom-up?
Start with a known ontology or graph and assign inputs to existing concepts. You gain shared identities and established labels, but the inventory can be too coarse or outdated.
Build concepts from the observed inputs. You can preserve new specialties, but repeated batches may invent duplicate names and inconsistent groupings.
The hybrid approach makes the external graph an anchor, not a ceiling. An initial mapping can be useful but too broad. The later curation step asks a different question: does this raw declaration really belong inside that canonical node?
Five stages, with a feedback loop
The implementation uses Gemini 1.5 Flash with temperature 0 and structured output. Here, “multi-agent” means separate prompted roles for linking, naming, reviewing, and merging. “Reflexion” is a loop that revisits rejected assignments using reasons and proposed replacement labels—not a guarantee that the model’s judgment becomes correct.
1. Reconciliation: context before identity
For each raw string, the system retrieves the top ten Wikidata candidates and their multilingual metadata. It adds the top fifteen co-occurring peer skills and top five professional categories from Malt’s profile graph. The LLM selects the best QID or QID combination and returns confidence, a concise rationale, and flags such as is_skill and is_compound.
Why add those neighbors? A short, ambiguous word rarely explains itself. Nearby skills and professional categories provide evidence about the intended field. Retrieval narrows the identity choices; context helps rank them. A model can still choose the wrong candidate, especially if the correct one never entered the candidate list.
The appendix also distinguishes direction-sensitive combinations. “English → French translation” is not interchangeable with “French → English translation.” Sorting all component identifiers into an unordered set would lose information.
2. Canonicalization: identity first, names second
Validated mentions are grouped by their unique QID combinations. The naming role creates one preferred label per target language, using a fallback order: professional, frequent platform wording first; an official Wikidata title second; a generated label only when necessary.
Each localized label carries provenance. The text describes malt, wikidata, and a generative origin; the appendix uses gemini for the last one. This is label-source metadata, not proof that the underlying concept or all its relations are correct.
3. Active curation: an initially plausible mapping can fail
The review role compares each raw member with its node’s preferred concept. The main text lists seven rejection criteria: ambiguity, specialization, semantic mismatch, non-skill, methodology, context, and sub-task. Rejections other than non-skills are grouped under NODE_CURATION for tracking.
The paper’s running example maps “gestion de projet web” to the broad project-management anchor. Curation then rejects that assignment as a specialization and proposes “Web Project Management.” Accepted members stay in the core node; meaningful outliers enter an orphan queue. The core label is rewritten using only accepted members. A node whose rejection rate exceeds 50% is flagged for human review.
4. Consolidation: local batches need global identity
The first batch seeds a baseline graph. Subsequent incoming nodes are compared with that growing baseline. Shared raw members or low edit distance between labels nominate candidate pairs; the LLM decides MERGE or KEEP_SEPARATE from their metadata. A merge records a surviving ID, and decisions are cached to avoid repeated work.
This avoids asking the LLM about every possible pair. With N nodes, an exhaustive undirected comparison requires N(N−1)/2 pairs. The heuristics reduce the candidate set, but also risk missing duplicates with very different spellings. The appendix protects specialization boundaries and fundamentally different software versions; it prefers a Wikidata ID over a synthetic ID when a merge is valid.
5. Orphan recovery: turn the mismatch into a new hypothesis
The main text derives a synthetic orphan identifier from the suggested preferred label. Inputs with the same suggested label are routed together in the next epoch. The loop repeats until the orphan queue is empty. In the intended outcome, “Web PM” and “gestion de projet web” become one specialized node rather than two fragments or one overly broad core node.
Conceptual teaching diagram, not the paper’s formal graph schema. Side arrows mean “names this concept”; the vertical arrow preserves the specialization boundary.
Watch aliases converge—and specializations survive
This browser-only toy uses a tiny, explicit lookup table in place of Wikidata retrieval and LLM judgments. It is an illustration of the control flow, not a reproduction or a benchmark. Unknown text is held for review rather than assigned an invented identity.
Recognized examples are those above, plus “Project Manager,” “Web project delivery,” “Python,” and “Python 3.” Edit the list to test duplicates or an unknown skill.
“Drift” exposes a weakness of label-derived IDs: equal meanings need not receive equal names. The toy merger recognizes web aliases explicitly; real systems must find candidate pairs first.
Routing trace
Three experiments to try
- Hybrid, stage 2 → stage 3: web mentions first share a broad anchor, then leave its core group. Rejection here means “wrong identity assignment,” not “worthless input.”
- Hybrid, stage 5, drift on, consolidation off: equivalent web aliases fragment into different synthetic nodes. Turn consolidation on to recover one node.
- Compare the baselines at stage 5: anchor-only loses specialization; fragment-only preserves wording but does not unify identity. These baselines are explanatory caricatures, not evaluated methods from the paper.
Large vocabulary, partial validation
The evaluation uses proprietary Malt declarations, not a public benchmark. The authors report a hand-annotated Wikidata gold standard curated by five domain experts without overlap. Crucially, they say the current evaluation covers initial retrieval and pre-consolidation phases; final post-consolidation behavior has not yet been formally benchmarked.
Raw expertise strings
27,743 resolved; 8,294 unmapped. The resolved share is approximately 77%.
Canonical skill nodes
15,010 initial semantic groupings were streamlined into this node inventory.
Preferred labels
Exactly five per node: French, English, German, Dutch, and Spanish.
Compression: 1 − 13,298 / 27,743 = 52.1%. The denominator is the mapped vocabulary, not all raw strings. The average is 2.08 mapped variations per canonical node. Compression measures consolidation, not whether every merge was semantically appropriate.
Gold-standard reconciliation
Paper-reported global metrics. “Found coverage” is described as precision among inputs for which retrieval was attempted. These metrics use different definitions; do not treat them as parts of one pie.
Usage is highly concentrated
Top 1,000 nodes: 82.74% of usage volume.
Top 5,000 nodes: 97.25% of usage volume.
The authors distinguish the lexical long tail (redundant spellings) from the semantic long tail (rare capabilities). A useful normalizer should compress the former without deleting the latter.
Domain differences matter
| Domain | Alignment coverage | Found coverage |
|---|---|---|
| Video Games | 88.1% | 91.8% |
| Industrial Engineering | 86.0% | 89.9% |
| Data & Analytics | 83.5% | 88.7% |
| Tech / Software | 81.7% | 86.8% |
| Marketing | 76.7% | 82.4% |
| Communication | 75.4% | 81.0% |
Source: paper Table 1. Structured fields have stronger reconciliation scores than fields with softer boundaries and shifting terminology. The paper does not provide domain sample sizes or uncertainty intervals here.
Provenance and curation
Across preferred labels, the reported origins include 67.28% platform usage and 22.08% Wikidata titles. After accounting for overlap, their union is 80.65%; the remaining 19.35% is generated. The two source percentages should not simply be added. Also, calling 80.65% of the “graph” factual is stronger than the demonstrated measurement, which tracks label provenance.
| Rejection category | Share of all rejected inputs |
|---|---|
| SCORE_REJECTED | 31.9% |
| NODE_CURATION | 28.1% |
| NOT_A_SKILL | 13.7% |
| WIKIDATA_NOT_FOUND | 11.4% |
| NOT_IN_NODE | 10.7% |
| ERROR_NOT_KNOWN | 4.2% |
Source: paper Table 2. The 13.7% figure is a share of rejections, not of all inputs. The rejection mix also shows that “unmapped” is not synonymous with “not a skill.”
Qualitative audits report 0.01–0.06 manually removed outliers per sub-graph. Audit size and the graph-level aggregation are not detailed enough to turn this into a general accuracy claim. The taxonomy is already integrated into candidate matching, but no measured matching-quality lift is reported.
“Self-healing” is a design ambition, not a theorem
External grounding reduces—not eliminates—errors
A valid QID can still be the wrong QID. JSON schemas constrain formatting, not truth. The reported 19.1% wrong-guess rate is a concrete reason not to read the paper’s “prevent hallucinations” language literally.
Hashes stabilize strings, not meanings
A deterministic hash guarantees the same string yields the same identifier. It cannot guarantee the model uses the same string for synonymous skills. Renaming a concept can also change a label-derived ID unless there is an identity registry.
An empty queue is not proof of correctness
The paper gives no formal convergence proof, epoch limit, or oscillation analysis. A loop can stop after making a bad merge; it can also keep splitting and renaming. Termination and semantic quality need separate tests.
Coverage is not linguistic fidelity
Five labels per node demonstrate complete locale fields, not perfect translations or preservation of regional distinctions. The authors explicitly flag English-centric normalization and gender-marked occupational terms as future concerns.
Details that an implementer must resolve
- Identity rules differ across text and appendix. The main method hashes suggested labels into orphan IDs; the orphan prompt requests generated
SYNTH_SKILL_XXXidentifiers. These are not an identical persistence mechanism. - Compound handling is inconsistent. The reconciliation prompt permits component QID combinations for examples such as React and Node.js, but its specificity rule says composite lists should be created only for truly directional mappings. This needs a single explicit policy before implementation.
- Rejection labels differ. The main seven-way taxonomy is not the same as the appendix’s criteria, which include
DOMAIN_SPECIALIZATIONandDISTINCT_TOOL_OR_LANGUAGE. Standardize the vocabulary and version its meaning. - The graph is underspecified. The narrative promises specialized branches and relational metadata but does not fully define edge types, relation-validation tests, or serialization. The figures also contain inconsistent project-management QIDs. Treat the identifier examples as illustrative, not authoritative lookup data.
- Deduplication can depend on batch order. Comparing newcomers with a baseline is efficient, but an early mistaken survivor can influence later decisions. The proposed cache needs a way to invalidate decisions after labels, members, or model policies change.
The authors’ future-work plan includes comparisons with ESCO and other AI methods, testing consolidation and human review, measuring tokens and latency, and analyzing failure patterns and manual intervention rates. Those are precisely the missing checks needed to evaluate the full architecture.
The kernel is routing, not a giant prompt
The following pseudocode compresses the paper’s five roles into an implementable outline. It is a teaching sketch; the paper’s appendix supplies actual prompts, not a complete public implementation or its proprietary dataset.
baseline, orphan_queue, merge_cache = load_state()
for epoch in bounded_epochs: # bounded: added safety measure
incoming = []
for raw in next_batch(orphan_queue):
candidates = retrieve_wikidata(raw, k=10)
context = profile_neighbors(raw, skills=15, categories=5)
proposal = link(raw, candidates, context) # schema-validated
route_invalid_or_uncertain(proposal)
incoming.append(proposal)
nodes = group_by_identity_and_direction(incoming)
nodes = name_in_locales(nodes, ["fr", "en", "de", "nl", "es"])
core, orphans = curate_members(nodes)
flag_for_review(core, rejection_fraction_above=0.5)
for node in core:
for target in candidate_duplicates(node, baseline):
decision = cached_or_judge(node, target, merge_cache)
apply_merge_or_keep(decision)
persist(node)
orphan_queue = route_by_suggested_label(orphans)
if not orphan_queue: break
# Queue exhaustion is operational completion, not semantic validation.A minimal teaching record
{
"id": "local:web-project-management",
"anchor": "external:project-management",
"preferred_labels": {
"en": {"text": "Web Project Management", "source": "generated"}
},
"raw_members": ["Web PM", "gestion de projet web"],
"proposed_relation": "narrower_than",
"curation_reason": "SPECIALIZATION",
"status": "proposed"
}Illustrative local schema, not the paper’s exact output. Only one language is shown for brevity. A local identity avoids claiming an unverified Wikidata identifier; “proposed” distinguishes model hypotheses from accepted graph facts.
Tests worth writing before automation
- Identity invariance: translations and harmless typos converge across batch orders; direction-reversed translation skills do not.
- Granularity preservation: a domain specialization remains distinct while retaining an explicit broader connection.
- Merge safety: aliases merge, Java and JavaScript do not, and fundamentally different framework versions remain separate.
- Loop stability: track orphan count, repeated assignments, renamed IDs, and held-for-review cases per epoch.
- Evidence separation: evaluate linking, member curation, localized names, duplicate detection, and relation correctness individually before testing downstream matching.
The useful move is to preserve the mismatch
A broad external anchor gives new text a semantic foothold. Curation then distinguishes genuine aliases from concepts that the anchor inventory represents too coarsely. The latter become structured hypotheses for graph growth. That combination is the paper’s most teachable contribution—even though the accuracy, stability, and cost of its complete consolidation loop still require evaluation.
One connection to context platforms: this paper offers a concrete pattern for evolving a domain concept layer over noisy source text: keep shared external identities where appropriate, preserve novel distinctions locally, and retain source metadata rather than silently replacing the originals.