Paper explained · entity identity & multilingual skill graphs

Ground the familiar.
Preserve the novel.

How a hybrid pipeline turns noisy skill declarations into shared concepts—without flattening every specialization into its nearest dictionary entry.

An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
Emma Jouffroy, Warren Jouanneau & Marc Palyart · Malt

Original paper · PDF · Full text and prompt appendix

The central idea: use Wikidata to stabilize established concepts, then treat semantically meaningful mismatches as candidates for new concepts—not merely as errors to discard. The architecture is promising; the published evaluation mainly tests retrieval and the pre-consolidation pipeline, not the complete refinement loop.
01 · The problem

A string is not a concept

A freelancer writes “gestion de projet,” another writes “Project Management,” and a third writes “Web PM.” A search system that compares only text can miss the first pair’s equivalence. A system that merges everything similar can make the opposite mistake: losing the web specialization.

The paper tackles this at Malt, a multilingual freelance marketplace. Its inputs are free-text expertise declarations: translations, misspellings, abbreviations, compound skills, job titles, and new jargon. The intended output is a skills knowledge graph and a simpler taxonomy derived from it for downstream matching.

Entity identity

A stable identifier answers which concept? Different labels can name one entity. A Wikidata QID is a language-independent identity, not an English word.

Canonical label

A preferred name answers what should the interface display? One concept can have preferred labels in French, English, German, Dutch, and Spanish.

Relation

A connection answers how do concepts differ or relate? “Web Project Management” can be narrower than “Project Management” without being its synonym.

A taxonomy emphasizes broader–narrower organization. A knowledge graph can also carry identities, multilingual names, raw members, source metadata, and other relations. The paper describes a metadata-rich graph from which a streamlined taxonomy is extracted; it does not specify a complete formal ontology or executable relation schema.

Why combine top-down and bottom-up?

Top-down

Start with a known ontology or graph and assign inputs to existing concepts. You gain shared identities and established labels, but the inventory can be too coarse or outdated.

Bottom-up

Build concepts from the observed inputs. You can preserve new specialties, but repeated batches may invent duplicate names and inconsistent groupings.

The hybrid approach makes the external graph an anchor, not a ceiling. An initial mapping can be useful but too broad. The later curation step asks a different question: does this raw declaration really belong inside that canonical node?

The crucial distinction: “related to” does not mean “identical to.” A specialized skill may be discoverable through a broad skill, yet still deserve its own identity.
02 · The method

Five stages, with a feedback loop

The implementation uses Gemini 1.5 Flash with temperature 0 and structured output. Here, “multi-agent” means separate prompted roles for linking, naming, reviewing, and merging. “Reflexion” is a loop that revisits rejected assignments using reasons and proposed replacement labels—not a guarantee that the model’s judgment becomes correct.

1ReconcileChoose an external anchor.
2CanonicalizeName the clustered concept.
3CurateSeparate aliases from outliers.
4ConsolidateMerge duplicate concepts.
5RecoverRoute orphans into another epoch.

1. Reconciliation: context before identity

For each raw string, the system retrieves the top ten Wikidata candidates and their multilingual metadata. It adds the top fifteen co-occurring peer skills and top five professional categories from Malt’s profile graph. The LLM selects the best QID or QID combination and returns confidence, a concise rationale, and flags such as is_skill and is_compound.

Why add those neighbors? A short, ambiguous word rarely explains itself. Nearby skills and professional categories provide evidence about the intended field. Retrieval narrows the identity choices; context helps rank them. A model can still choose the wrong candidate, especially if the correct one never entered the candidate list.

The appendix also distinguishes direction-sensitive combinations. “English → French translation” is not interchangeable with “French → English translation.” Sorting all component identifiers into an unordered set would lose information.

2. Canonicalization: identity first, names second

Validated mentions are grouped by their unique QID combinations. The naming role creates one preferred label per target language, using a fallback order: professional, frequent platform wording first; an official Wikidata title second; a generated label only when necessary.

Each localized label carries provenance. The text describes malt, wikidata, and a generative origin; the appendix uses gemini for the last one. This is label-source metadata, not proof that the underlying concept or all its relations are correct.

3. Active curation: an initially plausible mapping can fail

The review role compares each raw member with its node’s preferred concept. The main text lists seven rejection criteria: ambiguity, specialization, semantic mismatch, non-skill, methodology, context, and sub-task. Rejections other than non-skills are grouped under NODE_CURATION for tracking.

The paper’s running example maps “gestion de projet web” to the broad project-management anchor. Curation then rejects that assignment as a specialization and proposes “Web Project Management.” Accepted members stay in the core node; meaningful outliers enter an orphan queue. The core label is rewritten using only accepted members. A node whose rejection rate exceeds 50% is flagged for human review.

A subtlety in the appendix: the curation prompt defines equivalence partly by hiring relevance and explicitly favors inclusion. It keeps translations, typos, proficiency modifiers, job-title variants, and versions, while separating domain specializations. That is a search-indexing policy, not pure logical equivalence. Your choice of “same enough” changes the resulting graph.

4. Consolidation: local batches need global identity

The first batch seeds a baseline graph. Subsequent incoming nodes are compared with that growing baseline. Shared raw members or low edit distance between labels nominate candidate pairs; the LLM decides MERGE or KEEP_SEPARATE from their metadata. A merge records a surviving ID, and decisions are cached to avoid repeated work.

This avoids asking the LLM about every possible pair. With N nodes, an exhaustive undirected comparison requires N(N−1)/2 pairs. The heuristics reduce the candidate set, but also risk missing duplicates with very different spellings. The appendix protects specialization boundaries and fundamentally different software versions; it prefers a Wikidata ID over a synthetic ID when a merge is valid.

5. Orphan recovery: turn the mismatch into a new hypothesis

The main text derives a synthetic orphan identifier from the suggested preferred label. Inputs with the same suggested label are routed together in the next epoch. The loop repeats until the orphan queue is empty. In the intended outcome, “Web PM” and “gestion de projet web” become one specialized node rather than two fragments or one overly broad core node.

Alias merging versus specialization preservationTwo multilingual aliases point to Project Management. Two web-related aliases point to a separate Web Project Management concept, which has a narrower-than connection to Project Management.Gestion de projetProjektmanagementProject ManagementWeb Project ManagementWeb PMGestion de projet webnarrower than

Conceptual teaching diagram, not the paper’s formal graph schema. Side arrows mean “names this concept”; the vertical arrow preserves the specialization boundary.

03 · Interactive explanation

Watch aliases converge—and specializations survive

This browser-only toy uses a tiny, explicit lookup table in place of Wikidata retrieval and LLM judgments. It is an illustration of the control flow, not a reproduction or a benchmark. Unknown text is held for review rather than assigned an invented identity.

Recognized examples are those above, plus “Project Manager,” “Web project delivery,” “Python,” and “Python 3.” Edit the list to test duplicates or an unknown skill.


“Drift” exposes a weakness of label-derived IDs: equal meanings need not receive equal names. The toy merger recognizes web aliases explicitly; real systems must find candidate pairs first.

Routing trace

Three experiments to try

  1. Hybrid, stage 2 → stage 3: web mentions first share a broad anchor, then leave its core group. Rejection here means “wrong identity assignment,” not “worthless input.”
  2. Hybrid, stage 5, drift on, consolidation off: equivalent web aliases fragment into different synthetic nodes. Turn consolidation on to recover one node.
  3. Compare the baselines at stage 5: anchor-only loses specialization; fragment-only preserves wording but does not unify identity. These baselines are explanatory caricatures, not evaluated methods from the paper.
04 · What was measured

Large vocabulary, partial validation

The evaluation uses proprietary Malt declarations, not a public benchmark. The authors report a hand-annotated Wikidata gold standard curated by five domain experts without overlap. Crucially, they say the current evaluation covers initial retrieval and pre-consolidation phases; final post-consolidation behavior has not yet been formally benchmarked.

36,037

Raw expertise strings

27,743 resolved; 8,294 unmapped. The resolved share is approximately 77%.

13,298

Canonical skill nodes

15,010 initial semantic groupings were streamlined into this node inventory.

66,490

Preferred labels

Exactly five per node: French, English, German, Dutch, and Spanish.

Compression: 1 − 13,298 / 27,743 = 52.1%. The denominator is the mapped vocabulary, not all raw strings. The average is 2.08 mapped variations per canonical node. Compression measures consolidation, not whether every merge was semantically appropriate.

Gold-standard reconciliation

Alignment coverage
79.7%
Found coverage
84.9%
Wrong guess rate
19.1%

Paper-reported global metrics. “Found coverage” is described as precision among inputs for which retrieval was attempted. These metrics use different definitions; do not treat them as parts of one pie.

Usage is highly concentrated

Top 1,000 nodes: 82.74% of usage volume.
Top 5,000 nodes: 97.25% of usage volume.

The authors distinguish the lexical long tail (redundant spellings) from the semantic long tail (rare capabilities). A useful normalizer should compress the former without deleting the latter.

Domain differences matter

DomainAlignment coverageFound coverage
Video Games88.1%91.8%
Industrial Engineering86.0%89.9%
Data & Analytics83.5%88.7%
Tech / Software81.7%86.8%
Marketing76.7%82.4%
Communication75.4%81.0%

Source: paper Table 1. Structured fields have stronger reconciliation scores than fields with softer boundaries and shifting terminology. The paper does not provide domain sample sizes or uncertainty intervals here.

Provenance and curation

Across preferred labels, the reported origins include 67.28% platform usage and 22.08% Wikidata titles. After accounting for overlap, their union is 80.65%; the remaining 19.35% is generated. The two source percentages should not simply be added. Also, calling 80.65% of the “graph” factual is stronger than the demonstrated measurement, which tracks label provenance.

Rejection categoryShare of all rejected inputs
SCORE_REJECTED31.9%
NODE_CURATION28.1%
NOT_A_SKILL13.7%
WIKIDATA_NOT_FOUND11.4%
NOT_IN_NODE10.7%
ERROR_NOT_KNOWN4.2%

Source: paper Table 2. The 13.7% figure is a share of rejections, not of all inputs. The rejection mix also shows that “unmapped” is not synonymous with “not a skill.”

Qualitative audits report 0.01–0.06 manually removed outliers per sub-graph. Audit size and the graph-level aggregation are not detailed enough to turn this into a general accuracy claim. The taxonomy is already integrated into candidate matching, but no measured matching-quality lift is reported.

What the evidence supports: substantial multilingual normalization and reasonably strong initial linking on this platform’s data. What it does not yet establish: superiority over other graph-generation methods, end-to-end convergence correctness, downstream hiring outcomes, or production cost efficiency.
05 · Critical reading

“Self-healing” is a design ambition, not a theorem

External grounding reduces—not eliminates—errors

A valid QID can still be the wrong QID. JSON schemas constrain formatting, not truth. The reported 19.1% wrong-guess rate is a concrete reason not to read the paper’s “prevent hallucinations” language literally.

Hashes stabilize strings, not meanings

A deterministic hash guarantees the same string yields the same identifier. It cannot guarantee the model uses the same string for synonymous skills. Renaming a concept can also change a label-derived ID unless there is an identity registry.

An empty queue is not proof of correctness

The paper gives no formal convergence proof, epoch limit, or oscillation analysis. A loop can stop after making a bad merge; it can also keep splitting and renaming. Termination and semantic quality need separate tests.

Coverage is not linguistic fidelity

Five labels per node demonstrate complete locale fields, not perfect translations or preservation of regional distinctions. The authors explicitly flag English-centric normalization and gender-marked occupational terms as future concerns.

Details that an implementer must resolve

The authors’ future-work plan includes comparisons with ESCO and other AI methods, testing consolidation and human review, measuring tokens and latency, and analyzing failure patterns and manual intervention rates. Those are precisely the missing checks needed to evaluate the full architecture.

06 · A small implementation model

The kernel is routing, not a giant prompt

The following pseudocode compresses the paper’s five roles into an implementable outline. It is a teaching sketch; the paper’s appendix supplies actual prompts, not a complete public implementation or its proprietary dataset.

baseline, orphan_queue, merge_cache = load_state()
for epoch in bounded_epochs:              # bounded: added safety measure
    incoming = []
    for raw in next_batch(orphan_queue):
        candidates = retrieve_wikidata(raw, k=10)
        context = profile_neighbors(raw, skills=15, categories=5)
        proposal = link(raw, candidates, context)  # schema-validated
        route_invalid_or_uncertain(proposal)
        incoming.append(proposal)

    nodes = group_by_identity_and_direction(incoming)
    nodes = name_in_locales(nodes, ["fr", "en", "de", "nl", "es"])
    core, orphans = curate_members(nodes)
    flag_for_review(core, rejection_fraction_above=0.5)

    for node in core:
        for target in candidate_duplicates(node, baseline):
            decision = cached_or_judge(node, target, merge_cache)
            apply_merge_or_keep(decision)
        persist(node)

    orphan_queue = route_by_suggested_label(orphans)
    if not orphan_queue: break
# Queue exhaustion is operational completion, not semantic validation.

A minimal teaching record

{
  "id": "local:web-project-management",
  "anchor": "external:project-management",
  "preferred_labels": {
    "en": {"text": "Web Project Management", "source": "generated"}
  },
  "raw_members": ["Web PM", "gestion de projet web"],
  "proposed_relation": "narrower_than",
  "curation_reason": "SPECIALIZATION",
  "status": "proposed"
}

Illustrative local schema, not the paper’s exact output. Only one language is shown for brevity. A local identity avoids claiming an unverified Wikidata identifier; “proposed” distinguishes model hypotheses from accepted graph facts.

Tests worth writing before automation

  1. Identity invariance: translations and harmless typos converge across batch orders; direction-reversed translation skills do not.
  2. Granularity preservation: a domain specialization remains distinct while retaining an explicit broader connection.
  3. Merge safety: aliases merge, Java and JavaScript do not, and fundamentally different framework versions remain separate.
  4. Loop stability: track orphan count, repeated assignments, renamed IDs, and held-for-review cases per epoch.
  5. Evidence separation: evaluate linking, member curation, localized names, duplicate detection, and relation correctness individually before testing downstream matching.
Suggested hardening, beyond the paper: immutable concept IDs with label history, reversible merges, versioned candidate-pair caches, explicit abstention, and a maximum iteration budget make the refinement loop inspectable. They should not be confused with features already demonstrated by this study.
07 · Takeaway

The useful move is to preserve the mismatch

A broad external anchor gives new text a semantic foothold. Curation then distinguishes genuine aliases from concepts that the anchor inventory represents too coarsely. The latter become structured hypotheses for graph growth. That combination is the paper’s most teachable contribution—even though the accuracy, stability, and cost of its complete consolidation loop still require evaluation.

One connection to context platforms: this paper offers a concrete pattern for evolving a domain concept layer over noisy source text: keep shared external identities where appropriate, preserve novel distinctions locally, and retain source metadata rather than silently replacing the originals.