Research · Knowledge an agent can cite ·

Ontology-guided GraphRAG

Question answering over a knowledge graph whose only schema is one curated OWL ontology, which shapes what gets extracted, how the graph is typed and how an agent queries it.

Vector retrieval finds passages that look like the question. It cannot count, follow a relationship, or tell a rover from the mission that carried it. A knowledge graph can, but only if the graph has a schema that extraction obeys and querying understands. Without one, every document invents its own labels and nothing joins.

The approach here makes one curated OWL ontology the only schema. It is loaded into the graph, and everything else is derived from that copy: the schema the extraction model fills, the pruning that drops anything the ontology does not name, the superclass labels that make hierarchy queries work, and the schema a model writes queries against. Change the ontology and the change reaches extraction and querying with no code change.

The working version is a small fast start that runs locally in Docker or on Kubernetes, against Neo4j either self-hosted or as a managed service.

The parts

Component Role
Ontology A curated OWL file under version control: classes, hierarchy, and relationships with domain and range
Document store The source documents, in object storage or a local folder
Ingest job Loads the ontology, derives the extraction schema, extracts, embeds and writes the graph
Graph The ontology as classes, plus extracted entities, text chunks and their vectors, in Neo4j
Chat service Answers by writing a graph query from the ontology’s schema, or by vector search plus a hop into the graph
Models Claude for extraction, queries and answers; Titan for embeddings

How the ontology does its work

Publishing. The ingest job loads the ontology into the graph as classes, subclass links, domains and ranges. On a self-hosted Neo4j this uses the n10s plugin. The managed service has no n10s, so the same structure is written from rdflib instead, and a test requires the two to produce identical graphs.

Extraction. The ingest job then reads the ontology back out of the graph, not from the file, and builds the extraction schema from it. The graph is the single source, so no two consumers can drift apart by reading different copies. Each document is split into chunks, and the model extracts entities and relationships against that schema. Anything whose type or relationship the ontology does not name is pruned before it is written. Every surviving entity links back to the chunk and document it was read from.

Typing the graph. After writing, each entity gains the labels of its superclasses as well as its own class. A question about spacecraft then finds rovers, orbiters and landers without anyone spelling out the hierarchy in the query.

Validation. A conformance query checks that every relationship respects its domain and range. SHACL shapes check what the class hierarchy cannot state, such as cardinality and value rules.

Retrieval. The chat service has two paths. In graph mode, the model writes Cypher against a schema generated from the ontology, runs it, and answers from the rows, showing the answer, the query and the rows. In vector mode, it searches chunk embeddings and hops from each chunk to the entities it mentions. Questions that need a join, a count or the hierarchy belong on the graph path. Lookups work on either.

Design decisions

The ontology lives in the graph. Reading the Turtle file separately in each consumer was rejected because it lets them drift.

A strict extraction schema. The extraction uses the neo4j-graphrag pipeline with its schema held strict. A tool whose ontology support is a model reading the file was rejected, because nothing then enforces the ontology.

Entity resolution guarded by the ontology. Merging duplicate entities is harder than it looks. Name-only resolution leaves one thing as two nodes. The pipeline’s fuzzy resolver at its default threshold, on a scripted test, merged Voyager 1 with Voyager 2 and folded a probe, an orbiter and a lander into one node, because names alone score distinct things nearly as high as true variants. The chosen approach resolves after superclass labels are added, so a Rover and a Spacecraft of the same name can meet, and lets the ontology veto a merge: only compatible classes (one a superclass of the other) may merge, and never names that differ by a number. Merged nodes keep the longest name, with the others as aliases, after a test showed a short name such as “Webb” defeating a later name match.

Read-only answers. The chat cannot change the graph. A filter refuses any query that writes, and the query then runs in a read-only transaction as a second barrier.

Infrastructure. Terraform holds the platform (network, cluster, permissions, registries, secrets) and Kustomize holds the workloads. Helm was rejected in place of Terraform because it cannot create a cluster, permissions or secrets, and is not yet needed for two workloads. Model access uses workload identity, so no model keys sit in the cluster.

Self-hosted or managed

Managed Neo4j (Aura) Self-hosted Neo4j Community
Ontology loader rdflib in the ingest job n10s plugin in the image
SHACL validation Not in the database; would need pySHACL over an export Through n10s after each ingest
Cost floor From about USD 65 per GB a month on the professional tier One node and a small volume on the cluster
Network Public endpoint with TLS; private networking on higher tiers Inside the cluster
Backups and availability Managed None; Community has no clustering

The images and manifests are the same either way. One setting chooses the store.

What was measured

These results are on small corpora and come from single runs, so they are indications, not settled findings.

On a fast-start corpus of 11 public documents about space missions and 19 golden questions, run once per graph with Claude Sonnet 5.5:

Question type Graph mode Vector mode
Needs a join or the class hierarchy (10 questions) 9 and 10 passed (two graphs) 3 and 2 passed
Simple lookup (9 questions) 9 and 8 passed 9 and 8 passed

Vector mode failed by retrieval: four chunks cannot hold every mission flown on a given launch vehicle. On that corpus, every extracted relationship conformed to its domain and range, every class had instances, and SHACL raised only warnings about missing optional links. An ingest cost about USD 0.59 (22 model calls, about 201,000 input tokens). Two extraction runs differed in small ways, so resolvers are compared over one extraction rather than across runs.

A second test, on a subset of a published financial-industry ontology, showed the same conformance under a real model (82 of 82 relationships). A strict match against hand-written gold data found 60 of 74 entities, most misses being choices of name or class rather than invented facts.

Where it fits

It fits questions that need joins, counts or class hierarchies over facts scattered across documents, in a domain with an agreed vocabulary or one worth agreeing. It suits a team that wants the graph’s meaning under version control and reviewed like code.

It does not fit pure similarity search over documents with no shared entities, where vectors alone are cheaper. Data already in tables is better queried as tables, or through a virtual graph over them. Workloads that need SPARQL or RDF named graphs end to end belong in an RDF store.

Ontologies from elsewhere

Ontology induction is out of scope: the ontology is curated by a person. Tools that induce a draft ontology from tables can still feed this approach. The draft is exported as Turtle, a person curates it into the master, and the ingest job loads it as usual. Porting such a tool’s own RDF store to Neo4j was weighed and rejected, because Neo4j has no SPARQL endpoint and the tool’s RDF side would have to stay where it was, so nothing would be saved.

Open questions

The central claim, that graph mode beats vector retrieval on join and hierarchy questions, holds on one small corpus and needs repeat runs and a larger one. Acronyms score poorly against their full names, so aliases need an alias list in the ontology, which is not built yet. Per-user entitlements on graph facts are not in scope: anyone who can reach the chat sees the whole graph.