Knowledge Graphs

A knowledge graph stores entities and typed relations as a graph. In search, it supplies structure that text alone cannot: aliases, entity disambiguation, typed filters, relation traversal, and explanations. It is a natural partner for graph-based retrieval and hybrid search.

Triples and graphs

RDF-style knowledge graphs use triples:

A SPARQL-like graph pattern retrieves subjects that satisfy joins over triples. For example, “papers that use BM25 and are evaluated on MS MARCO” is a two-edge conjunction over the same subject.

Worked example

The table below is a tiny graph: each row is one subject-predicate-object edge. The example then asks for a conjunctive graph pattern rather than a keyword match.

SubjectPredicateObject
paperAusesbm25
paperAevaluated_onmsmarco
paperBusesdense_retrieval
paperBevaluated_onmsmarco

A graph query for “papers that use BM25 and are evaluated on MS MARCO” requires both edges to share the same subject:

SELECT ?paper WHERE {
  ?paper uses bm25 .
  ?paper evaluated_on msmarco .
}

The result is paperA. Lexical search could find both papers for MS MARCO, but the graph query enforces the relation pattern: the same paper must both use BM25 and be evaluated on the dataset.

Where it fits

Knowledge graphs help literature-management search systems connect papers, methods, datasets, authors, and notes. They also help Elasticsearch-style systems with entity enrichment: a document mentioning Robertson can link to the intended researcher, not just the surface string.

Caveats

Schema design is the hard part. Overly rigid ontologies slow ingestion; overly loose predicates become unqueryable. Entity resolution errors are costly because one wrong merge contaminates every downstream traversal. Keep provenance on edges so users can inspect why a relationship exists.

References