Knowledge graph
Imagine a vast, interconnected web of facts, a digital brain that understands not just words, but the relationships between them. This is the essence of a knowledge graph—a powerful technology that transforms raw data into meaningful, interconnected insights. It's the silent architect behind many of the smart systems we use every day, making our digital world remarkably more intelligent. Knowledge graphs represent information as interconnected entities and their relationships, offering a deeper understanding of data's meaning. They are the backbone of modern AI, search engines, and recommendation systems, powering intelligent interactions and discoveries. Advances in machine learning and large language models are continually expanding their capabilities, from scientific research to dynamically organizing information.
AI Summary
Imagine a vast, interconnected web of facts, a digital brain that understands not just words, but the relationships between them. This is the essence of a knowledge graph—a powerful technology that transforms raw data into meaningful, interconnected insights. It's the silent architect behind many of the smart systems we use every day, making our digital world remarkably more intelligent.
- Knowledge graphs represent information as interconnected entities and their relationships, offering a deeper understanding of data's meaning.
- They are the backbone of modern AI, search engines, and recommendation systems, powering intelligent interactions and discoveries.
- Advances in machine learning and large language models are continually expanding their capabilities, from scientific research to dynamically organizing information.
What is a Knowledge Graph?
At its core, a knowledge graph is a specialized kind of database that uses a graph structure to store and manage information. Instead of rigid tables, it organizes data like a network, with individual pieces of information—called 'entities'—connected by lines that define their 'relationships'.
More Than Just Data
What makes a knowledge graph truly powerful is its focus on semantics—the meaning and context behind the data. It's not just storing 'Paris' and 'France'; it's storing 'Paris is the capital of France,' capturing the actual relationship between these entities. This rich web of meaning allows machines to understand connections much like humans do.
Ubiquitous Technology
You interact with knowledge graphs constantly, perhaps without even realizing it. They power the sophisticated results you get from Google and Bing, answer your questions via Siri or Amazon Alexa, and even help social networks like LinkedIn and Facebook understand your connections and interests.
Expanding Horizons
Initially a cornerstone of search and recommender systems, knowledge graphs are now at the forefront of data science and machine learning. They're invaluable in complex fields like genomics and proteomics, where understanding intricate biological relationships is key to scientific discovery.
A Legacy of Understanding
The concept of a 'knowledge graph' isn't as new as it might seem. The term itself was first used in 1972 by linguist Edgar W. Schneider, laying early groundwork for organizing information modularly. By the late 1980s, projects were already exploring semantic networks that used restricted relationships to facilitate complex data operations.
Early Topic-Specific Graphs
Some of the earliest practical applications focused on specific domains. In 1985, WordNet began mapping semantic relationships between words, essentially building a knowledge graph of language itself. GeoNames followed in 2005, capturing the intricate connections between geographic locations and entities.
General Knowledge Repositories Emerge
A significant leap occurred in 2007 with the creation of DBpedia and Freebase. DBpedia automatically extracted structured data directly from Wikipedia, while Freebase aggregated public datasets from various sources. Though not explicitly called 'knowledge graphs' at the time, these projects were foundational to the development of large-scale, general-purpose knowledge bases.
Google's Game-Changer
The term truly entered the mainstream in 2012 when Google unveiled its own 'Knowledge Graph.' Building upon resources like DBpedia and Freebase, along with structured data from countless websites, Google used this graph to enhance search results, offering direct answers and related facts beyond simple keyword matching. Its popularity online cemented the term in public consciousness.
Industry-Wide Adoption
Following Google's lead, numerous tech giants like Facebook, LinkedIn, Amazon, and Microsoft began widely advertising their use of knowledge graphs. These sophisticated structures became essential for understanding user behavior, personalizing experiences, and managing vast amounts of interconnected data across their platforms.
LLMs and Dynamic Graphs
More recently, the rise of large language models (LLMs) has sparked renewed interest. Knowledge graphs are now seen as a way to give LLMs structured information, improving their reasoning and factual accuracy. Cutting-edge systems like Microsoft Research's GraphRAG even integrate LLM-generated graphs to enhance how they retrieve and summarize information.
What Makes a Knowledge Graph?
Despite their widespread use, there isn't one single, universally accepted definition for a knowledge graph. However, most experts agree on a set of core characteristics that distinguish them from simpler databases, especially when viewed through the lens of the Semantic Web.
Key Characteristics
Flexible relations among knowledge: They define abstract classes and relationships, describe real-world entities, allow arbitrary interconnections, and cover diverse domains. General structure: They are networks of entities, their semantic types, properties, and relationships, often using categorical or numerical values for properties. Supporting reasoning: They integrate information into an ontology and apply a 'reasoner' to derive new, implicit knowledge that wasn't explicitly stored.
A Simpler View
For those seeking a more straightforward understanding, consider this: a knowledge graph is essentially a digital structure that represents knowledge as concepts and the intricate relationships between them, or 'facts.' Critically, it can include an 'ontology'—a formal framework of concepts and categories—that allows both humans and machines to understand and logically reason about its contents.
Beyond Storage: Reasoning and Inference
One of the most profound capabilities of a knowledge graph is its ability to go beyond merely storing data. By formally representing semantics and often incorporating ontologies as a schema layer, knowledge graphs enable logical inference. This means they can deduce implicit knowledge—facts not directly stated—rather than just answering queries for explicit knowledge.
Knowledge Graph Embeddings
To bridge the gap between structured knowledge graphs and the world of machine learning, researchers developed 'knowledge graph embeddings.' These are latent feature representations—numerical vectors—that capture the meaning of entities and relationships within the graph. Just as 'word embeddings' allow computers to understand the context of words, knowledge graph embeddings enable machine learning models to process complex relationships.
Graph Neural Networks (GNNs)
The magic behind generating these powerful embeddings often lies with Graph Neural Networks, or GNNs. These deep learning architectures are uniquely suited to process graph-structured data. They learn by analyzing the connections between nodes (entities) and edges (relationships), making them adept at predicting values, discovering patterns, and ultimately, enabling sophisticated knowledge graph reasoning.
The Challenge of Entity Alignment
As knowledge graphs proliferate across various fields, a significant challenge arises: how do we know if 'Barack Obama' in one graph is the same 'Barack Obama' in another, differently structured graph? This is the problem of 'entity alignment'—identifying corresponding entities across disparate knowledge graphs, a crucial step for integrating vast amounts of information.
Strategies for Alignment
Solving entity alignment involves sophisticated techniques. Researchers look for similar substructures within graphs, analyze semantic relationships, compare shared attributes, or combine all three. The recent breakthroughs in large language models are also proving invaluable, as their ability to generate meaningful linguistic embeddings can greatly assist in matching entities.
Virtual vs. Stored KGs
Knowledge graphs can exist in different forms. Some are stored in specialized 'graph databases' like Neo4j, optimized for network-like data. Others, known as 'virtual knowledge graphs,' don't store data directly but instead provide a graph-based view over existing relational databases or data lakes, dynamically generating graph responses to queries based on predefined mappings. Each approach has its strengths, catering to different data management needs.
Article
Knowledge graph
Example conceptual diagram
In knowledge representation and reasoning, a knowledge graph is a knowledge base that uses a graph-structured data model or topology to represent and operate on data. Knowledge graphs are often used to store interlinked descriptions of entities – objects, events, situations or abstract concepts – while also encoding the free-form semantics or relationships underlying these entities.
Since the development of the Semantic Web, knowledge graphs have often been associated with linked open data projects, focusing on the connections between concepts and entities. They are also historically associated with and used by search engines such as Google, Bing, and Yahoo; knowledge engines and question-answering services such as WolframAlpha, Apple's Siri, and Amazon Alexa; and social networks such as LinkedIn and Facebook.
Recent developments in data science and machine learning, particularly in graph neural networks, representation learning, and machine learning, have broadened the scope of knowledge graphs beyond their traditional use in search engines and recommender systems. They are increasingly used in scientific research, with notable applications in fields such as genomics, proteomics, and systems biology.
History
Knowledge graph
In December 1968, the Human Resources Research Organization (HumRRO) delivered a presentation for the Society for General Systems Research's Annual Meeting in which they described the "Subject-Matter Map," and formalized the idea of representing knowledge as a graph. In 1972, the Austrian linguist Edgar W. Schneider cited HumRRO's presentation in his discussion of how to build modular instructional systems for courses, and coined the term "Knowledge Graph." In the late 1980s, the University of Groningen and University of Twente jointly began a project called Knowledge Graphs, focusing on the design of semantic networks with edges restricted to a limited set of relations, to facilitate algebras on the graph. In subsequent decades, the distinction between semantic networks and knowledge graphs was blurred.
Some early knowledge graphs were topic-specific. In 1985, Wordnet was founded, capturing semantic relationships between words and meanings – an application of this idea to language itself. In 2005, Marc Wirk founded Geonames to capture relationships between different geographic names and locales and associated entities. In 1998, Andrew Edmonds of Science in Finance Ltd in the UK created a system called ThinkBase that offered fuzzy-logic based reasoning in a graphical context.
In 2007, both DBpedia and Freebase were founded as graph-based knowledge repositories for general-purpose knowledge. DBpedia focused exclusively on data extracted from Wikipedia, while Freebase also included a range of public datasets. Neither described themselves as a 'knowledge graph' but developed and described related concepts.
In 2012, Google introduced their Knowledge Graph, building on DBpedia and Freebase among other sources. They later incorporated RDFa, Microdata, JSON-LD content extracted from indexed web pages, including the CIA World Factbook, Wikidata, and Wikipedia. Entity and relationship types associated with this knowledge graph have been further organized using terms from the schema.org vocabulary. The Google Knowledge Graph became a complement to string-based search within Google, and its popularity online brought the term into more common use.
Since then, several large multinationals have advertised their use of knowledge graphs, further popularising the term. These include Facebook, LinkedIn, Airbnb, Microsoft, Amazon, Uber and eBay.
In 2019, IEEE combined its annual international conferences on "Big Knowledge" and "Data Mining and Intelligent Computing" into the International Conference on Knowledge Graph.
The development of large language models expanded interest in knowledge graphs as a way to structure information from unstructured text, with advances in language processing enabling their automatic or semi-automatic generation and expansion. The term knowledge graph has since broadened to include the dynamically constructed and adaptive graph structures, which support retrieval, reasoning, and summarization in generative systems. Microsoft Research's GraphRAG (2024) exemplified this development by integrating LLM-generated graphs into retrieval-augmented generation.
Definitions
Knowledge graph
There is no single commonly accepted definition of a knowledge graph. Most definitions view the topic through a Semantic Web lens and include these features:
• Flexible relations among knowledge in topical domains: A knowledge graph (i) defines abstract classes and relations of entities in a schema, (ii) mainly describes real world entities and their interrelations, organized in a graph, (iii) allows for potentially interrelating arbitrary entities with each other, and (iv) covers various topical domains. • General structure: A network of entities, their semantic types, properties, and relationships. To represent properties, categorical or numerical values are often used. • Supporting reasoning over inferred ontologies: A knowledge graph acquires and integrates information into an ontology and applies a reasoner to derive new knowledge.
There are, however, many knowledge graph representations for which some of these features are not relevant. For those knowledge graphs, this simpler definition may be more useful:
• A digital structure that represents knowledge as concepts and the relationships between them (facts). A knowledge graph can include an ontology that allows both humans and machines to understand and reason about its contents.
Implementations
In addition to the above examples, the term has been used to describe open knowledge projects such as YAGO and Wikidata; federations like the Linked Open Data cloud; a range of commercial search tools, including Yahoo's semantic search assistant Spark, Google's Knowledge Graph, and Microsoft's Satori; and the LinkedIn and Facebook entity graphs.
The term is also used in the context of note-taking software applications that allow a user to build a personal knowledge graph.
The popularization of knowledge graphs and their accompanying methods have led to the development of graph databases such as Neo4j, GraphDB and AgensGraph. These graph databases allow users to easily store data as entities and their interrelationships, and facilitate operations such as data reasoning, node embedding, and ontology development on knowledge bases.
In contrast, virtual knowledge graphs do not store information in specialized databases. They rely on an underlying relational database or data lake to answer queries on the graph. Such a virtual knowledge graph system must be properly configured in order to answer the queries correctly. This specific configuration is done through a set of mappings that define the relationship between the elements of the data source and the structure and ontology of the virtual knowledge graph.
Using a knowledge graph for reasoning over data
Knowledge graph
A knowledge graph formally represents semantics by describing entities and their relationships. Knowledge graphs may make use of ontologies as a schema layer. By doing this, they allow logical inference for retrieving implicit knowledge rather than only allowing queries requesting explicit knowledge.
In order to allow the use of knowledge graphs in various machine learning tasks, several methods for deriving latent feature representations of entities and relations have been devised. These knowledge graph embeddings allow them to be connected to machine learning methods that require feature vectors like word embeddings. This can complement other estimates of conceptual similarity.
Models for generating useful knowledge graph embeddings are commonly the domain of graph neural networks (GNNs). GNNs are deep learning architectures that comprise edges and nodes, which correspond well to the entities and relationships of knowledge graphs. The topology and data structures afforded by GNNs provide a convenient domain for semi-supervised learning, wherein the network is trained to predict the value of a node embedding (provided a group of adjacent nodes and their edges) or edge (provided a pair of nodes). These tasks serve as fundamental abstractions for more complex tasks such as knowledge graph reasoning and alignment.
Entity alignment
Two hypothetical knowledge graphs representing disparate topics contain a node that corresponds to the same entity in the real world. Entity alignment is the process of identifying such nodes across multiple graphs.
As new knowledge graphs are produced across a variety of fields and contexts, the same entity will inevitably be represented in multiple graphs. However, because no single standard for the construction or representation of knowledge graph exists, resolving which entities from disparate graphs correspond to the same real world subject is a non-trivial task. This task is known as knowledge graph entity alignment, and is an active area of research.
Strategies for entity alignment generally seek to identify similar substructures, semantic relationships, shared attributes, or combinations of all three between two distinct knowledge graphs. Entity alignment methods use these structural similarities between generally non-isomorphic graphs to predict which nodes correspond to the same entity.
In 2023, researchers found success in using large language models (LLMs) in the task of entity alignment. This was in particular thanks to their effectiveness at producing syntactically meaningful embeddings.
As the amount of data stored in knowledge graphs grows, developing dependable methods for knowledge graph entity alignment becomes an increasingly crucial step in the integration and cohesion of knowledge graph data.