Skip to content

SEO

How does Google process information from Wikipedia in Knowledge Graph?

Read the articleQuestions and answers

Article cover: How does Google process information from Wikipedia in Knowledge Graph?

One of the biggest challenges facing Google when it comes to semantic search is the identification and retrieval of entities, their attributes and other information from data sources such as websites. Such information usually does not have the right structure and is not free from errors. The current Knowledge Graph (Knowledge Graph), which is a kind of “semantic centre of Google”, is largely based on structured content from Wikidata, Wikimedia and Wikipedia.

Processing semi-structured data

Semi-structured data are information that is not explicitly marked up in accordance with general markup standards such as RDF or schema.org. They have an implicit structure from which structured data can be retrieved via a workaround.

Retrieving information from semi-structured data sources can be carried out with the help of a pattern-based extractor. It is able to identify sections based on a repeating structure of the same components from which information is retrieved.

Data exploration sources for Google Knowledge Graph presented by Olaf Kopp, Aufgesang GmbH
Data exploration sources for Google Knowledge Graph presented by Olaf Kopp, Aufgesang GmbH
Technology Processing semi-structured data
  1. 01Semi-structured dataNo standards, implicit structure
  2. 02Pattern-based extractorIdentification of repeating structure
  3. 03Information retrievalExtraction and organisation of data

A workaround method for retrieving structured information from unstructured sources.

Processing semi-structured data from Wikipedia

Retrieving information from semi-structured data sources is carried out on the basis of a special pattern-based extractor. It is able to identify content sections based on a repeating structure and retrieve information from them.

Wikipedia is a very attractive source of information due to the similar, repetitive structure in every entry. In addition, the entries are regularly checked by editors. On top of that, Wikipedia is based on the MediaWiki CMS. As a result, the content is provided with basic tags and can be easily retrieved via XML, SQL or as html.

The structure of a typical Wikipedia article provides a model for classifying entities by category, identifying attributes and retrieving information for featured snippets and knowledge panels.

A very similar or even identical structure of a specific article on Wikipedia looks, for example, as follows:

The title of each Wikipedia article reflects the name of the entity. In the case of ambiguous titles, the type of entity is added to the title to clearly distinguish it from other entities with the same name but a different meaning.

The information box located in the top right-hand corner of a Wikipedia article provides structured data about a specific entity. Introductory content is quite often found in the knowledge panel for a given entity.

Internal links found on Wikipedia provide Google with information about which future topics or other entities are semantically related to a given entity.

Google’s use of Wikipedia special pages

Of course, Google use not only the main content of the article and its elements for their purposes, but also special pages:

  • List and category pages allow classifications based on entity classes and types.
  • Special pages make it possible to identify synonyms.
  • Pages with terminology explanations allow for the identification of multiple meanings.
Attribute retrieval Retrieving attributes from Wikipedia as a starting point
  1. 01START: Retrieving the attribute and valueAge: 43 → Entity: Human
  2. 02Uncontrolled importNo quality verification at this stage.
  3. 03Requirement: high qualityThe source document must be correct.
  4. 04Reliable source: WikipediaThe basis of important information.

It is crucial to use reliable sources such as Wikipedia to ensure data quality.

Retrieving attributes from Wikipedia as a starting point

The technology developed by Google is geared towards the continual acquisition of new facts and attributes relating to the entity. This method starts with obtaining the attribute and value from the page. For example, age is the attribute and the number 43 is the value. If, in addition, information appears that a given person is, for example, an actor, Google can determine that the entity is a human being.

During data acquisition by the importing module, no quality checks are carried out. That is why the source document must be characterised by high quality and accuracy. The value of the acquired information depends mainly on the source data. Therefore, reliable sources of information, such as Wikipedia, should be used.

How is information about entities collected?

At present, it appears that Google essentially acquires all information about entities from RDF-compliant structured data sources and from semi-structured sources such as Wikipedia.

In order to collect information such as attributes, types and classes of entities, and links to neighbouring entities, an entity profile must first be created. The profile is previously marked with the entity name and URL, which allows for unique assignment.

The profile is then supplemented with information about the given entity from various data sources. For example, information may come from Wikipedia or DBpedia, and the entity is linked to the same entity in Freebase.

Information from Wikipedia… Information from Wikipedia in featured snippets and knowledge panels
  1. 01Instant answersFast search results.
  2. 02Entity acquisitionIdentification of relevant information.
  3. 03Concise descriptionShort, substantive snippets.

Fast and precise delivery of knowledge in response to modern user queries.

Due to the increasing importance of voice search, modern search engines try to deliver results instantly, without the user having to browse through many results pages. To provide such fast performance, the meaning of the searched term needs to be determined and the relevant information must be acquired from structured and unstructured data sources.

Rank Math panel in the WordPress editor: Google result preview with title and description, keyword field and a list of basic SEO tests
Example The preview in the SEO panel shows the title and description as they may appear in the results — overly long texts are visible before the post is published. Rank Math in WordPress (local CMS), own screenshot

The solution to this problem is entity retrieval. Its task is to identify the relevant entities in the directory in response to a written or spoken query. They are then sorted in list form according to the level of match to the query. To provide an answer, a snippet is needed that briefly and concisely explains the given entity (Entity).

Such descriptions are known and observed in the form of snippets in search listings (featured snippets) and in the form of entity descriptions in knowledge panels. Most often they are acquired from Wikipedia or DBpedia. Sometimes, in featured snippets, information is acquired from unstructured data sources such as glossaries, blogs, magazines, etc. At this point, however, Google prefer descriptions from Wikipedia, and use other sources when Wikipedia offers no information and something has to be used to “fill the gap”.

In terms of snippets, Google trust descriptions from Wikipedia. One of the reasons is the clear structure of the content there. Such articles provide a concise and substantive description of specific topics.

How Google acquire information from unstructured content on websites for featured snippets is a matter of speculation. There are many different theories. Perhaps it involves focusing on the subject, predicate and object appearing in the paragraph.

The frequency with which information from Wikipedia is used in featured snippets suggests that Google are not yet satisfied with the results of acquiring unstructured data, or that attempts to manipulate the data are still beyond Google’s control.

Wikipedia as evidence of the existence of entities

The most reliable way of treating something as an entity is an entry in Wikipedia, Wikidata or submitting it to Google.

Google, however, reserve the right to review entries and remove them from databases if the Wikidata entry does not contain the appropriate reference sources. To prevent manipulation, the entry must be verified on the basis of at least 1/3 of the sources. A page in Wikipedia or a Wikimedia entry appears to be an important source here.

In Wikidata, Entity attributes are rather itemised, whereas Wikipedia describes Entity in detailed text. In other words, an entry on Wikipedia is a detailed description of an Entity and, as an external document, is an important source for Knowledge Graph (knowledge graph).

Wikipedia entries play a dominant role in many Knowledge Graph fields, as a source of information, and are used by Google together with Wikidata entries as evidence of Entity properties. Without an entry in Wikipedia or Wikidata, the Entity field or even the Knowledge Panel (knowledge panel) does not appear.

However, it is worth remembering that a Wikipedia entry is rejected in the case of most companies and people, because in the eyes of many Wikipedia users such entries have low social significance. A useful alternative is to create a profile in Wikidata. The frequency and repeatability of consistent data across various trusted sources make it easier for Google to identify Entities more accurately.

FAQ

Frequently asked questions

How does Google use Wikipedia to build Knowledge Graph?

Google extracts information from Wikipedia about entities, their attributes, classes and relationships with other entities. Wikipedia articles are an important source for Google because they have a repeatable structure and are easier to process.

Does Google trust Wikipedia more than other sources for Featured Snippets?

Yes, the article indicates that Google prefers descriptions from Wikipedia because they are concise and well organised. Other sources are used when Wikipedia does not provide the needed information.

Why is the structure of Wikipedia articles important for Google?

Because a typical article has a repeatable layout: the title matches the entity name, the infobox contains structured data, and the introduction often appears in the knowledge panel. Such a structure makes it easier to identify categories, attributes and relationships.

When does Google reach for information from special Wikipedia pages?

Google also uses pages with lists, categories, term explanations and other special pages. They are used for classification, identifying synonyms and distinguishing multiple meanings.

How does Google extract entity attributes from Wikipedia?

The process starts by retrieving an attribute and its value, for example age and the number 43. On this basis, Google can also infer the type of entity, for example that it is a person.

Can you have a knowledge panel in Google without an entry in Wikipedia?

The article suggests that an entry in Wikipedia or Wikidata is very important for the Entity and Knowledge Panel to appear. Without such entries, this element usually does not appear.

Contents