One of the problems associated with internet search engines such as Google is duplicate content, that is content duplication. The concept of duplication means that similar content appears in many places (under different URL addresses) across the web. This can have an adverse impact on a website’s rankings, and the problem worsens when users start linking to different versions of the same content. It is therefore worth familiarising yourself with the various cases of content duplication and finding a solution for each of them.
Content duplication – what is it?
Duplicate content is content available under multiple URL addresses across the web. Because one address shows the same content as another, search engines do not know which of these addresses should appear in search results. It is above all worth focusing mainly on the technical causes of content duplication and the solutions to this problem.
DiagramThe colour variant has its own address with a parameter, but the canonical link points to the main product page — the signals from both addresses are consolidated on one URL.Source: Google Search Central, CC BY 4.0
Content duplication is partly like stopping at a crossroads where signposts point in two different directions leading to the same destination. So which route should you choose? The reader only wants an answer to the question asked, but the search engine has to choose which page should appear in search results, because it is obvious that it does not want to show the same content twice.
Let’s assume that an article about “keyword x” appears at the following addresses:
Next, let’s assume that the article has been picked up by many bloggers and some of them link to the first URL address, while the others link to the second. At this point, the search engine problem shows its true nature.
Duplicate content is a problem because the links mentioned promote different URL addresses. If all links led to the same URL, then the chances of ranking for the phrase “keyword X” would increase.
SEO blogContent duplication – what is it?
01One piece of content, different addressesAvailable under multiple URLs
02The search engine is lostIt does not know which URL to choose
03A duplicated destinationWhich route is the right one?
Main challenge: search engines must decide which version of the same content to show in results, which is like choosing one road at a crossroads leading to the same goal.
Different types of content duplication
There are two types of content duplication:
internal content duplication – occurs when one main domain creates duplicate content across many internal URL addresses (on the same website).
External content duplication – also known as cross-domain duplication, occurs when two or more different domains have the same copy of a page indexed by search engines.
Both internal and external content duplication can occur in exact or near-exact form.
An example of using Screaming Frog and the Near Duplicates option to analyse internal duplication
Why should you prevent content duplication on a website?
Content duplication has a negative impact on rankings. Search engines do not know which page to suggest to users. As a result, all pages seen by search engines as duplicates are, at best, at risk of a drop in rankings.
If problems with duplicate content are very serious, for example if only a small amount of original content goes hand in hand with word-for-word copied content, then you can even expect action from Google in connection with an attempt to deceive users. If content is to achieve a high position, it is important to make sure that each page offers an adequate amount of unique content.
This is not only a search engine problem. If users are looking for a specific page, it may be frustrating for them that they cannot find what they are looking for. Therefore, just like in many other aspects of SEO, it is important to make sure that duplicate content issues are resolved so that the user experience is positive.
SEO & content strategyWhy should you prevent duplicate content on the site?
01Search engine confusionDifficulty choosing a page
02Ranking dropRisk of losing positions
03Risk of Google actionPotential penalties for deception
04Poor user experienceFrustration at the lack of results
Summary: unique content is key to high rankings and user satisfaction, protecting against penalties and ranking drops.
Is duplicate content detrimental from an SEO perspective?
Officially Google does not impose penalties for duplicate content. However, the search engine filters identical content with the same meaning. The effect of this may be a deterioration in a site’s position in Google rankings (positions in the SERP).
Duplicate content confuses Google and forces the search engine to decide which of the identical pages should be ranked in the top results. Regardless of who created the content, there is a high probability that the original, first page will not be selected to appear in the top results.
This is one of the reasons why duplicate content is bad from an SEO perspective. There are some other obvious reasons why duplicated content is bad.
John Muller on duplicate content and its potential value
How do you check a site for duplicate content?
If you have content-rich pages that do not rank highly in search results, you should check whether the content has not been copied and used on another website. This can be done in several ways.
SEO & contentHow do you check a site for duplicate content?
01Choose a passageCopy a few sentences from the content.
02Search in quotesCheck the exact match in Google.
03Use CopyscapeA free tool for domain analysis.
Unique content is key to high positions in search results.
Exact Match search
Copying a few sentences of content from one of your websites, which should then be put in quotation marks and searched in Google. By using quotation marks, you indicate to Google that you want to receive exactly that text. If many results appear, it means that someone has copied that content.
Implementation of quotation marks on the example of a poetic work by Adam Mickiewicz
Copyscape
Copyscape is a free tool that checks site content for duplication on other domains. If the text from a site has been copied, then the URL infringing it will appear in the results.
Identifying duplicate content issues
You may not always be aware of duplicate content on a site or in the content. Using Google is one of the simplest ways to detect duplicated content.
There are many search operators that are helpful in such situations. If you want to find all the URL addresses on your site that contain an article with the keyword X, enter the following search query in Google:
site:example.com intitle:”Keyword X”
Google will then show subpages on the example.com site that contain such a keyword. The more specific the intitle part of the query is, the easier it will be to filter out duplicate content. The same method can be used to check for content duplication across the web. Let’s assume the article title was “Keyword X – why it is excellent”. In that case, you should search for:
intitle:”Keyword X – why it is excellent”
Google will present all pages matching this title. Sometimes it is even worth checking pages for one or two full sentences from the article, because some people who copy content may change the title. In some cases, when carrying out such searches, Google may display a notice on the last results page.
Example of using the site intitle operator on the example of the keyword “winter tyres”
Practical solutions for duplicate content
Once you have decided which URL is the canonical URL of the content, you should begin this process. This means informing search engines about the canonical version of the page and allowing them to find it as quickly as possible.
Avoiding duplicate content
Some of the above causes can be fairly easy to fix:
Are session IDs present in the URL addresses?
They are usually displayed in the system settings.
Is comment pagination being used in WordPress?
This function should be deactivated.
Are there issues related to the WWW prefix and its absence?
You should choose one of the options and stick to it by means of redirects.
Is indexing of addresses with a query string blocked on the site?
It is worth noindexing addresses related to filtering, sorting or the internal search engine. Unless we have some specific goal for such an entity to exist in the index.
FAQ
Frequently asked questions
01How does duplicate content affect a page’s ranking in Google?
Search engines may not know which URL to choose, which puts duplicate pages at risk of losing positions. In serious cases, Google may also treat such a problem as an attempt to mislead users.
02Does Google penalise duplicate content?
Officially, Google does not impose penalties for content duplication. However, it does filter out identical content, which can reduce pages’ visibility in search results.
03What is internal content duplication?
This is a situation where the same content appears under multiple URL addresses within one domain. The issue concerns one website, but many of its versions.
04How does external content duplication differ from internal duplication?
External duplication occurs when identical content is found on two or more different domains. Internal duplication concerns multiple URL addresses within the same domain.
05How can you check whether content has been copied to another site?
You can paste a few sentences from the article in quotation marks and search for them in Google as an exact match. The free Copyscape tool is also helpful, as it shows whether the text appears on other domains.
06Which technical issues are worth fixing to reduce content duplication?
It is worth choosing one canonical URL and informing search engines about it. You should also avoid, among other things, session IDs in URLs, comment pagination in WordPress, www version issues, and indexing URLs with a query string when filtering and sorting.