Skip to content

Technical SEO

Crawl budget optimisation

At the kick-off I collect information about the type and structure of the site, its environments, and the process for implementing and publishing changes. We clarify the indexing priorities and the scope: domains, subdomains, language versions and page types. We also establish whether server logs are available and what the full set of materials for analysis is.

Area
Technical SEO
Process
3 stages
Scope
6 sections · 8 min
Quote
Free
About the service

Defining the scope and input data for optimising your site’s crawl budget

About the crawl budget optimisation service

A few service details
  • Setting indexing priorities
  • Defining the scope of domains and subdomains
  • Defining the site’s language versions
  • Defining page and URL types
  • Establishing server log availability
  • Agreeing the materials: sitemaps, robots
Work process

The crawl budget optimisation process: from scope to implementation

I start our cooperation with a kick-off and by collecting information about the site and the way changes are implemented. We then agree the indexing priorities and the materials available for analysis. After that I go through the diagnosis, the change plan and the implementation, verifying the settings as I go

  1. 01/ 03

    Kick-off and scope

    I collect data on the type of site, its structure, environments and publication process, and we agree which areas the analysis and actions cover.

  2. 02/ 03

    Data and priorities

    We clarify the domains, subdomains, language versions and URL types, and agree the indexing priorities and the availability of logs, sitemaps, URL rules and the robots configuration.

  3. 03/ 03

    Diagnosis and implementation

    Based on the data collected, I prepare a diagnosis of crawl budget allocation and a change plan, and after implementation I verify that the settings are consistent with the agreed plan.

Defining the scope of crawl budget optimisation

Defining the scope of crawl budget optimisation consists of collecting key information about the site and specifying precisely which areas are to be covered by the analysis and the actions. At the kick-off stage I collect data on the type of site, its structure, its environments and what the process for implementing and publishing changes looks like. In parallel we agree how to understand indexing priorities in the context of your website and which sections are critical to keep in the index. As a result, the subsequent steps (diagnosis, change plan and implementation) are grounded in the real technical and organisational context.

As part of the scope we clarify which resources the work covers: domains, subdomains, language versions and the page types to be analysed in terms of crawling and indexing. We also determine which URL types are to be treated as priorities and which will potentially be restricted or excluded in the subsequent stages. This is also the stage at which it is decided whether we analyse server logs — if they are available, the diagnosis of crawl budget allocation is more accurate. If there are no logs, or they cannot be shared, the further work is based on indexing data, sitemaps and crawl observations.

Access and materials for crawl budget analysis

Access and materials for crawl budget analysis are the set of data without which it is not possible to reliably assess where the bot wastes time and how to direct the crawl to the right URLs. The work requires, above all, server logs (if available), sitemaps and a list of the key sections of the site that are to be the indexing priority. Information about the rules for generating URLs is important too, because these are what often determine the creation of technical variants (e.g. parameters). In addition, access to the robots configuration and the ability to verify the implementation are necessary, in order to confirm that the settings are consistent with the change plan.

The range of input materials directly affects the precision of the diagnosis and the way the analysis is carried out.

When logs are missing, it is harder to confirm unambiguously the real distribution of the crawl at the level of URL types, and the recommendations are then based on indexing data and URL tests. That is why at the beginning we establish what data can be obtained and in what form it can be shared for analysis.

  • Server logs (if available).
  • Sitemaps.
  • List of the key sections of the site and the indexing priorities.
  • Information about the rules for generating URLs.
  • Access to the robots configuration and the ability to verify the implementation.

Server log analysis and crawl budget efficiency

Server log analysis makes it possible to assess most accurately how the crawl budget is really distributed at the level of URL types and server responses. In practice I check in the logs the frequency of bot visits, the distribution of crawling across categories of URLs and which HTTP statuses the visited URLs return most often. I also verify redirect loops and situations in which the bot lands on non-canonical URLs or on resources that should not absorb a significant part of the crawl. Such a diagnosis makes it easier to point out the places where the bot wastes time and what technically lies behind this distribution.

If server logs are not available, the assessment of crawl budget allocation does not include confirmation of actual bot visits to specific URL types. The analysis is then based on indexing data, sitemaps, crawl observations and URL tests, which makes the conclusions about how the crawl is “used up” less clear-cut.

The decision to analyse logs is therefore one of the arrangements that affects the level of detail of the diagnosis.

Collecting crawl and indexing data

Collecting crawl and indexing data consists of obtaining and organising information about the URLs that affect where the bot goes and what ultimately has a chance of being indexed. I organise the data on HTTP statuses, canonicalisation, indexing directives and which URLs are in the sitemaps. In parallel I analyse internal linking patterns, because they often determine whether the bot reaches priority pages or technical variants. The scope of this inventory grows with the size and complexity of the site, including the number of URLs, parameters, filters and language versions.

As part of the input data I create an inventory of page types and URL patterns, such as categories, products, pagination or URLs with filter parameters, together with an assessment of their role in indexing. I then compare the number and quality of URLs in the sitemaps with the URLs actually indexed, in order to identify sections with excess or insufficient indexing. I also verify whether the sitemaps contain only URLs that are canonical, indexable and have the correct HTTP status, and whether they are logically divided, e.g. by content type. Finally, I check the consistency of robots.txt, noindex, canonical and X-Robots-Tag headers with the sitemaps, and assess whether the linking generates parameterised URLs on a mass scale, creates paths that are too deep or leaves orphaned URLs.

Diagnosing crawl budget waste

Diagnosing crawl budget waste consists of gathering the irregularities identified into a coherent list and assigning them to specific causes. Based on the earlier observations, I group the problems by areas such as duplication, parameters, server errors, redirect chains, inefficient sitemaps and weak internal linking, among others. This breakdown makes it easier to assess which mechanisms “use up” the crawl on unwanted URLs and which limit the bot’s access to priority pages. The result is an organised picture of where the losses really arise and which URL types generate them most.

The outcome of the diagnosis is a condensed list of problems affecting the crawl, indicating the specific classes of URLs and server behaviours that absorb the bot’s resources. The list may include, among other things, non-canonical URLs, parameterised variants, soft 404s, 3xx/4xx/5xx responses, empty pages, inconsistent canonicals, and pagination or filter loops. On this basis I prepare the starting point for the “index vs. exclude” decision, i.e. determining which URL types are to remain indexable as valuable and which should be restricted in the crawl or excluded. Establishing this policy is crucial, because it determines the further recommendations and the way they are implemented.

  • Grouping problems by cause (e.g. duplication, parameters, server errors, redirects, sitemaps, linking).
  • List of the places where the bot wastes time (e.g. non-canonical URLs, soft 404s, 3xx/4xx/5xx, filter/pagination loops).
  • Decision support on which URL types are to be indexed and which restricted or excluded.

Implementation plan and prioritisation of changes

The implementation plan and prioritisation of changes means arranging the recommendations in order of execution, together with the reasoning for “what to change, where and why”. I create the plan on the basis of the agreed policy for URL types (what is to be indexed and what restricted in the crawl), so that the technical and configuration actions are consistent with one another. For each change I also describe the risks and dependencies, so as not to cut off pages needed for indexing and not to introduce contradictory signals. I match the order of work to how the release process works on your site and to the technical dependencies.

The result of the planning is an implementation backlog, i.e. a list of tasks ready to be handed over for implementation and verification. Each task includes a priority, a description of the implementation, an acceptance criterion and a way of checking the effect, e.g. which URLs are to stop being crawled and which are to become canonical. At this stage I take into account that some changes may require development work or modifications in the CMS, and that the pace of delivery will depend on the availability of resources and the release cycle. The scope of the backlog may also grow with the size and complexity of the site, e.g. with a large number of URLs, parameters, filters or language versions.

  • Change plan: what to change, where, why, with what risk and in what order.
  • Implementation backlog: priority, description of the implementation, acceptance criterion and method of verification.
  • Taking into account the technical dependencies, the implementation process and the availability of the team making the changes.
Testimonials
What clients and the industry say

Feedback from clients and industry people I have worked with on SEO projects.

Damian Salkowski

I had the chance to work with Kuba at Kulturalnie o SEO, an event I organise. Kuba did a great job as a speaker and received high marks from the audience. He showed professionalism and broad knowledge. In other projects at Vestigio, Kuba shows enormous commitment, a willingness to explore and implement new ideas, and excellent organisation of his work.

Damian SalkowskiCEO of SENUTO
01 / 08
Jakub Dzikowski
Jakub DzikowskiSEO freelancer & consultant

You talk to me, not a salesperson

Tell me what you want to achieve. I’ll reply personally

I work as a freelancer: the same person reads your message, prepares the quote and then runs the project. No sales team in between.