Glossary · Technical SEO

What is crawl budget?

Crawl budget is the amount of crawling a search engine is willing and able to do for a website. Google's guidance emphasizes large or rapidly changing sites, rather than ordinary small business websites.

Updated

What is crawl budget?

Crawl budget can matter when a site creates very large numbers of URLs or changes important content frequently. Most small service websites should first fix discoverability and indexing errors. Blaming a crawl budget limit can distract from a noindex rule or an unavailable service page that needs a direct repair.

See the relationships

Crawl capacity and demand meet at the host

Actual crawling is shaped by what a crawler can fetch and what it wants to fetch.

Crawl budget: related considerationsConceptual connections between Crawl budget and Host capacity, Crawl demand, URL inventory, Observed requests. Connections group considerations; they do not represent measured effects or mandatory sequence. Explanations follow below.Crawl budgetHost capacityCrawl demandURL inventoryObserved requests
  • Host capacity

    Availability and serving performance constrain acceptable fetching.

  • Crawl demand

    Known URL usefulness and change patterns influence revisiting.

  • URL inventory

    Generated variants and necessary resources affect what is available to fetch.

  • Observed requests

    Logs and webmaster reports describe fetching under their own coverage rules.

Indexing decisions follow fetching and are not the same measurement.Conceptual illustration informed by Crawl Budget Management.

A service-business example

Illustrative example, not client data. An HVAC website’s filter widget creates many combinations of service, suburb, and sort order. Most versions repeat the same cards. The useful repair is to review those URL patterns and decide which are needed for customers, rather than purchasing a crawler service that promises to spend Google’s budget differently.

When is crawl budget a useful diagnosis?

Crawl budget becomes a useful diagnosis when a site’s scale, change rate, or observed serving problems make crawling capacity relevant. Google’s crawl-budget guidance primarily addresses large or frequently updated sites. Its size examples are rough guidance rather than fixed eligibility thresholds.

Primary evidence: crawl-budget guidance. Accessed October 8, 2026.

  • Many service websites have more direct problems.

    A missing navigation link, an accidental noindex header, or a broken service route can explain why an important page is absent. These causes deserve inspection before the business concludes that Google assigned too little crawling attention to the entire host.

  • Start with a concrete question.

    Is a newly published explanation unknown to Google? Is an already known page waiting to be fetched? Did a request fail? Was a fetched page excluded during indexing? These are distinct stages. An issue in the last stage is not automatically evidence of inadequate crawl capacity.

  • Scale also needs a useful definition.

    Count meaningful public content separately from parameter combinations and utility routes. A contractor site with a modest service inventory can expose a much larger apparent inventory through filters. That situation concerns uncontrolled address generation, not necessarily a shortage affecting the useful service pages.

  • A defensible budget investigation links the observed delay or access problem to capacity, demand, or inventory evidence.

    Without that connection, the term can become a vague explanation for anything missing from search. Keep the original symptom visible throughout the investigation so the proposed work addresses the actual business problem.

How do crawl capacity and crawl demand differ?

Crawl capacity concerns what Google’s infrastructure can fetch without overwhelming the host. Crawl demand concerns what the relevant crawler wants to revisit or discover. Google describes crawl budget through both elements. A healthy server can have spare capacity while Google has little current demand for additional requests.

How do crawl capacity and crawl demand differ?
Point to considerExplanation and application
Capacity relates to connections and response behavior.If the host becomes slow or returns server errors, Google can reduce crawling. Stable responses make more work technically feasible where demand exists. This is a serving consideration, not a certificate that the content is useful or that every known address deserves processing.
Demand reflects factors such as the known inventory, changes, and the information’s value for the relevant product.A site move can require reprocessing at new addresses. Unnecessary duplicate paths can expand what the crawler perceives as available. Different Google products may have different reasons to request the same host.
The distinction changes the implementation decision.More hosting resources may help when a demonstrated server-capacity issue prevents fetching. They do not automatically create demand for nearly identical pages. Conversely, cleaning up duplicate inventory can make the architecture clearer without repairing a host that repeatedly drops connections.
Document the hypothesis before changing settings.A capacity hypothesis should identify serving failures or a documented hostload observation. An inventory hypothesis should identify concrete patterns exposing unnecessary paths. A demand question requires careful interpretation and cannot be settled by buying a faster server or submitting the same unchanged page repeatedly.

What does host scope mean for the investigation?

Google’s crawl-budget guidance treats a site as a unique hostname for this purpose. An application’s publishing inventory, a business’s brand, and the crawler’s host boundary are therefore different concepts. Check the actual hosts involved before combining observations or assuming that changes on one host explain requests on another.

  • A public website may use a separate booking host, image CDN, or resource subdomain.

    Those resources can affect the customer experience and rendering. Their request histories may not appear in the same report. Record the dependency relationship without collapsing every host into one indistinguishable crawl total.

  • Property scope in Search Console adds another boundary.

    A domain property and a narrower URL-prefix property can expose different views. Google’s Crawl Stats documentation, accessed October 8, 2026, explains the scope of requested URLs and resources. Read that scope alongside the displayed totals.

  • Logs also need host labels.

    An export from the origin might omit cached edge responses or include traffic for several hostnames. A report about one service host should not silently include unrelated requests to staging or another application. Filter according to the question being tested and retain the original boundaries.

  • This matters during migrations.

    Requests to old and new hosts can coexist while addresses are being reprocessed. A change in one host’s request count is not a complete migration result. Inspect old-to-new routes, final destinations, and important pages rather than treating the total as a measure of completed transfer.

How should the intended URL inventory be defined?

The intended inventory is the set of public destinations the business wants customers and search engines to use. Define it by purpose before comparing it with discovered addresses. Separate substantive service pages, supporting advice, collection pages, duplicate delivery variants, and utility routes that serve another operational task.

  • Use the publishing database as one starting source.

    Add routes created outside it, such as filters, search pages, and application endpoints. A content inventory alone can miss dynamically generated combinations. A homepage-started crawl alone can miss unlinked pages. Combining sources reveals the patterns that a single inventory method cannot expose.

  • For each pattern, record what changes in the content.

    A sort parameter may rearrange the same cards without adding information. A service filter may produce a genuinely useful collection or only another empty result. The decision should follow that difference instead of assuming that every parameter address is unnecessary.

  • Review orphan pages separately.

    An important guide absent from navigation is a discovery gap even if the site’s total inventory is controlled. Making the existing explanation reachable can be more valuable than narrowing unrelated filter combinations. Inventory control and discovery repair often belong to the same review but require different edits.

How can parameter combinations consume unnecessary work?

Parameter combinations can create many addresses without creating equally many useful explanations. A crawler discovers the addresses through links or other sources and may revisit them as part of its known inventory. The problem is the mismatch between distinct addresses and distinct user value, not the mere presence of a query string.

  • Consider a service directory with sorting and filtering.

    If each selection links to every possible next combination, the interface can expose a large network of repeated results. Some combinations may be useful for browsing but unnecessary as standalone search destinations. Others may be empty or differ only in card order.

  • Inspect representative combinations manually.

    Compare the main content, the intended audience, and whether a person could reasonably arrive on the page as an independent answer. Determine whether the feature needs crawlable public routes for all combinations or whether useful component pages can provide the relevant discovery paths more clearly.

  • Canonical annotations can identify equivalent pages while preserving access.

    They do not prevent fetching by themselves, because the crawler needs to read the relationship. Google’s budget guidance discusses consolidation and appropriate crawl restrictions as different inventory tools. Choose between them according to the page’s function and desired processing behavior.

  • Avoid broad exclusions that catch genuine service explanations.

    A parameter used for sorting may also appear in a route with distinct content. Test the application’s actual behavior before publishing a rule. The repair should reduce unnecessary address generation without breaking the browsing paths customers rely on.

How should robots.txt be used for inventory control?

Robots.txt can limit crawling of paths the business does not want processed through ordinary automatic crawling. It should express a stable access intention rather than serve as a temporary attempt to move attention between pages. Google’s budget guidance explicitly warns that blocking some routes does not automatically reallocate their crawling to other routes.

  • Review the exact path pattern and host before applying a rule.

    Google’s robots specification, accessed October 8, 2026, describes group selection and matching. A rule intended for repeated filter routes can unintentionally affect a useful page if both share the same path structure.

  • Do not use a crawl restriction to express every indexing decision.

    A blocked resource can remain known to Google without its content being fetched. If the business needs a supported page-level exclusion to be processed, access to that instruction matters. The desired outcome determines which control should be applied.

  • The robots.txt definition explains this access boundary.

    Keep private information behind authentication regardless of the published crawl preference. A crawler instruction is not a confidentiality control, and a large crawl report is not a reason to expose customer data or relax unrelated security requirements.

  • Test both intended exclusions and adjacent allowed pages after a change.

    Include required scripts and stylesheets when the rule can affect resources. A smaller requested inventory is not a useful outcome if the remaining service pages render without their essential content because a shared dependency was accidentally blocked.

Why is noindex different from a crawl-saving control?

Noindex tells a supporting search engine not to retain the page in its index after reading the instruction. Reading requires a request. Google’s crawl-budget guidance therefore distinguishes noindex from measures aimed specifically at avoiding unwanted fetching. The choice should follow the intended indexing and access behavior rather than a single cleanup slogan.

Why is noindex different from a crawl-saving control?
Point to considerExplanation and application
A public confirmation page may need to remain available while staying out of search.Noindex can fit that purpose. The fact that it still requires processing does not make the page’s exclusion wrong. Operational utility and crawling efficiency are related considerations, but the latter should not erase the former’s legitimate requirements.
An already indexed unwanted route creates another question.Preventing future access can stop Google from reading a new exclusion. Review the current processing state before changing controls. A plan that removes the only readable directive can leave the business with less useful evidence about whether the intended search exclusion took effect.
Read noindex for its supported implementation locations.Check both HTML and HTTP headers. A publishing setting can affect a whole template, and an inherited staging header can exclude service pages the business wanted indexed. A request budget discussion should not distract from that immediate configuration error.
Record why a control was selected.If the purpose is search exclusion for a public utility, state that. If the purpose is stable crawl restriction for unnecessary generated paths, state that separately. This makes later changes reviewable and prevents a migration from copying one policy into a context where it no longer fits.

What should happen to removed pages and broken destinations?

Removed pages need responses that accurately represent their current state. A relevant permanent move differs from a resource that no longer exists without a replacement. Google’s budget guidance recommends appropriate missing-resource responses for permanently removed content and warns that soft-error pages can continue consuming unnecessary crawling work.

  • A redesign should begin with an old-to-new mapping for genuine moves.

    Update current internal links to the final destinations. Retain useful old routes where external links or bookmarks still need an equivalent destination. Sending every retired address to the homepage does not establish that the old content’s task has a relevant replacement.

  • An obsolete promotion with no continuing equivalent may need a genuine missing-resource response.

    A deleted service record that should still exist may need restoration. These decisions require business context. A developer should not choose them merely to eliminate an error count without checking whether the service is still offered.

  • A soft 404 occurs when content behaves like an error despite a successful response.

    The application may return its normal shell for any unknown route. Inspect a working service address and a deliberately missing one together. If both receive the same generic success shell, review route handling.

How do server response problems affect capacity?

Server response problems can make crawling less feasible by increasing the work required per request or causing failures. Google’s budget guidance discusses response stability, latency, server errors, and rate-limiting signals. Investigate these observations at the delivery layer before assuming the solution is new content or a different sitemap submission.

  • Check whether failures are concentrated during particular operations.

    A publishing job, backup, expensive filter request, or external data dependency may coincide with slow responses. The timing connection is a hypothesis until verified. Compare request evidence and application diagnostics rather than treating a nearby deployment timestamp as proof of causation.

  • Distinguish the origin from the edge.

    An edge cache may serve a page quickly while origin requests remain expensive. A security service may reject requests before the origin sees them. Both situations can affect interpretation of server logs. The responsible layer determines which team and configuration need attention.

  • Time to First Byte helps describe part of navigation timing, but it is not a direct measure of content usefulness.

    Inspect the timing breakdown and response behavior. A quick empty shell can still leave a customer and renderer waiting for the actual explanation through later data requests.

  • More server resources are appropriate only when evidence supports a capacity problem and the business case justifies the change.

    A repeated application error may need a code repair instead. An unnecessary generated inventory may need route controls. Keep the implementation proportional to the observed bottleneck rather than purchasing capacity to mask an avoidable failure.

What role can HTTP caching play?

HTTP caching can let a crawler reuse unchanged information instead of transferring the same document repeatedly. Google’s crawler infrastructure overview describes support for validators such as ETag and Last-Modified. Correct implementation depends on whether the underlying content actually changed. Test unchanged and changed representations as separate response cases.

Primary evidence: crawler infrastructure overview. Accessed October 8, 2026.

  • An entity tag identifies a particular representation for conditional requests.

    If the representation is unchanged, the server can respond accordingly. A last-modified value describes a relevant update time. These mechanisms should reflect real content state, not a freshly generated timestamp on every request or build.

  • Check how the application creates validators.

    A publishing change should alter the representation where needed. An unchanged request should not appear newly updated solely because a runtime timestamp was inserted into the document. A cache that never recognizes real updates can also serve stale service information after an editor changes an important detail.

  • Test normal and conditional responses with the hosting team.

    Confirm that the unchanged case behaves appropriately and the changed case delivers the current representation. A nominal cache setting is not evidence that the public response follows it. CDNs, application middleware, and content APIs can each contribute their own behavior.

  • Do not use caching to freeze information customers need current.

    Service availability, contact details, and substantive instructions require a suitable update path. The efficiency benefit should accompany accurate delivery. A smaller transfer is a poor outcome if the business can no longer publish a corrected explanation reliably.

How should crawling metrics be read?

Crawling metrics describe requests, responses, and covered resources. They do not directly describe indexed pages, qualified calls, or useful content. Read totals alongside their scope and category definitions. A change in requests can come from redirects, repeated resources, or a migration rather than increased attention to distinct service explanations.

  • Use log-file analysis for verified observations from available request records.

    Establish retention, sampling, timezone, and CDN coverage. Filter by resource pattern and response. Verify claimed crawler identity where needed. An apparent Googlebot pattern based only on user-agent names can include unrelated programs imitating the header.

  • Crawl Stats provides Google’s own history, with different coverage limits.

    Example URLs are samples rather than a complete inventory. The report also counts certain failed or considered requests under its documented rules. Do not require its totals to match an incomplete origin export before accepting that either has useful evidence.

  • Compare periods that cover the relevant change.

    A migration can temporarily increase reprocessing. A quiet content period can produce fewer requests without a serving problem. The metric needs the publishing and routing context. Retain the dates, host boundaries, and URL patterns with the explanation rather than reporting a percentage without a meaningful baseline.

  • Assess the original symptom separately.

    Did the intended service page become discoverable? Did its response become reliable? Did Google’s later processing reflect the corrected page? Those observations are more useful than a claim that the business increased its hidden budget, especially when no direct allocation measurement was available.

What diagnostic sequence avoids premature budget claims?

A useful diagnostic sequence moves from the affected page to the broader host only when the evidence warrants it. Begin with the intended public URL and its business purpose. Verify the response, discovery routes, and indexing instructions before attributing absence to Google’s overall capacity or demand.

  • Check the Page indexing report and inspect important examples individually.

    Separate discovered but unfetched pages from fetched pages excluded for another reason. An unexpected canonical relationship requires a content and signal comparison. A noindex exclusion requires an instruction review. Neither is repaired by discussing crawl capacity alone.

  • Next examine shared patterns.

    If many important routes are unreachable, review host availability and common configuration. If generated variants dominate the known inventory, assess their necessity. If the evidence concerns only an unlinked guide, repair its introduction rather than applying a site-wide crawl restriction to unrelated paths.

Continue the public-page review

Our website SEO checker provides an initial public-page review. It does not measure Google’s hidden crawl allocation or provide server logs. Use the evidence methods above to establish the cause, then connect the repair with the wider priorities of SEO services.

Questions about Crawl budget

Do most small websites need crawl-budget optimization?

Usually not. Google directs this guidance toward large, frequently changing sites; small sites should first investigate discoverability, access and indexing conditions.

Google Search documentation ↗
How do crawl capacity and demand differ?

Capacity concerns how much crawling a host can handle. Demand concerns which resources the crawler wants to revisit; both affect actual fetching.

Google Search documentation ↗
Does noindex save crawl requests?

Noindex must be read to take effect, which requires a request. It is an indexing control rather than a way to stop crawling that URL.

Google Search documentation ↗
What do server errors and slow responses change?

Serving problems can reduce Google's crawl capacity. Investigate response behavior and host health rather than assuming every unfetched URL needs a different crawl-budget setting.

Google Search documentation ↗

Continue learning

Try a relevant tool

  • Website SEO checker →

    Check individual URL access and directives before attributing a larger fetching problem to crawl budget.

Sources

Crawl Budget Management | Google Crawling Infrastructure  |  Crawling infrastructure  |  Google for Developers ↗Accessed October 8, 2026Crawl Stats report - Search Console Help ↗Accessed October 8, 2026How Google Interprets the robots.txt Specification | Google Crawling Infrastructure  |  Crawling infrastructure  |  Google for Developers ↗Accessed October 8, 2026Google Crawler (User Agent) Overview | Google Crawling Infrastructure  |  Crawling infrastructure  |  Google for Developers ↗Accessed October 8, 2026

Published . Definitions and examples link to their supporting sources. Our SEO methodology →

SEO · Content · Local · Web Design

Connect the website work to your business.

We assess the pages, search demand, and customer actions that matter to your business, then explain where to focus the work.