Glossary · Technical SEO

What is a web crawler?

A web crawler is software that requests web resources to discover or update information. Different crawlers have different purposes, access rules, and ways of identifying themselves.

Updated

Why Web crawler matters for a service business

A web crawler discovers information by requesting URLs. For a service website, links and sitemaps provide routes to the pages buyers need. A crawler can fetch public content for search indexing, monitoring, or another purpose, so understand the program’s role before deciding whether to allow or restrict it.

Orphan pages can be absent from a configured internal-link crawl while remaining known through other discovery sources.

How does a web crawler move through a website?

A web crawler requests known addresses, reads responses, and discovers additional addresses through the information it receives. Its behavior depends on its purpose and implementation. A search crawler needs enough access to understand public content, while a site-audit crawler may only inspect the pages reachable within its configured starting points.

  • The crawler does not begin with a complete copy of your website’s publishing database.

    It needs discovery routes. Those routes may include previously known addresses, ordinary links, and a submitted sitemap. Publishing a record in a content system creates a page, but it does not automatically create a useful route to that page.

  • Consider a heating company’s repair guide.

    The publishing system knows the guide exists. A customer browsing the service section may not know it exists because no page introduces it. A crawler starting at the homepage faces the same gap. This is an architecture issue even before anyone discusses search rankings.

  • An audit should therefore distinguish the complete publishing inventory from the linked inventory.

    The first says which pages exist according to the business. The second says which pages the chosen crawler discovered through the routes it followed. The difference deserves review, not an automatic instruction to put every missing item into the main menu.

What happens between discovery and indexing?

Discovery, fetching, rendering, and indexing describe different events. A crawler can know an address without successfully requesting it. A fetched page can still be unsuitable for indexing. Google describes these distinctions in its explanation of how Search works. The evidence needed depends on the stage under investigation.

Primary evidence: explanation of how Search works. Accessed October 8, 2026.

  • A discovery record establishes that an address became known to the system.

    It does not establish that the server delivered the page. The fetch may later fail because of a network problem, an access rule, or an unavailable host. Check the response evidence before diagnosing a content issue at this stage.

  • A successful fetch establishes that some response arrived.

    The response might contain the complete explanation, a login screen, a generic shell, or an error message. A successful status alone cannot tell you which of these the crawler received. Inspect the returned document and its essential dependencies.

  • Rendering can create the content that an initial response did not contain.

    Google’s processing supports JavaScript rendering, but that does not make every implementation equally reliable. A required data request can fail. A blocked script can prevent important text from appearing. Another crawler may not execute the same scripts at all.

  • Indexing is a later selection and processing decision.

    Google may select another representative for equivalent content or exclude a page for another reason. A log entry showing a successful visit cannot settle that decision. Investigate the appropriate platform report rather than describing every successful request as an indexed service page.

  • The Page indexing report helps examine Google’s processing outcomes.

    Keep its categories separate from observations made by a local audit crawler. They can support the same investigation, but they describe different systems with different evidence and reporting boundaries.

  • A response comparison can begin with a full public GET saved as evidence.

    Browser inspection then adds the rendered state. The two observations reveal whether the explanation arrived in the original document or only after later execution. They should not be merged into a single claim that every crawler sees the same output.

curl --silent --show-error --location \
  --dump-header crawl-response-headers.txt \
  --output crawl-response.html \
  'https://example.com/service/'

This is illustrative command syntax using a reserved example address. Replace the destination with the public resource under review. The saved HTML does not execute scripts. If the browser shows additional service text, inspect the network request responsible for that addition and verify its public response separately.

A difference is a diagnostic lead, not proof of invisibility in search. A script-rendering crawler may process the later text, while a response-only audit may not. Check the relevant provider’s rendering evidence before deciding that the business needs an implementation change.

Follow the process

Fetching and indexing are different stages

Discovery supplies candidate addresses and requests reveal their responses. Processing, further link discovery and index consideration are related stages with different evidence.

  1. Discovery

    Known addresses, links and sitemaps provide candidate URLs.

  2. Request

    The crawler receives the server's response and applicable content.

  3. Processing

    The search system interprets the delivered information and renders where appropriate.

  4. Further discovery

    New crawlable links supply additional candidate URLs.

  5. Index consideration

    Successful crawling alone does not decide inclusion or search appearance.

A configured audit crawl and a search engine's indexing record are not interchangeable inventories.Conceptual illustration informed by In-Depth Guide to How Google Search Works.

Are all automated requests the same kind of crawler?

Automated requests come from programs with different jobs, permissions, and triggers. Identifying the program’s role helps explain why it visited a page and which controls it respects. Google’s crawler and fetcher overview separates common crawlers, special-case crawlers, and user-triggered fetchers.

Primary evidence: crawler and fetcher overview. Accessed October 8, 2026.

Are all automated requests the same kind of crawler?
Point to considerExplanation and application
A common search crawler follows discovery and refresh behavior for its product.A user-triggered fetch can happen because somebody runs a verification or inspection tool. An advertising-related request can come from another Google product. Treating all of these as ordinary Google Search visits can produce misleading crawl reports.
The user-agent header provides an identity claim.It is useful for preliminary filtering, but it is not proof of ownership. Any client can send a familiar name. If a security decision depends on identity, verify the source through the relevant provider’s documented method rather than trusting the header by itself.
A local SEO audit application has another role.It may request pages with its own user agent and follow its own limits. An exclusion in its results may reflect configuration, authentication, a crawl-depth limit, or a response difference. The omission does not establish that Google also failed to visit the page.
Ask which question the observed program can answer.An audit crawler can reveal broken routes and exposed markup. A verified search crawler request can establish a fetch observation. Neither observation independently proves demand for the service or the quality of a resulting lead. Those questions require separate business and search-performance evidence.

Reliable discovery links use ordinary anchor elements with usable destination addresses. Google documents this in its crawlable-link guidance. A visual card or clickable-looking label is not enough when its destination exists only inside an event handler or an application-specific attribute.

Primary evidence: crawlable-link guidance. Accessed October 8, 2026.

  • Inspect the actual anchor markup.

    The destination should appear in the href attribute and resolve to a public address. A script can enhance the interaction without removing that destination. This lets ordinary navigation retain a clear route even when an enhancement does not execute or a particular crawler does not support it.

  • The relationship between pages also matters.

    A repair explanation linked from a relevant service section has a sensible role for customers. A long footer containing every route may provide technical links without helping anyone decide where to go. Discovery should follow the site’s informational structure rather than an arbitrary list of addresses.

  • Review internal linking when a useful page has no relevant introduction.

    Choose the source page according to the buyer’s task. A water-heater maintenance guide may belong near the water-heater service explanation. It does not need to appear under unrelated drainage advice simply because that page already receives visits.

  • Broken links interrupt discovery and customer navigation together.

    Request the destination instead of judging the markup alone. A well-formed href pointing to a missing page still creates a bad route. If the content moved, update the current link and assess whether the old address needs a relevant redirect.

What can a sitemap contribute?

A sitemap can expose intended public addresses that a crawler might otherwise discover slowly or miss through existing links. It supplements navigation rather than replacing it. Google’s sitemap overview describes this discovery role and its limits. Normal navigation still provides the customer’s browsing context.

Primary evidence: sitemap overview. Accessed October 8, 2026.

  • The list should reflect preferred destinations.

    A service route that redirects elsewhere is usually a poor inventory entry when the final destination is known. An excluded confirmation page has a different purpose from a public repair guide. Review these categories before exporting every record that happens to exist in the publishing database.

  • Submitting the file also creates another response to verify.

    The file must be reachable and contain the intended URLs. A submission receipt does not mean every listed page was fetched or selected for search. Follow up on important pages individually when their status remains uncertain.

  • The XML sitemap is particularly useful alongside an independent page inventory.

    Compare it with a navigation crawl. A route appearing only in the sitemap may be deliberately outside browsing, or it may be an important explanation that an editor forgot to link. The difference needs a purpose-based decision.

  • Avoid artificial change metadata.

    If a page’s substantive content did not change, a fresh build alone does not make the information newly updated. Keep publishing signals connected to real content changes. This makes the file more defensible as a description of the site rather than a device for requesting attention without new information.

How do access controls affect crawling?

Access controls determine whether a client receives the requested resource. Robots instructions communicate crawl preferences to cooperating programs. Authentication and security rules enforce access in different ways. Before changing either, identify the resource, the requesting program, and the purpose of the restriction.

  • A robots.txt rule can prevent a cooperating crawler from requesting a public path.

    It does not secure private customer records. If a document should require permission, use actual access control. Hiding an address from a crawl report is not a substitute for protecting the underlying information.

  • Crawler access can also be interrupted by a firewall challenge.

    A customer may pass the challenge through a normal browser interaction, while an automated request receives the challenge page. The underlying service page can therefore appear healthy to the business owner even though the crawler never receives its content.

  • Check whether the restriction affects the HTML document or a dependency.

    A blocked stylesheet or script can change rendering without blocking the main request. A data endpoint can be essential when a page loads its explanation after navigation. Inspect the dependency chain instead of stopping after the first successful document response.

  • Avoid broad exceptions based only on a familiar bot name.

    Verification should precede privileged treatment. Google’s request-verification instructions, accessed October 8, 2026, provide documented checks for its own crawlers and fetchers. Other providers require their own current evidence.

How should a business inspect a suspicious request?

A suspicious request should be investigated through its recorded address, source, response, and timing. Start with the exact request that prompted concern. A pattern such as repeated missing paths may have a different explanation from an isolated request to a current service page.

  • Retain the original log observation before applying a filter.

    Note the timezone and delivery layer. An edge service can answer requests without contacting the origin. Origin records alone may therefore miss successful responses. A conclusion about absent crawling needs a clear statement of which records were actually available.

  • Read the user-agent claim, then verify identity when it matters.

    Do not infer that a request belongs to a provider because its hostname merely contains the provider’s name as text. Follow the provider’s complete verification procedure, including any reverse and forward lookup requirements or published address ranges.

  • Next examine the delivered response.

    A recorded success code may accompany a challenge or empty application shell. Compare the response body with the expected service information. If the request redirected, follow the final address and inspect it as well. The initial response can be correct while the destination is broken.

  • Use log-file analysis for repeatable patterns across a defined period.

    Record what the evidence supports: verified requests to specified paths, particular statuses, and observed failures. Avoid extending those observations into claims about all crawlers, every customer, or subsequent indexing decisions.

An illustrative diagnosis: a guide behind a search widget

Another illustrative diagnosis: the audit and browser disagree

How should crawl findings be prioritized?

Prioritize crawling problems according to the customer task and the evidence of failure. An inaccessible primary service explanation deserves attention before an obsolete campaign variant. A report’s largest category is not automatically its most important issue, particularly when many entries describe intentional utility routes.

  • Create an intended public inventory first.

    Identify the service pages, supporting advice, and contact destinations the business expects customers to use. Record which pages should remain outside search or ordinary browsing. This gives a crawl finding a meaningful baseline instead of assuming every known address deserves equal exposure.

  • Group failures by responsible cause.

    Several routes may fail because one template points to an old host. Another group may be blocked by a deployment rule. Repairing the shared cause can be more useful than treating each address as an unrelated editorial problem. Keep representative examples to verify the change.

  • Check the final user experience as well.

    A crawler can reach a thin page that does not explain the offered work. A newly repaired route can point customers to an unrelated destination. Accessibility is necessary for discovery, but it does not establish that the page deserves a place in the customer’s decision process.

  • For larger or rapidly changing sites, investigate crawl budget only when scale and observed patterns make that question relevant.

    A small contractor site with a missing navigation link usually needs a direct architecture repair before an advanced allocation discussion. Choose the simplest explanation supported by the evidence.

What belongs in a crawl change record?

A crawl change record should explain the affected resource, the observed failure, the repair, and the remaining uncertainty. It lets an editor, developer, and business owner assess the same issue without confusing a technical response test with a later search-processing outcome.

  • Record the original route and the tested destination.

    If a redirect changes, retain the old-to-new mapping. If an access rule changes, record its intended scope. If navigation changes, identify the source page and the new link. These details make the work reproducible during a later redesign or security review.

  • Include the testing environment.

    A logged-in browser, an origin request, and an edge request may receive different responses. Note the relevant identity checks and whether rendering was performed. These boundaries explain why a test passed and prevent a future reader from assuming it covered conditions that were never examined.

  • State the immediate result precisely.

    The tested crawler reached the page, the dependency became accessible, or the redirect reached its intended destination. State the later question separately. Google’s selected representative, indexing state, and search performance may need additional evidence after processing. A clear handoff protects the business from treating an incomplete diagnosis as a completed outcome.

How does a crawl audit avoid configuration blind spots?

A crawl audit only covers the routes and behaviors its configuration allows. Review the starting addresses, excluded patterns, authentication settings, and rendering mode before interpreting its missing pages. A deliberate limit can explain an omission without establishing a problem in the public site.

  • Compare a narrow service-section crawl with the wider publishing inventory.

    If the crawler begins inside one branch, it may never follow links that exist only in another branch. Expand the scope when that comparison is the question being investigated. Do not describe a restricted sample as a complete site inventory.

  • Check destination hosts as well.

    A booking route can leave the main website for an external scheduling service. The audit may stop at that boundary. The customer journey still needs a working link and useful destination, even though the external service falls outside the original crawl’s ownership and configuration.

  • Retain those exclusions in the audit record.

    This allows a later reviewer to reproduce the observation and identify which routes require another method. An explicit boundary is more useful than a large address count with no explanation of what the crawler omitted.

Continue the public-page review

Use the website SEO checker for a preliminary public-page review. Follow the manual discovery, response, and identity procedures above for evidence outside its scope. Our technical SEO service connects those findings to implementation priorities and the pages customers need.

Questions about Web crawler

Does crawling guarantee indexing?

No. Fetching a URL is one stage; indexing and serving decisions follow separate processing.

In-Depth Guide to How Google Search Works ↗
How does Google discover URLs?

Google can discover new pages through links from known pages and through sitemaps. A page must then be fetched and processed; discovery alone does not guarantee indexing.

Google Search crawling and indexing stages ↗
Are all bots Googlebot?

No. Search engines, auditing tools and other programs have different purposes; a claimed user-agent can also be spoofed.

Google Crawler (User Agent) Overview ↗
How do I verify a Google crawler request?

Use Google's documented IP or DNS verification methods. A user-agent string alone is insufficient proof.

Google: verifying crawler requests ↗

Continue learning

Try a relevant tool

Sources

In-Depth Guide to How Google Search Works | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026Google Crawler (User Agent) Overview | Google Crawling Infrastructure  |  Crawling infrastructure  |  Google for Developers ↗Accessed October 8, 2026SEO Link Best Practices for Google | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026What Is a Sitemap | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026Verify Requests from Google Crawlers and Fetchers | Google Crawling Infrastructure  |  Crawling infrastructure  |  Google for Developers ↗Accessed October 8, 2026IndexNow protocol documentation ↗Accessed October 8, 2026

Published . Definitions and examples link to their supporting sources. Our SEO methodology →

SEO · Content · Local · Web Design

Connect the website work to your business.

We assess the pages, search demand, and customer actions that matter to your business, then explain where to focus the work.