Glossary · Technical SEO

What is robots.txt?

Robots.txt is a file that tells supported crawlers which paths they may request. It controls crawling and is not a reliable way to keep a URL out of search results.

Updated

Why Robots.txt matters for a service business

Robots.txt controls which resources cooperating crawlers may request. A service site can use it to manage unnecessary crawl paths, but should keep essential page resources accessible. A business that mistakes a disallow rule for an indexing block can leave a URL in search while preventing Google from reading the instruction needed to exclude it.

An X-Robots-Tag instruction must be fetched to be read; a disallowed request can prevent Google from seeing the header intended to exclude the resource.

What does robots.txt control?

Robots.txt tells cooperating automatic crawlers which paths they may request on its covered host. It does not directly remove every known address from search, and it does not secure private resources. Google’s robots.txt introduction separates crawl management from indexing exclusion and access protection.

Primary evidence: robots.txt introduction. Accessed October 8, 2026.

  • The crawler first needs the file so it can interpret the applicable rules.

    It then compares the intended request with those rules. A disallowed page can remain publicly available to ordinary visitors. This explains why a site can appear normal in the owner’s browser while Google’s automatic fetching is restricted.

  • The difference matters after a launch.

    A staging restriction copied to the public host can affect service pages without altering their titles or visible text. Rewriting those explanations would not change the access rule. The correct diagnosis starts with the exact public file and the paths matched by its relevant group.

  • A known blocked address can still have a search presence without its content being fetched.

    Google’s documentation explicitly distinguishes that case from supported page-level exclusion. If the business wants a public utility absent from indexing, choose a readable indexing instruction rather than assuming a crawl restriction has the same effect.

Where must the file be located?

The file belongs at the top-level robots.txt location for the relevant host and protocol. A similarly named file inside a content directory does not govern that directory automatically. Google’s robots specification guidance defines location, port, and hostname scope. Scope follows the public file location, not the business’s brand.

Primary evidence: robots specification guidance. Accessed October 8, 2026.

  • Request the exact location used by the public service URL.

    A file on a staging subdomain does not apply to the main site. A file on the www host does not automatically apply to the bare host. The site may redirect those hosts consistently, but that behavior needs inspection rather than assumption.

  • Protocol matters as well.

    The rule file’s scope follows its documented location. A business that changed from HTTP to HTTPS should inspect the public arrangement and redirects. Do not infer that an old host’s file is the policy Google reads for every newer route simply because the addresses look related.

  • Check nonstandard ports when they are genuinely part of the public application.

    The specification distinguishes the relevant host, protocol, and port. Most service-site audits can focus on the ordinary public destination, but the boundary is important when a testing or application endpoint has a different address arrangement.

  • Record the requested file address and response.

    A copied local configuration is not evidence of what the public host delivers. The application may generate the file dynamically, the CDN may cache it, or the deployment may still serve an older static version. Inspect the final public representation before editing the wrong source.

Choose the right approach

A blocked request can hide an indexing instruction

The crawler must fetch a resource to read its page or response-header directives. Blocking the request can therefore leave the address known while preventing Google from seeing the noindex intended to exclude it.

Robots.txt: related considerationsConceptual connections between Robots.txt and Crawl allowed, Crawl disallowed, Known URL, Correct control. Connections group considerations; they do not represent measured effects or mandatory sequence. Explanations follow below.Robots.txtCrawl allowedCrawl disallowedKnown URLCorrect control
  • Crawl allowed

    The crawler can fetch the resource and read supported page or header instructions.

  • Crawl disallowed

    The crawler cannot normally read the resource's noindex instruction.

  • Known URL

    A discovery reference can keep the address known despite blocked content.

  • Correct control

    Choose a supported indexing instruction when exclusion is the actual objective.

Choose crawl restrictions and indexing instructions according to the actual outcome needed.Conceptual illustration informed by Robots.txt Introduction and Guide.

How are user-agent groups selected?

A user-agent group associates following rules with a named crawler or a broader wildcard. Google’s crawler chooses the most specific applicable group under its documented interpretation. The order of groups does not make the last group win, and a specific group is not automatically combined with the global wildcard group.

How are user-agent groups selected?
Point to considerExplanation and application
This is a common source of surprises.A business may add a named Googlebot group to permit one resource while expecting every restriction from the wildcard group to remain inherited. Google’s specification instead treats the applicable specific group according to its own rules. Inspect the whole group rather than one added allow line.
If several specific groups apply to the same user agent, Google combines their relevant rules internally.A repeated group later in the file can therefore contribute an additional restriction. Searching only the first occurrence of a crawler name is insufficient when different systems append sections to the same generated file.
Use the documented product token rather than inventing one from a full browser-looking user-agent string.Googlebot smartphone and desktop variants share the relevant robots product token. The file is not a method for selectively targeting those variants as though they were independent named groups.
Test each crawler policy intentionally maintained by the business.A result for Googlebot does not establish how an unrelated client interprets its own matching rules or whether it cooperates with the protocol. Keep provider-specific behavior attached to the provider’s current documentation rather than generalizing one parser’s result to every automated visitor.

How do allow and disallow paths match?

Allow and disallow rules compare the rule path with the requested URL path under the documented parser behavior. Path values are case-sensitive in Google’s interpretation. A prefix can match more than a folder name, so a visually short rule may cover neighboring routes the editor never intended to restrict.

The following is illustrative configuration. It is not a proposed production file. The utility prefix needs review against real route names before use, because matching a prefix without its intended boundary can affect other resources with a similar beginning.

User-agent: *
Disallow: /utility/

That slash boundary can matter. A broader prefix such as /utility can match addresses beginning with the same text beyond the intended directory. Conversely, a rule expecting a capitalized path may miss a lowercase route. Compare representative exact paths rather than judging a rule from a natural-language summary.

  • Google supports limited wildcard behavior.

    A star can match a sequence of characters, and a dollar sign can anchor an end. These are not permission to assume that every regular-expression feature works in the file. Use the documented syntax, then test the matching cases relevant to the site.

  • Encoding can matter for non-ASCII route characters.

    Google’s specification explains comparison using percent-encoded forms and normalization of raw UTF-8 paths. When an exclusion appears inconsistent across encoded forms, inspect the actual requested address and parser behavior before replacing a useful public route or publishing a broader wildcard.

Which conflicting rule takes precedence?

Google selects the most specific matching path rule according to its documented rule-length interpretation. When matching rules conflict with equal specificity, it uses the least restrictive applicable rule. This differs from assuming that disallow always wins or that the final line in the file overrides everything above it.

  • A broad disallow can coexist with a more specific allowance.

    The allowance must actually match the requested resource. A typo or wrong path boundary can leave the broad restriction active. Read the requested URL, applicable group, and competing rules together rather than examining the intended exception in isolation.

  • The following configuration is illustrative.

    It demonstrates a permitted resource below a restricted path. It needs testing against the application’s actual resource names and query behavior before any production use.

User-agent: Googlebot
Disallow: /assets/
Allow: /assets/public-navigation.js

The example also demonstrates a design concern. Restricting a resource directory while adding isolated exceptions can be harder to maintain than allowing the public dependencies the page actually needs. When a build changes filenames, an old exception may stop matching. Prefer a policy whose scope aligns with the stable resource architecture.

Keep parser tests with the exact matching cases. Include the intended restricted path, the intended exception, and an adjacent allowed route. A correct explanation of precedence cannot compensate for testing the wrong user-agent group. The outcome depends on both group selection and path matching.

What formatting errors can change the file’s meaning?

Formatting errors can cause intended rules to be ignored or grouped differently. The file should be a supported plain-text representation with the documented encoding and line syntax. An HTML response, invalid field, or misunderstood group boundary can make the delivered policy differ from the text an editor thought they published.

  • Google’s specification describes fields, comments, and grouping.

    A comment begins after the comment marker. An empty path in a rule does not behave like a root restriction. Unsupported fields are not automatically access controls. Review the actual parser rules rather than relying on a configuration example copied from an unrelated crawler.

  • Comments are public too.

    Use them to explain a rule’s purpose without publishing credentials or private customer details. A comment does not make that information inaccessible to visitors.

  • In particular, Google does not support crawl-delay as a robots control.

    Adding that line should not be reported as an implemented Googlebot rate setting. If request load is the concern, investigate serving capacity and Google’s current methods for handling it instead of treating an ignored directive as evidence of a repair.

  • A sitemap line has its own discovery role and does not create a new rule group for the parser.

    A generated file that interleaves user-agent declarations, sitemap locations, and restrictions can therefore surprise its owner. Read the grouping behavior when a policy appears to affect more crawlers than intended.

  • Check the response’s type and contents directly.

    A frontend fallback may return the website’s HTML shell at the robots.txt address. Google can attempt to parse what arrived, but that is not a reliable substitute for the intended policy file. Repair the route so the public representation matches the configured text.

How do response errors affect robots processing?

The file’s HTTP response affects how Google proceeds with crawling. Successful delivery, client errors, server errors, and network failures have different documented handling. Google’s specification explains these branches, which is why an audit needs the response status as well as any visible text returned at the robots.txt location.

  • Google treats 4xx responses other than 429 as if no valid robots.txt file exists, so those responses impose no crawl restrictions.

    A 429 response, server errors and network failures follow different handling. Check Google’s documented failure timeline and cached-rule behavior before interpreting a temporary outage. A missing policy file and an unavailable host require different repairs.

  • For an illustrative outage investigation, keep the robots response and an ordinary page response from the same host and time.

    A 404 at robots.txt with a successful service page suggests a missing policy resource. A failed connection to both points toward a serving problem. Those observations call for different repairs even if the crawler report mentions robots access.

  • Record the incident separately from an intentional restriction.

    Restoring the host may restore delivery of an unchanged disallow rule; it does not automatically permit the service page. After recovery, inspect the delivered file and rerun the matching test for the intended URL. This checks both the serving failure and the policy that becomes readable again.

  • A redirect at the file location needs its own trace.

    The specification describes limits and behavior when fetching rules through redirects. Logical page redirects through scripts are not a replacement for the supported HTTP response arrangement. A browser eventually showing text does not establish that the crawler obtained it through the same route.

  • Avoid changing a response code solely to make the report look clean.

    If the intended file is present, deliver it correctly. If a serving failure exists, repair the responsible network, host, or application behavior. A fake success page that hides the unavailable policy can make subsequent diagnosis less clear.

  • The Crawl budget definition explains why serving reliability matters to fetching capacity.

    A robots-file failure can affect more than one service page because it sits before ordinary page access decisions. Compare the failure timing with host availability evidence instead of editing unrelated service content.

How can caching affect a rule change?

Caching can cause a previously fetched policy to remain relevant after the origin file changes. Google’s specification describes its caching behavior, while an edge platform can also preserve an older public representation. Check both delivery and processing boundaries before concluding that a newly edited rule was immediately read by every crawler.

  • First verify the current public file.

    Compare its contents with the intended generated output. If the edge still serves old text, the implementation has a delivery issue to resolve. A local file diff does not establish that the public policy changed. Record the request conditions and available cache evidence.

  • Next separate Google’s previous observation from the live representation.

    A platform report may describe an earlier restriction. A request log may show the file being fetched at another time. These observations can coexist with a correct current response. The change record should identify which version each observation could have covered.

  • Do not repeatedly rewrite rules merely to provoke an update.

    Choose a stable policy, deliver it correctly, and use appropriate platform checks. Frequent temporary toggles can make it harder to establish what a later page request was allowed to fetch. That uncertainty reduces the usefulness of both logs and inspection reports.

  • Log-file analysis can help when the relevant records are available.

    It cannot prove absent requests from an incomplete origin export, particularly if the edge answers the file directly. The evidence needs its delivery-layer scope and retention period to support a precise conclusion.

Why must rendering resources remain accessible?

Rendering resources can be necessary for a crawler to see the intended public explanation and navigation. Blocking the main HTML is not the only way to disrupt a page. A restriction on a script, stylesheet, or data request can leave the document reachable while its rendered information is incomplete.

Why must rendering resources remain accessible?
Point to considerExplanation and application
Google’s JavaScript processing guidance, accessed October 8, 2026, explains that blocked pages or files are not rendered through the normal process.Inspect dependency paths when a service explanation appears in the browser but is missing from the crawler’s tested output.
A public navigation script may create important destinations.A service-data endpoint may provide the main text. A stylesheet may affect whether content is visible or legible. Identify which dependency actually changes the observed output before allowing an entire directory or assuming every blocked file needs public access.
JavaScript SEO provides the broader distinction between initial response and rendering.Test a public visit without the owner’s session or preexisting cache. A successful logged-in view can conceal a dependency permission problem that ordinary customers and automated requests encounter.

How should robots.txt interact with indexing directives?

A crawl rule and an indexing directive need a coherent relationship. Google must reach a page to read a meta tag or response header on it. Blocking that access can hide a new exclusion or canonical preference. The correct arrangement follows the intended outcome rather than a habit of adding every available restriction together.

  • Google’s noindex implementation guidance, accessed October 8, 2026, explicitly requires crawler access to the instruction.

    A public utility meant to leave search may therefore need an accessible response carrying noindex. Disallowing it first can prevent the desired new directive from being processed.

  • Noindex is not a crawl-saving mechanism in itself.

    The crawler requests the page to read it. That can be appropriate for the indexing requirement even when the business also wants a controlled generated inventory. Define which problem the route presents before choosing the response and crawl policy.

  • Canonical preferences also require readable evidence.

    A blocked duplicate can prevent Google from reading its relationship to a preferred page. Canonical tags describe equivalent-content preferences; they do not work as a privacy control. Avoid using robots restrictions to force representation without inspecting the actual page relationship.

  • Make the requirements explicit in the implementation record.

    A path permanently unwanted for automatic crawling, a public utility excluded from search, and a private document require different arrangements. The broad word block does not capture those differences, and a rule can appear to succeed in one sense while failing another.

How do sitemap declarations fit into the file?

A sitemap declaration points crawlers to a sitemap location. It is a discovery signal, separate from allow and disallow path rules. The destination should be a complete address for a valid sitemap or index file, and its availability and entries need their own checks rather than being assumed correct because a line exists.

  • A generated site can declare more than one sitemap where that fits its inventory.

    The line is not tied to one user-agent group under Google’s described processing. Do not place it among rules expecting it to reset the group or act as an access exception for every URL the sitemap lists.

  • The XML sitemap definition explains its inventory role.

    Preferred public URLs should be consistent with the site’s intended indexing instructions. A listed page can still be disallowed or excluded. The sitemap does not override those controls or require Google to treat an entry as indexed.

  • Open the declared file and test representative entries.

    A stale host, broken sitemap path, or list of old redirects can weaken the intended inventory description. Fix the relevant destination rather than adding another line to the robots file that points to the same unavailable resource.

  • Treat discovery and access separately during launch checks.

    The declaration can remain correct while a copied root restriction blocks all service requests. The access rules can be correct while the declaration points to an obsolete list. Both deserve direct tests because neither proves the other is working.

An illustrative diagnosis: a specific group changes a launch policy

Continue the public-page review

Start with the website SEO checker for a preliminary page review. Request the host file and test its actual rules separately. Our SEO services connect that access policy with the public resources and service routes the business needs.

Questions about Robots.txt

Does disallow remove a URL from Google?

Not reliably. A blocked URL can remain known and appear in search without Google crawling its content.

Robots.txt Introduction and Guide ↗
Can noindex be placed in robots.txt?

Google does not support noindex as a robots.txt directive. Use a supported page meta tag or HTTP header while allowing the crawler to read it.

Block Search Indexing with noindex ↗
Which user-agent group applies?

Google chooses the most specific matching user-agent group. Inspect that group's rules rather than assuming the wildcard group controls every crawler.

How Google Interprets the robots.txt Specification ↗
How do Allow and Disallow conflicts resolve?

The longest matching path determines the rule; for equally specific conflicting Allow and Disallow rules, Google uses the less restrictive rule.

How Google Interprets the robots.txt Specification ↗

Continue learning

Practical reading

Try a relevant tool

Sources

Robots.txt Introduction and Guide | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026How Google Interprets the robots.txt Specification | Google Crawling Infrastructure  |  Crawling infrastructure  |  Google for Developers ↗Accessed October 8, 2026Understand JavaScript SEO Basics | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026Block Search Indexing with noindex | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026

Published . Definitions and examples link to their supporting sources. Our SEO methodology →

SEO · Content · Local · Web Design

Connect the website work to your business.

We assess the pages, search demand, and customer actions that matter to your business, then explain where to focus the work.