Website

Check the headers, not just the HTML: the noindex you can't see in page source

Page source can miss an indexing restriction in response headers. Follow the published Solar Software Company example and check headers, HTML, cache layers, and conflicting directives.

By Ritik Namdev · Published · Updated · 7 min read

Most SEO checks look at a page’s HTML: the title, the meta robots tag, the canonical. But search engines also read HTTP response headers, and a header can tell Google not to index a page while the HTML looks perfectly fine.

Our Solar Software Company case study reports this problem on a solar design SaaS website we work with.

What went wrong

Two separate problems hid large parts of the site from Google in August and September 2026:

  1. A site-wide header. For several weeks every page was served with X-Robots-Tag: noindex, nofollow, noarchive. Nothing in the page source showed it.
  2. An over-broad safeguard. A content gate designed to noindex pages containing risky claims matched one sentence in a shared template, so it noindexed over a thousand blog posts at once.

How we found it

The HTML was clean, so the checks had to go one layer down:

  • curl -I https://example.com/page/ requests headers with HTTP HEAD. It is a useful first check. Confirm the normal GET response too, because servers can handle these methods differently. A noindex directive in either the applicable response header or the HTML can prevent indexing.
  • Search Console’s URL Inspection shows “Indexing allowed? No: ‘noindex’ detected in ‘X-Robots-Tag’ http header” when this is the cause.
  • Counting pages by indexability state in the build output showed how many pages the content gate had excluded.

What we changed

  • The header was removed (resolved 17 September 2026).
  • The content gate was scoped to the specific sentence instead of the whole template. Gated pages fell from 1,812 to 213, and URLs in the sitemap rose from 1,361 to 2,723.
  • A regression test now fails the build if the gate over-matches again.

How to check your own site

  • Test the normal GET response for your homepage and each important template. Look for X-Robots-Tag as well as the HTML robots directive.
  • Check that headers set at the CDN or hosting layer match what your code intends. Staging rules sometimes leak into production.
  • Treat every automated rule that can noindex pages as code that needs a test.
  • After any deploy that touches headers, routing or templates, inspect a few URLs in Search Console.

How do you inspect the response visitors actually receive?

Use GET when confirming an indexing directive. HEAD can reveal a problem, but it does not prove that the normal page response has the same rules.

This command downloads the page body to a local file and prints the response headers:

curl -sS -D - -o /tmp/page-response.html https://example.com/page/

Replace the example URL with the exact public URL you want indexed. Keep the protocol, hostname, path, and trailing slash intact. These details can change which route or redirect you test.

Inspect the HTTP status first. A successful response, a redirect, and an error need different diagnoses. Then look for every X-Robots-Tag header. Multiple header lines can carry separate directives; reading only the first can miss a restriction.

Next, inspect the saved HTML for robots meta tags. An index instruction in one place does not cancel an applicable noindex instruction elsewhere. Google applies the more restrictive rule when directives conflict.

For a redirected URL, inspect the chain and its final destination:

curl -sS -L -D - -o /tmp/page-response.html https://example.com/page/

The output includes headers from each response. Keep each header block with its status and location. Otherwise, a directive on an intermediate response can be mistaken for a directive on the final page.

These checks explain the live response. They do not show what Google last crawled. Compare them with the last crawl information in URL Inspection before assuming the issue is still present.

What if your code and public headers disagree?

Trace the response through the systems that can change it. A clean source file does not establish that the public response is clean.

Start with the page model and shared layout. Check whether the directive comes from frontmatter, a generated tag, or a condition applied to an entire template. Next, inspect hosting rules, middleware, and CDN configuration. A hosting rule can add a header after the application generates the HTML.

A cached response can also retain an older rule. Compare the exact public URL after the intended change. Record cache-related headers alongside the robots headers so the engineering team can distinguish an old response from a current configuration problem.

If you can test the origin separately, compare it with the public response. Keep the hostname and routing context correct. A test against a different host may reach a different application or security policy and give a misleading result.

Do not remove every restriction merely because some pages should be indexed. Account screens, staging routes, internal search pages, and unfinished content may need deliberate exclusions. Identify the affected page family and narrow the unintended rule.

How do crawling rules affect the diagnosis?

Google must be able to crawl a page to see its noindex directive. Blocking the same URL in robots.txt can prevent that instruction from being read.

That distinction matters when retiring an indexed page. A crawl block is not a substitute for a readable noindex instruction. Choose the intended outcome first: prevent crawling, prevent indexing, redirect a replacement, or return an appropriate removal status.

A canonical tag also serves a different purpose. It indicates a preferred representative among similar URLs. It should not be treated as a switch that reverses a noindex rule.

Review these signals together when troubleshooting indexing exclusions. Check the response, directives, canonical target, internal links, and sitemap entry. A mismatch between them creates avoidable ambiguity even when each individual setting looks plausible.

What should you verify after the fix?

Verify both the affected pages and pages that should remain excluded. Testing only the newly indexable group can miss a rule that was broadened too far in the opposite direction.

For each affected template, retain one representative URL and one relevant edge case. Include an intentionally excluded page where that exclusion is part of the design. Record the expected status, robots directives, and canonical target for each.

Run those checks against the built output and the actual public response. A build test can catch template mistakes. A public-response check catches hosting and cache behavior that the build cannot observe.

Then use Google Search Console to inspect representative URLs. A live test and the indexed-page record answer different questions. The live test checks current accessibility; the indexed record reflects Google’s processing history.

Removing an unintended restriction restores eligibility. It does not establish that Google will index the page, choose its canonical, or rank it. Continue checking the Page indexing report, important URL inspections, and the consistency of your sitemap and internal links.

What should the regression test assert?

Assert the intended rule for each page family. A test that checks only whether any noindex exists is too blunt for a site with legitimate exclusions.

For an indexable article, check the successful response, absence of an applicable noindex directive, and expected canonical. For an excluded internal page, check that its deliberate restriction remains present. Include both HTML directives and response headers where the test environment can observe them.

Keep the fixture small enough to understand. When a shared template changes, a failing test should identify the page family and unexpected directive. That makes the failure actionable instead of encouraging someone to disable the test.

The practical lesson is straightforward: inspect the whole response, preserve deliberate exclusions, and verify the public result. A clean page source is only one part of the evidence.

Sources

SEO · Content · Local · Web Design

Build a website search engines can trust.

We'll look at your business and search landscape and show you where to start.