Why Google Isn't Indexing Your Pages: 12-Point Diagnostic Guide
Diagnose why Google isn't indexing your pages with this 12-point diagnostic flowchart. Cover noindex tags, canonical conflicts, crawl budget, thin content, and server errors.
Google's index is selective. Of the trillions of URLs it discovers, Google actively chooses NOT to index a significant percentage — and the reasons range from explicit directives you set (noindex tags, robots.txt blocks) to implicit quality signals (thin content, duplicate content, low authority). When you see "Crawled - currently not indexed" or "Discovered - currently not indexed" in Google Search Console, the cause is always one of 12 specific technical or content-quality failures. This diagnostic guide walks through each cause in order of likelihood, with exact validation commands and production fixes for each.
Start at Checkpoint 1. If that checkpoint passes, move to the next. The first failure you hit is almost certainly your primary indexation blocker.
The 12-Point Diagnostic Sequence
| Checkpoint | Category | Common Status in GSC |
|---|---|---|
| 1 | Noindex directive | Excluded by 'noindex' tag |
| 2 | Robots.txt block | Blocked by robots.txt |
| 3 | Canonical conflict | Duplicate, Google chose different canonical |
| 4 | Server errors | Server error (5xx) |
| 5 | Redirect issues | Page with redirect |
| 6 | Crawl budget exhaustion | Discovered - currently not indexed |
| 7 | Orphan pages | Crawled - currently not indexed |
| 8 | Thin or duplicate content | Crawled - currently not indexed |
| 9 | Manual action | Manual actions (GSC) |
| 10 | JavaScript rendering failure | Crawled - currently not indexed |
| 11 | URL parameter pollution | Crawled - currently not indexed |
| 12 | New site / low authority | Discovered - currently not indexed |
Checkpoint 1: Noindex Directive
The most common cause of intentional non-indexation — and the most frequent accidental cause on staging-to-production migrations.
How to Check
# Check for noindex meta tag in HTML
curl -s https://example.com/your-page/ | grep -i "noindex"
# Check for X-Robots-Tag in HTTP headers
curl -I https://example.com/your-page/ 2>/dev/null | grep -i "x-robots-tag"Expected output when noindex is present:
<meta name="robots" content="noindex, follow" />or
X-Robots-Tag: noindexCommon Causes
- Staging configuration leaked to production: CMS or framework has a "discourage search engines" setting enabled.
- Plugin or middleware conflict: SEO plugins, caching plugins, or custom middleware adding noindex headers globally instead of selectively.
- Environment variable misconfiguration:
ROBOTS_META=noindexset in.envfile from staging, carried to production deployment.
How to Fix
# WordPress: Check Settings → Reading → "Discourage search engines"
# Ensure checkbox is UNCHECKED on production
# Next.js: Check for noindex in layout or page metadata
grep -r "noindex" app/ pages/ --include="*.tsx" --include="*.ts"
# Nginx: Check for X-Robots-Tag header injection
grep -r "X-Robots-Tag" /etc/nginx/<!-- ✅ Correct: Only noindex pages that should not be indexed -->
<!-- On search result pages -->
<meta name="robots" content="noindex, follow" />
<!-- On indexable pages — either omit the tag or explicitly set index -->
<meta name="robots" content="index, follow" />Checkpoint 2: Robots.txt Block
If robots.txt disallows Googlebot from crawling a URL, Google cannot access the page content and will not index it.
How to Check
# Download and inspect robots.txt
curl -s https://example.com/robots.txt
# Check if your specific URL path is blocked
# Look for Disallow directives that match your URL patternUse the robots.txt Tester in Google Search Console to verify whether a specific URL is blocked.
Common Causes
# ❌ Overly broad disallow that blocks legitimate content
User-agent: *
Disallow: /blog/ # Blocks ALL blog content
# ❌ Wildcard pattern that catches more than intended
User-agent: *
Disallow: /*? # Blocks ALL URLs with query parametersHow to Fix
# ✅ Precise disallow rules that protect crawl budget without blocking content
User-agent: *
Disallow: /search
Disallow: /admin/
Disallow: /*?sessionid=
Disallow: /*?sort=
Allow: /blog/
Allow: /products/💡 Critical distinction:
robots.txtprevents crawling, not indexing. If a blocked page has external backlinks, Google may index the URL (showing it in search results with no snippet) even though it cannot crawl the content. To prevent indexing, usenoindex— notrobots.txt.
Checkpoint 3: Canonical Conflict
If your page declares a canonical URL that points to a different page, Google will index the canonical target instead of your page.
How to Check
# Extract canonical tag from the page
curl -s https://example.com/your-page/ \
| grep -oP '<link[^>]*rel="canonical"[^>]*href="\K[^"]+'
# Verify the canonical points to the page itself (self-referencing)
# If it points elsewhere, Google indexes the canonical target insteadCommon Causes
<!-- ❌ All pages canonicalize to homepage (most destructive bug) -->
<link rel="canonical" href="https://example.com/" />
<!-- ❌ Canonical points to a different, non-equivalent page -->
<link rel="canonical" href="https://example.com/old-page/" />
<!-- ❌ Protocol mismatch (HTTP canonical on HTTPS page) -->
<link rel="canonical" href="http://example.com/your-page/" />
<!-- ❌ Trailing slash mismatch -->
<link rel="canonical" href="https://example.com/your-page" />
<!-- When the actual URL is https://example.com/your-page/ -->How to Fix
<!-- ✅ Self-referencing canonical on every indexable page -->
<link rel="canonical" href="https://example.com/your-page/" />
<!-- Match protocol, domain, and trailing slash exactly -->For a deep dive into canonical issues, see our guide on canonical tags and how to avoid duplicate content.
Checkpoint 4: Server Errors (5xx)
If Googlebot receives a 5xx error when crawling your page, it cannot index the content. Persistent 5xx errors cause Google to reduce crawl frequency and eventually drop the page from the index.
How to Check
# Simulate Googlebot's request
curl -I -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
https://example.com/your-page/
# Check for 5xx status codes
# HTTP/2 500 → Internal Server Error
# HTTP/2 502 → Bad Gateway
# HTTP/2 503 → Service UnavailableCommon Causes
- Database connection timeouts under high crawl load
- Memory exhaustion from rendering complex pages
- Rate limiting that incorrectly blocks Googlebot's IP ranges
- CDN misconfigurations returning 502/503 during deployments
How to Fix
# Check server error logs for the specific URL pattern
grep "500\|502\|503" /var/log/nginx/error.log | tail -20
# Verify the page loads under simulated crawler load
ab -n 100 -c 10 -H "User-Agent: Googlebot" https://example.com/your-page/Checkpoint 5: Redirect Issues
Pages that redirect are not indexed — Google indexes the redirect target instead. Redirect chains (3+ hops) may cause Google to abandon the crawl entirely.
How to Check
# Follow all redirects and show each hop
curl -sIL https://example.com/your-page/ 2>&1 | grep -E "^HTTP|^location"
# Expected output for a redirect chain:
# HTTP/2 301
# location: https://example.com/intermediate-page/
# HTTP/2 301
# location: https://example.com/final-page/
# HTTP/2 200How to Fix
Every redirect should go directly from source to final destination in a single hop. Collapse redirect chains using 301 (permanent) or 308 (permanent, preserves method) redirects. For more details, see our guide on redirect chain auditing.
Checkpoint 6: Crawl Budget Exhaustion
For large sites (100K+ URLs), Googlebot may discover your page but not crawl it within a reasonable timeframe due to crawl budget constraints.
How to Diagnose
In Google Search Console, navigate to Pages → filter by "Discovered - currently not indexed." This status means Google knows the URL exists (from a sitemap or internal link) but has not yet allocated crawl budget to fetch it.
How to Fix
- Submit the URL directly via the URL Inspection tool in GSC
- Ensure the page has strong internal links (not an orphan page)
- Improve server TTFB to increase Googlebot's crawl rate
- Remove crawl traps that waste budget on low-value URLs
Checkpoint 7: Orphan Pages (No Internal Links)
Pages with zero inbound internal links receive no internal link equity and are crawled infrequently, making indexation unlikely.
How to Diagnose
Compare your sitemap URLs against your crawled URL set. URLs present in the sitemap but not discoverable via link-following are orphan pages.
# Quick orphan page check
curl -s https://example.com/sitemap.xml | grep -oP '<loc>\K[^<]+' | sort > sitemap.txt
# Compare with a crawl of your site's internal linksHow to Fix
Add contextual internal links from topically related pages. Ensure the page appears in at least one navigation element (category listing, related posts, breadcrumbs). See our detailed guide on orphan pages and how to find them.
Checkpoint 8: Thin or Duplicate Content
Google may crawl a page and choose not to index it if the content does not provide sufficient unique value — it is too short, too similar to other indexed pages, or does not satisfy the search intent for any query. Our guides to finding and fixing thin content and to running a duplicate content checker cover both cases in depth.
How to Diagnose
# Check word count of the page's main content
curl -s https://example.com/your-page/ \
| python3 -c "
import sys
from html.parser import HTMLParser
class TextExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.text = []
self.skip = False
def handle_starttag(self, tag, attrs):
if tag in ('script', 'style', 'nav', 'footer', 'header'):
self.skip = True
def handle_endtag(self, tag):
if tag in ('script', 'style', 'nav', 'footer', 'header'):
self.skip = False
def handle_data(self, data):
if not self.skip:
self.text.append(data.strip())
parser = TextExtractor()
parser.feed(sys.stdin.read())
words = ' '.join(parser.text).split()
print(f'Word count: {len(words)}')
"Indexation thresholds (approximate):
| Content Depth | Word Count | Indexation Likelihood |
|---|---|---|
| Very Thin | < 200 words | Low — unlikely to be indexed unless highly unique |
| Thin | 200–500 words | Moderate — needs strong authority signals |
| Adequate | 500–1,500 words | Good — if content is unique and relevant |
| Comprehensive | 1,500+ words | High — if content provides genuine value |
How to Fix
- Add unique, substantial content to thin pages (minimum 500 words of original, non-boilerplate content)
- Consolidate duplicate pages using canonical tags or 301 redirects
- Delete pages that cannot be improved to a useful quality level (return 410 Gone)
Checkpoint 9: Manual Action
Google may issue a manual action (penalty) against specific pages or your entire site for violating Google's spam policies.
How to Check
Navigate to Security & Manual Actions → Manual actions in Google Search Console. If a manual action exists, Google describes the specific violation and affected pages.
Common Causes
- Unnatural links (paid links, link schemes)
- Thin content with no added value
- Cloaking (showing different content to Googlebot vs. users)
- User-generated spam (spammy comments, forum posts)
How to Fix
Address the specific violation described in the manual action notice, then submit a reconsideration request via GSC. Processing takes 2–4 weeks.
Checkpoint 10: JavaScript Rendering Failure
If your page relies on client-side JavaScript to render content, Googlebot's Web Rendering Service (WRS) must successfully execute the JavaScript to see the content. Rendering failures cause "Crawled - currently not indexed" because Googlebot sees an empty or incomplete page.
How to Diagnose
Use the URL Inspection tool in GSC → click "Test Live URL" → compare the "Screenshot" and "HTML" tabs. If the rendered HTML is missing content that appears in the browser, JavaScript rendering is failing.
# Compare server-rendered HTML vs. JavaScript-rendered content
# Server HTML (what Googlebot first receives):
curl -s https://example.com/your-page/ | wc -c
# If the server HTML is tiny (< 5KB) and the page is content-rich,
# critical content is likely loaded via client-side JavaScriptHow to Fix
- Implement server-side rendering (SSR) or static site generation (SSG)
- Ensure critical content is in the initial HTML response, not loaded via API calls after page load
- Avoid
setTimeoutorrequestIdleCallbackfor rendering critical content — Googlebot has a limited rendering budget - For framework-specific guidance, see our JavaScript SEO guide
Checkpoint 11: URL Parameter Pollution
URLs with excessive query parameters may be treated as duplicate content or deprioritized by Google's URL normalization algorithms.
How to Diagnose
# Check how many parameter variations Google has discovered
# In GSC: Pages → filter by URL containing "?" → count results
# In server logs:
grep "Googlebot" /var/log/nginx/access.log \
| awk -F'?' '{if(NF>1) print $2}' \
| tr '&' '\n' | cut -d= -f1 \
| sort | uniq -c | sort -rn | head -10How to Fix
- Add canonical tags pointing to the parameter-free URL
- Block non-essential parameters in
robots.txt - Implement server-side parameter stripping (redirect parameterized URLs to clean versions)
Checkpoint 12: New Site or Low Domain Authority
Brand-new websites and pages on low-authority domains face a natural indexation delay. Google has limited crawl resources and prioritizes sites it already trusts.
How to Diagnose
- Domain age: Sites under 6 months old experience slower indexation
- Backlink profile: Sites with zero referring domains have minimal crawl demand
- Content volume: Sites with fewer than 20 pages may not trigger frequent crawling
How to Fix
- Submit your sitemap in Google Search Console
- Build genuine backlinks from relevant, authoritative sources
- Publish content consistently to increase crawl demand
- Use the URL Inspection tool to request crawling of important pages
- Ensure every page has strong internal links from existing indexed pages
How BugViso Runs This Diagnostic Automatically
BugViso's audit engine executes the equivalent of this 12-point diagnostic across every crawled page in a single automated scan.
The AI Search Readiness (GEO) Engine checks indexability by detecting noindex directives in both <meta name="robots"> and the X-Robots-Tag HTTP header, plus Googlebot disallow rules in robots.txt — covering Checkpoints 1, 2, and the intersection of crawler access and indexability.
The Canonicalization & Crawl-Budget Protection audit validates canonical tag correctness across every crawled page — flagging self-referencing failures, protocol mismatches, and the destructive homepage-canonicalization bug that silently de-indexes entire site sections (Checkpoint 3).
The Multi-Page Site Crawl discovers URLs via sitemap.xml and rendered links, then the Internal Link Graph & Orphan Pages module identifies pages with zero inbound internal links and excessive click depth (Checkpoints 6 and 7).
The Duplicate Content Detection module runs SimHash near-duplicate analysis and exact-hash comparison to identify thin and duplicate content across the crawl (Checkpoint 8).
The React Hydration Engine parses console/exception streams for hydration mismatch signatures — a common cause of JavaScript rendering failures that prevent content indexation (Checkpoint 10).
Run a free BugViso scan to execute this entire diagnostic sequence automatically and get a prioritized remediation playbook for every indexation issue found.
Frequently Asked Questions
How long should I wait before worrying about non-indexed pages?
For new pages on established sites, wait 1–2 weeks after submission before investigating. For new pages on new sites, wait 4–6 weeks. If a page remains unindexed after these windows despite being submitted in GSC and having internal links, work through this diagnostic starting at Checkpoint 1.
Can I force Google to index a page?
You can request crawling via the URL Inspection tool in GSC, but you cannot force indexation. Google ultimately decides whether to index based on content quality, technical accessibility, and its assessment of user value. Fixing all 12 checkpoints maximizes your indexation probability but does not guarantee it.
Why does Google say "Crawled - currently not indexed"?
This status means Google successfully crawled the page (no technical blocking) but chose not to include it in the index. The most common causes are: thin content (Checkpoint 8), duplicate content (Checkpoint 8), orphan page with no internal links (Checkpoint 7), or content that does not serve any identified search intent. Improving content depth and building internal links are the primary fixes.
Should I delete pages that Google refuses to index?
Only if the pages serve no business purpose. If a page is genuinely useful for users (e.g., a niche product page, a support article), improve its content quality and internal linking rather than deleting it. If a page is truly thin or duplicate with no path to improvement, either consolidate it via redirect or return a 410 (Gone) status code to remove it cleanly.
Do meta tags like "index, follow" help get pages indexed?
The index, follow directive is the default behavior — Google indexes and follows links by default. Explicitly adding <meta name="robots" content="index, follow"> does not provide any boost over having no meta robots tag at all. It only matters as a correction if a noindex directive exists elsewhere (e.g., in an HTTP header or inherited from a template).
Conclusion
Non-indexed pages always fail one of 12 specific checkpoints — from explicit noindex directives and canonical conflicts to thin content and JavaScript rendering failures — and running through this diagnostic systematically, which is exactly what a free BugViso audit automates across canonical validation, orphan page detection, duplicate content analysis, and crawl accessibility, is the fastest path to getting every important page into Google's index.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.