noindex vs nofollow vs canonical: The Complete Decision Matrix

Clear decision matrix for noindex, nofollow, and canonical tags. Learn which directive to use for every scenario with code examples and common mistake fixes.

BugViso

14 min read

Three directives control how Google processes your pages — noindex, nofollow, and rel="canonical" — and misusing any one of them creates cascading indexation problems that are difficult to diagnose after the fact. The confusion is understandable: all three appear to "manage" how pages show up in search, but they operate on entirely different mechanisms, affect different stages of Google's pipeline, and produce different side effects when combined incorrectly. Using noindex when you need a canonical consolidation de-indexes a page that should rank. Using nofollow when you need noindex does nothing to prevent indexation. Using a canonical when you need noindex signals a preference Google may override.

This guide provides the definitive decision matrix — a scenario-by-scenario lookup table that maps your specific situation to the correct directive, with code examples, common mistakes, and the exact behavior each directive triggers in Googlebot's crawling, rendering, and indexing pipeline.

The Three Directives: What Each Actually Does

noindex — Prevents Indexation, Does Not Prevent Crawling

The noindex directive tells Googlebot: "Crawl this page, read its content, but do not include it in the search index." Googlebot still fetches and renders the page — it just doesn't store it in the searchable index.

html
<!-- Implementation: meta robots tag -->
<meta name="robots" content="noindex" />

<!-- Implementation: X-Robots-Tag HTTP header (for non-HTML resources) -->
X-Robots-Tag: noindex

Key behaviors:

  • Googlebot continues to crawl the page periodically to verify the noindex directive is still present
  • Link equity on the page does flow outward to linked pages (unless combined with nofollow)
  • The page will be removed from the index within days to weeks after noindex is added
  • noindex is a directive — Google obeys it, not just treats it as a hint

The nofollow attribute tells Googlebot: "Do not pass PageRank equity through this link." Since March 2020, Google treats nofollow as a hint rather than a directive — Google may still follow the link and index the destination page.

html
<!-- Page-level: all outbound links on this page are nofollowed -->
<meta name="robots" content="nofollow" />

<!-- Link-level: this specific link is nofollowed -->
<a href="/user-generated-page" rel="nofollow">User Page</a>

<!-- Link-level: also prevents sponsored/UGC equity passing -->
<a href="/sponsored-page" rel="sponsored">Sponsor</a>
<a href="/ugc-comment-link" rel="ugc">Comment Link</a>

Key behaviors:

  • nofollow does NOT prevent a page from being indexed — it only affects link equity transfer
  • Google may still follow and index the destination URL through other discovery paths (sitemaps, other links)
  • Page-level nofollow affects all outbound links on the page
  • Link-level nofollow affects only the specific <a> element it's applied to

rel="canonical" — Declares the Preferred Version of Duplicate Content

The canonical tag tells Googlebot: "This page is a duplicate or variant of another URL — please index that URL instead, and consolidate ranking signals toward it."

html
<!-- Self-referencing canonical (every indexable page should have one) -->
<link rel="canonical" href="https://example.com/products/widget-pro" />

<!-- Cross-domain canonical (less commonly used, less reliably obeyed) -->
<link rel="canonical" href="https://other-site.com/original-article" />

Key behaviors:

  • Canonical is a strong signal, not a directive — Google may override it if its own signals disagree
  • The canonical URL should return 200 OK and contain substantially similar content
  • Googlebot continues to crawl the non-canonical URL periodically
  • Link equity on the non-canonical page is consolidated toward the canonical URL
  • Self-referencing canonicals are a best practice on every indexable page

The Decision Matrix: 18 Scenarios Mapped to Correct Directives

#ScenarionoindexnofollowcanonicalCorrect Approach
1Paginated list page (page 2+)○✗Self-referencingEach page canonicalizes to itself; products on each page are unique
2URL parameter variant (?sort=price)○✗→ Base URLCanonical to the unparameterized version
3Staging/dev environment page✓○✗noindex + password protection + robots.txt Disallow
4Internal search results✓✗✗noindex, follow + robots.txt Disallow for crawl savings
5Expired job/event listing✗✗✗Return HTTP 410 Gone status code
6Thank-you / confirmation page✓✗✗noindex — no search value, contains user data
7Login / account dashboard✓✗✗noindex — private content, no search value
8Print-friendly page variant✗✗→ Main URLCanonical to the standard version
9AMP variant✗✗→ Non-AMPCanonical to the non-AMP version (per Google's AMP guidance)
10HTTP → HTTPS duplicate✗✗→ HTTPS301 redirect + canonical to HTTPS
11www → non-www duplicate✗✗→ Preferred301 redirect to preferred domain
12User-generated content page○On linksSelf-referencingnofollow on user-submitted links; noindex only if content is thin
13Thin tag/category page (0–2 items)✓✗✗noindex, follow until the page has substantial content
14PDF / downloadable document○✗✗Use X-Robots-Tag: noindex header if you don't want PDFs indexed
15Faceted navigation (multi-filter)○✗→ Base categoryCanonical to the single-facet or base category URL
16Syndicated / republished content✗✗→ OriginalCross-domain canonical to the original publisher
17Legal pages (privacy, terms)✗✗Self-referencingIndex them — they demonstrate legitimacy and E-E-A-T
18Archived / outdated blog post✗✗Self-referencingKeep indexed; add an editorial note with date — historical content retains value

Key: ✓ = Use this directive, ✗ = Do not use, ○ = Optional depending on context

The Five Most Dangerous Misconfigurations

Mistake 1: Using noindex When You Need a Canonical

The scenario: Your product page exists at both /products/widget and /products/widget?ref=homepage. You want to consolidate them.

html
<!-- ❌ WRONG: noindex on the parameter variant -->
<!-- /products/widget?ref=homepage -->
<meta name="robots" content="noindex" />

Why it's wrong: noindex prevents the parameter variant from being indexed, but it does NOT consolidate ranking signals toward /products/widget. Any backlinks, social shares, or internal links pointing to the ?ref=homepage URL lose their equity entirely instead of being transferred to the canonical.

html
<!-- ✅ CORRECT: canonical to the clean URL -->
<!-- /products/widget?ref=homepage -->
<link rel="canonical" href="https://example.com/products/widget" />

The canonical tag consolidates equity — noindex discards it.

Mistake 2: Using nofollow to Prevent Indexation

The scenario: You have a user-generated content page that you don't want indexed.

html
<!-- ❌ WRONG: nofollow does NOT prevent indexation -->
<meta name="robots" content="nofollow" />

Why it's wrong: nofollow only affects outbound link equity. The page itself can still be discovered via sitemaps, other internal links, or external links — and will be indexed if Google finds it valuable enough.

html
<!-- ✅ CORRECT: noindex prevents indexation -->
<meta name="robots" content="noindex, follow" />

Mistake 3: Combining noindex With a Canonical to a Different URL

The scenario: A paginated page is both noindex and canonicalizes to the first page.

html
<!-- ❌ CONFLICTING: noindex + canonical to different URL -->
<!-- /products?page=3 -->
<meta name="robots" content="noindex" />
<link rel="canonical" href="https://example.com/products" />

Why it's wrong: These directives conflict. noindex says "don't index this page." The canonical says "index /products instead and consolidate signals." Google's documented behavior is to treat this as ambiguous — it may follow one or the other. In practice, Google often respects the noindex and ignores the canonical, meaning no equity consolidation occurs.

The fix depends on your goal:

html
<!-- Goal: Don't index page 3, DO consolidate equity → Use canonical only -->
<link rel="canonical" href="https://example.com/products" />

<!-- Goal: Don't index page 3, DON'T care about equity → Use noindex only -->
<meta name="robots" content="noindex, follow" />

Pick one directive. Don't combine them on the same page.

Mistake 4: Using robots.txt Disallow Instead of noindex

The scenario: You want to remove a page from search results.

text
# ❌ WRONG: Disallow prevents crawling, NOT indexation
User-agent: *
Disallow: /private-page

Why it's wrong: If Google has already discovered /private-page through any means (an old sitemap, an external backlink, a cached internal link), the robots.txt Disallow prevents Googlebot from crawling the page — but Google may still index the URL based on anchor text, title tag from cached data, or surrounding link context. The page can appear in search results with a "No information available for this page" snippet.

html
<!-- ✅ CORRECT: noindex tells Google to remove it from the index -->
<!-- Requires the page to be crawlable so Google can read the directive -->
<meta name="robots" content="noindex" />

The paradox: for noindex to work, the page must be crawlable. If you block it with robots.txt, Googlebot can never see the noindex tag.

Mistake 5: Mass-Applying noindex to Thin Content Without Fixing the Root Cause

The scenario: Google Search Console reports 500 pages as "Crawled - currently not indexed." The team adds noindex to all 500 pages.

Why it's wrong: Adding noindex to content Google was already declining to index is redundant — it just formalizes what Google already decided. The root cause (thin content, poor internal linking, duplicate content) remains unfixed, and the noindex prevents those pages from ever recovering even if the content is later improved.

The correct approach: Investigate why Google rejected each page category. Improve content depth, add internal links, or consolidate near-duplicates. Only apply noindex to pages that genuinely should never appear in search (login pages, confirmation pages, internal tools).

Implementation Reference: Code Patterns

Meta Robots Tag (HTML)

html
<!-- Index and follow (default behavior — tag is optional) -->
<meta name="robots" content="index, follow" />

<!-- Don't index, but follow links for equity flow -->
<meta name="robots" content="noindex, follow" />

<!-- Don't index AND don't follow links -->
<meta name="robots" content="noindex, nofollow" />

<!-- Index the page but don't follow outbound links -->
<meta name="robots" content="index, nofollow" />

<!-- Target specific bots -->
<meta name="googlebot" content="noindex" />
<meta name="bingbot" content="noindex" />

X-Robots-Tag HTTP Header (for PDFs, images, API responses)

nginx
# Nginx: noindex for all PDF files
location ~* \.pdf$ {
    add_header X-Robots-Tag "noindex, nofollow" always;
}

# Nginx: noindex for specific paths
location /internal-tools/ {
    add_header X-Robots-Tag "noindex" always;
}
python
# FastAPI: X-Robots-Tag on API responses
from fastapi import Response

@app.get("/api/preview/{id}")
async def preview(id: str, response: Response):
    response.headers["X-Robots-Tag"] = "noindex"
    return {"preview": "data"}

Canonical Tag Patterns

html
<!-- Self-referencing canonical (every indexable page) -->
<link rel="canonical" href="https://example.com/products/widget-pro" />

<!-- Dynamic canonical in Next.js -->
<!-- pages/products/[slug].tsx -->
<Head>
  <link rel="canonical"
    href={`https://example.com/products/${slug}`}
  />
</Head>
typescript
// Next.js App Router: canonical in metadata
export async function generateMetadata({ params }) {
  return {
    alternates: {
      canonical: `https://example.com/products/${params.slug}`,
    },
  };
}

How BugViso Audits Directive Configurations

BugViso's Advanced SEO Intelligence engine audits the canonical and indexability directives on every page crawled during a multi-page site scan.

The canonicalization and crawl-budget protection module compares each page's actual URL against its declared <link rel="canonical">, flagging three high-risk misconfigurations: protocol mismatches (HTTP canonical on an HTTPS page), homepage canonicalization (subpages that incorrectly canonicalize to /, effectively requesting their own de-indexation), and missing self-referencing canonicals on indexable pages.

The AI Search Readiness (GEO) engine flags noindex directives via both <meta name="robots"> and the X-Robots-Tag HTTP header. If a page that appears in your sitemap or receives strong internal links also carries a noindex directive, BugViso surfaces the conflict — the page is simultaneously signaled as important (via sitemap/links) and excluded (via noindex).

The duplicate content detection engine identifies near-duplicate pages via SimHash analysis. When two pages have >90% content similarity but neither uses a canonical tag pointing to the other, BugViso reports the duplication cluster — these are the exact scenarios where a canonical tag is needed but missing, causing Google to make its own (often incorrect) canonical selection.

Interaction Matrix: What Happens When Directives Combine

CombinationGooglebot BehaviorRecommendation
noindex aloneCrawls page, reads content, does not index✅ Correct for pages that should never appear in search
nofollow aloneMay still follow links (hint), may still index page⚠️ Rarely useful alone — usually combine with noindex
noindex, nofollowDoes not index, does not pass equity through links✅ Correct for private pages with user-submitted links
noindex, followDoes not index, does pass equity through links✅ Correct for paginated/filtered pages you want equity to flow from
canonical to selfReinforces this URL as the preferred indexable version✅ Best practice on every indexable page
canonical to other URLConsolidates equity toward the canonical target✅ Correct for duplicate/variant pages
noindex + canonical to selfContradictory — noindex wins, canonical is ignored❌ Remove one directive
noindex + canonical to otherContradictory — Google may follow either❌ Choose canonical-only OR noindex-only
robots.txt Disallow + noindexGooglebot can't crawl, so it never sees noindex❌ Remove the Disallow if you want noindex to work

Frequently Asked Questions

Yes — by default. A noindex page without nofollow allows link equity to flow through its outbound links to other pages. If you want to block equity flow, use noindex, nofollow. Most practitioners use noindex, follow because they want the page excluded from search but still want its internal links to contribute to the site's link graph.

How long does it take for noindex to remove a page from Google?

Typically 1–4 weeks after Googlebot's next crawl of the page. You can accelerate removal by requesting re-crawl via GSC's URL Inspection tool and using the Removals tool for urgent removals (temporary 6-month hide while noindex takes effect permanently).

Can Google override my canonical tag?

Yes. Canonical tags are strong signals, not directives. If Google's algorithms determine that the declared canonical URL is incorrect (e.g., the canonical target returns a 404, contains completely different content, or has weaker signals than the declaring page), Google will choose its own canonical. Check GSC's URL Inspection tool to see which canonical Google selected.

Should every page have a canonical tag?

Yes — every indexable page should have a self-referencing canonical tag pointing to its own URL. This prevents Google from arbitrarily selecting a different URL as the canonical (e.g., a parameter variant that leaked into the crawl graph). Self-referencing canonicals are a defensive best practice.

What's the difference between rel="nofollow", rel="sponsored", and rel="ugc"?

All three hint to Google that link equity should not pass through the link. nofollow is the generic directive. sponsored specifically identifies paid/advertising links (required by Google's guidelines). ugc identifies user-generated content links (comments, forum posts). Google treats all three as equity-blocking hints, but the distinction helps Google understand the reason the link is marked.

Does noindex prevent a page from appearing in Google Discover or Google News?

Yes — noindex excludes the page from all Google surfaces, including Search, Discover, and News. If you want the page excluded from Search but available in Discover, there is no directive-level way to achieve this; noindex is all-or-nothing.

Conclusion

The correct directive depends on your goal — noindex removes pages from search, canonical consolidates duplicates, and nofollow reduces equity transfer — and applying the wrong one creates the silent indexation problems that BugViso's multi-page audit surfaces by comparing every page's canonical declaration, noindex directives, and duplicate content signals against its actual crawl position in the site-wide link graph.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.