Thin Content SEO: How to Find and Fix Low-Value Pages
Thin content SEO: why word count misleads, how to find pages with little unique value, and when to expand, merge, noindex or delete. Data from 1,332 pages.
Thin content is a page that offers little or no unique value for the query it targets: a tag page listing one post, a city page with only the place name swapped, a product page copied from the manufacturer, or a stub nobody finished. Low word count is a symptom, not the definition. A 90-word contact page isn't thin. A 900-word page that says nothing new is.
That distinction matters for how you audit. When we measured the main content of 1,332 pages on 57 websites, 32.6% of pages had fewer than 300 words, and 87.7% of sites had at least one such page among their most visible URLs. Most of those pages are fine. The job is to find the few that are genuinely empty of value and decide what to do with each.
This guide covers how Google frames thin content, what our crawl showed, a script that measures unique words (not words padded out by menus and footers), and a decision matrix for every thin page you find.
What Google Actually Means by "Thin"
Google's spam policies target the deliberate forms: doorway pages, scraped content, thin affiliate pages that add nothing to the merchant's own description, and scaled content abuse, meaning many pages generated mainly to rank rather than to help. Its guidance on creating helpful, reliable content covers the broader case: would someone who landed on this page feel they'd found what they were looking for?
Most thin pages on ordinary sites aren't spam. They're accidents of the CMS or the content process:
- Taxonomy pages: tags and categories with one or two posts and no introduction.
- Template-generated pages: locations, integrations or product variants with a few words swapped.
- Stubs: "Coming soon", placeholder service pages, empty FAQ entries.
- Boilerplate-heavy pages: 900 words on the page, 850 of them navigation, footer and cookie text.
- Duplicated descriptions: supplier copy repeated across hundreds of product pages.
Their cost is indirect but real. Pages like these are the usual residents of Search Console's "Crawled โ currently not indexed" report. They also dilute internal links and give AI answer engines nothing worth citing.
๐ก Word count is a filter, not a verdict. Use it to build a shortlist, then judge each page by one question: does it answer something better than the page that already ranks?
What We Found: Main-Content Word Counts on 1,332 Pages
We sampled 57 websites from the Tranco top-sites list (ranks 1,001โ50,000), rendered each homepage and up to 25 pages linked from it in Chromium, and counted words in the main content only. BugViso's content-root logic picks the <main> landmark, the skip-link target or a dominant <article>, and falls back to <body>. Sites whose pages mostly redirected off-site (to consent walls, for instance) were excluded.
| Main-content words | Pages | Share |
|---|---|---|
| Under 50 | 121 | 9.1% |
| Under 100 | 168 | 12.6% |
| Under 300 | 434 | 32.6% |
| Under 600 | 767 | 57.6% |
| Median page | โ | 497 words |
| Site-level finding (57 sites) | Result |
|---|---|
| Sites with โฅ1 sampled page under 300 words | 50 (87.7%) |
| Sites where half or more of sampled pages were under 300 words | 17 (29.8%) |
| Median share of a site's sampled pages under 300 words | 30.4% |
| Median homepage vs median inner page | 541 vs 494 words |
We sorted the short inner pages by URL pattern:
| Type of page under 300 words | Share of short inner pages | Usually a problem? |
|---|---|---|
| Content and product pages ("other") | 73.7% | Often: review each one |
| Category, tag, listing and archive pages | 13.7% | Often: no unique intro, one or two items |
| Login, account, app and download pages | 6.7% | Rarely: their purpose isn't text |
| Contact, about and legal pages | 5.8% | Rarely, unless they're placeholders |
Two takeaways. First, a short page is extremely common even on well-known sites, so "under 300 words" alone would flag a third of a typical site. Second, about one in eight short pages was a taxonomy or listing page. That's the most reliable thin-content category: easy to find by URL pattern, and usually best fixed by merging or noindex rather than by writing more.
โ ๏ธ Limits: we sampled pages linked from homepages (the best-maintained part of a site) from one location, and word counts depend on how well the main-content element can be identified. Deeper archives usually contain more thin pages, not fewer.
Why Raw Word Counts Mislead
A page's raw word count includes everything a template repeats on every URL: mega-menus, footers, cookie notices, newsletter boxes, "related posts". On a content-light page, that boilerplate can be most of the text.
The fix is to measure unique words. Collect every sentence across a sample of pages, treat any sentence that appears on a large share of them as boilerplate, and count what's left on each page. That's what the script below does.
We ran it on our own site first, and it flagged one page:
url,title,total_words,unique_words,boilerplate_pct,flag
https://bugviso.com/about,About BugViso โ AI-Powered Website Audit & QA,219,216,1,THIN
https://bugviso.com/contact-sales,Contact Sales & Enterprise Solutions | BugViso,627,627,0,
https://bugviso.com/scan,Free Website Scan & Bug Finder โ Instant Audit | BugViso,737,737,0,Our About page has 216 unique words. Whether that's a problem depends on its job. An About page doesn't need to rank for a competitive query, but it is an E-E-A-T trust page. A fuller page (who builds the product, how audits are validated, contact details) helps users and AI systems decide whether to trust everything else. So it goes on our "expand" list, not the "delete" list.
Find Thin Pages With Unique-Word Counts
#!/usr/bin/env python3
"""Thin content finder: words of *unique* main content per page.
Raw word counts lie: a 900-word page can be 850 words of mega-menu, footer and
cookie text repeated on every URL. This script strips sentences that appear on
more than 30% of the sampled pages (site boilerplate) and counts what is left.
Flags: THIN < 300 unique words and no strong non-text purpose
TEMPLATE > 70% of the page's text is boilerplate shared with other pages
Usage: pip install httpx beautifulsoup4
python3 thin_content_finder.py https://example.com/sitemap.xml --limit 300 > thin.csv
"""
import argparse, asyncio, csv, re, sys
from collections import Counter
import httpx
from bs4 import BeautifulSoup
WORD = re.compile(r"[A-Za-z0-9]+(?:['-][A-Za-z0-9]+)*")
SPLIT = re.compile(r"(?<=[.!?])\s+|\n+")
async def sitemap_urls(client, url, limit):
urls = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", (await client.get(url)).text)
for child in [u for u in urls if u.endswith(".xml")][:20]:
urls += re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", (await client.get(child)).text)
return list(dict.fromkeys(u for u in urls if not u.endswith(".xml")))[:limit]
def sentences(html):
soup = BeautifulSoup(html, "html.parser")
for tag in soup(["script", "style", "noscript", "svg", "template"]):
tag.decompose()
title = soup.title.get_text(strip=True) if soup.title else ""
text = (soup.body or soup).get_text("\n")
sents = [re.sub(r"\s+", " ", s).strip() for s in SPLIT.split(text)]
return title, [s for s in sents if len(WORD.findall(s)) >= 3]
async def main(sitemap, limit):
async with httpx.AsyncClient(timeout=20, follow_redirects=True,
headers={"User-Agent": "Mozilla/5.0 (compatible; thin-audit/1.0)"}) as client:
urls = await sitemap_urls(client, sitemap, limit)
sem = asyncio.Semaphore(6)
async def get(u):
async with sem:
try:
r = await client.get(u)
if r.status_code == 200 and "html" in r.headers.get("content-type", ""):
return u, *sentences(r.text)
except httpx.HTTPError:
pass
return None
pages = [p for p in await asyncio.gather(*(get(u) for u in urls)) if p]
df = Counter(s for _, _, sents in pages for s in set(sents))
cutoff = max(2, int(len(pages) * 0.3))
w = csv.writer(sys.stdout)
w.writerow(["url", "title", "total_words", "unique_words", "boilerplate_pct", "flag"])
for url, title, sents in sorted(pages, key=lambda p: p[0]):
total = sum(len(WORD.findall(s)) for s in sents)
unique = sum(len(WORD.findall(s)) for s in sents if df[s] < cutoff)
boiler = 0 if not total else round(100 * (total - unique) / total)
flag = "THIN" if unique < 300 else ("TEMPLATE" if boiler > 70 else "")
w.writerow([url, title[:80], total, unique, boiler, flag])
thin = sum(1 for _, _, s in pages if sum(len(WORD.findall(x)) for x in s if df[x] < cutoff) < 300)
print(f"# {len(pages)} pages, {thin} under 300 unique words, boilerplate cutoff = sentence on >= {cutoff} pages",
file=sys.stderr)
if __name__ == "__main__":
ap = argparse.ArgumentParser()
ap.add_argument("sitemap")
ap.add_argument("--limit", type=int, default=300)
a = ap.parse_args()
asyncio.run(main(a.sitemap, a.limit))Open the CSV in a spreadsheet and sort by unique_words. Then add two columns from your analytics, organic clicks (last 12 months) and backlinks, because the right fix depends on them.
The TEMPLATE flag catches a different problem: pages with plenty of words, nearly all of them shared. A 1,000-word page that is 85% boilerplate is effectively a 150-word page. The script needs at least a few dozen pages to tell boilerplate from content, so run it on a sitemap rather than a handful of URLs.
The Thin Content Decision Matrix
Every flagged page gets exactly one of five outcomes:
| Situation | Action | Status / tag |
|---|---|---|
| Has traffic or links, and the topic deserves its own page | Expand: add unique detail, data, examples, FAQs | 200 |
| Overlaps a stronger page on the same topic | Merge the useful parts into the stronger page | 301 to it |
| Needed by users but not as a search result (thin tags, filtered views, thank-you pages) | Keep, but noindex | <meta name="robots" content="noindex"> |
| No traffic, no links, no purpose | Delete | 410 (or 404) |
| Purpose isn't text (login, contact, tool, checkout) | Leave it | 200 |
Some rules of thumb:
- Don't pad. Adding 400 generic words to a thin page makes a longer thin page. Add what's missing: specifications, prices, steps, original data, answers to the questions people ask.
- Merge before you delete anything with backlinks or traffic. Deleting throws those signals away; merging with a 301 keeps them.
- Fix the template, not the page. If 300 location pages are thin, the problem is the template's data model. Add unique fields (local team, opening hours, service area, reviews, photos) or consolidate into one page per region.
noindexisn't a cure-all. It removes the page from results but keeps it crawlable, so Google still spends requests on it. For large volumes of genuinely useless URLs, deletion orrobots.txtblocking may be better. Our guide on when to use noindex, nofollow or canonical covers the trade-offs.
<!-- โ
Tag archive with one post: useful for browsing, not for search -->
<meta name="robots" content="noindex, follow">For content that's thin because it's outdated rather than empty, see our process for updating old blog posts for SEO. When two pages are thin because they split one topic between them, treat it as keyword cannibalization and merge.
Measure the result
Give each batch of changes 8โ12 weeks before judging it, since recrawling and reprocessing take time. Track three numbers for the affected URLs: how many are indexed (Search Console's Pages report), their combined organic clicks, and how many still show as "Crawled โ currently not indexed". A successful clean-up usually shows fewer indexed URLs overall but more clicks per indexed URL, because search engines are spending their attention on the pages that deserve it.
Thin Content and AI Search
AI answer engines retrieve passages, not pages. A page with 120 words of real content gives a retrieval system one small chunk with few facts, so there's little to quote and little reason to cite it. Pages that answer a question completely, with specific numbers, steps and definitions, are the ones that get extracted. Our guide to content extractability for AI search engines covers the structure that helps.
The overlap between "thin for Google" and "thin for AI" is almost total. If a page has no unique substance, neither system has a reason to surface it.
How BugViso Flags Thin Content
BugViso measures the main content of every page it renders, using the same content-root selection described in the methodology above, so menus and footers don't inflate the count. Its AI Search Readiness (GEO) engine raises a "Thin content for AI citation" suggestion when a page's main content is under 300 words, alongside its content-extractability checks for question-style headings, lists, tables and summary blocks.
On a multi-page crawl, the Duplicate Content Detection engine compares the same main-content text across pages with exact hashing and SimHash, which catches the templated variant of thin content: pages that aren't just short but nearly identical to each other. The internal link graph shows which short pages are orphans or buried deep, which often explains why they're "Crawled โ currently not indexed". You can run a free BugViso scan to see which of your pages are flagged.
Mistakes to Avoid
- Mass-deleting by word count. You'll delete contact pages, tools and product pages that rank on structured data.
- Ignoring the template. Fixing 20 pages by hand while the CMS generates 2,000 more.
- Rewriting manufacturer copy with a thesaurus. It's still the same information as every other retailer's page. Add your own: photos, sizing notes, comparisons, reviews.
- Using AI to pad. Generated filler lengthens the page without adding information gain. That's exactly the pattern Google's scaled-content policy describes.
- Forgetting internal links after a merge. Update links to point at the surviving URL, so you don't leave internal redirects behind.
FAQ
How many words does a page need to avoid being thin?
There's no minimum. Google doesn't use a word-count threshold, and many 100-word pages rank well because they answer a short question completely. Use about 300 words of main content as a review trigger, not a rule, then judge each page by whether it offers anything a searcher can't get elsewhere.
Does thin content cause a Google penalty?
Ordinary thin pages don't get a manual penalty. They tend to go unindexed, rank poorly and drag down how Google assesses the site's overall helpfulness. Manual actions are for deliberate cases: doorway pages, scraped content, thin affiliate pages at scale.
Should I noindex or delete thin pages?
noindex pages that users need but searchers don't, such as thin tag archives, filtered views and thank-you pages. Delete (410) pages with no users, no traffic and no links. Merge (301) anything with traffic or backlinks into the strongest related page.
Are category and tag pages thin content?
They can be. A category with a unique introduction, useful filters and many items is a strong page. A tag with one post and no text is thin. In our data, taxonomy and listing pages made up 13.7% of short inner pages.
How do I find thin content in Google Search Console?
Start with Pages โ "Crawled โ currently not indexed" and "Discovered โ currently not indexed". Those lists often overlap heavily with thin and duplicate pages. Our indexing diagnostic guide walks through the rest.
Conclusion
Thin content is about missing value, not missing words. Use unique-word counts to build a shortlist, then expand, merge, noindex or delete each page on its merits and fix the templates that create them, starting with the pages a BugViso crawl flags as thin or near-duplicate.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness โ with fixes you can ship today.