WooCommerce thin content fix

Your thin content problem isn't your product descriptions. Google isn't penalising you — it's confused.

Thin content in WooCommerce is almost never about description length. It is about archive pages — filter combinations, tag archives, brand pages and parameter URLs — that Google encounters with no signal telling it what they are or which page is the real one. The result is not a penalty. It is confusion. And confusion has a fix.

Free audit. Real Search Console data. Your exact thin content picture in minutes.

What Google sees on one category4 URLs
/product-category/shoes/No canonical
/product-category/shoes/?filter_store=nikeNo canonical
/product-category/shoes/?filter_store=nike,adidasNo canonical
/product-category/shoes/?filter_store=nike,adidas&rating_filter=4No canonical

Four addresses. One page. Same grid, same template, same headings — and nothing on any of them saying which is the original.

The myth and the reality

Stop padding your product descriptions. That's not what's causing this.

When store owners see "thin content" mentioned in SEO forums, the instinct is to write longer product descriptions. Add more words. Hit 300. Hit 500. Hire a copywriter. That is the wrong fix for the wrong problem, and it costs time and money while the real issue continues untouched.

Thin content as Google understands it is not primarily about word count. It is about pages that exist with no unique value — pages serving the same content as other pages on your site, with nothing differentiating them or telling Google which one matters.

Your products almost certainly have unique names, attributes, images and real differentiating content. Your archive pages are the problem.

Google calls this thin content — not because your words are short, but because it cannot find a reason to treat these pages as meaningfully different from each other or from your real category pages.

What you're told to fixProduct description length

120 words → 500 words. The page was never the problem, so nothing moves.

What's generating the signalArchive URLs
/shoes/?filter_store=nike /shoes/?rating_filter=4 /shoes/page/2/ /product-tag/sale/

Untouched by any amount of copywriting.

Google isn't penalising you

What Google actually does when it finds thin archive URLs — and why it isn't a penalty.

A manual action

Deliberate manipulation, reviewed by a human

  • Hundreds of thousands of templated pages built to rank — "plumber in Chicago", "plumber in Houston" — with a form and swapped keywords
  • Created on purpose, to game rankings
  • A reviewer at Google identifies it and applies a penalty
  • Site-wide or page-level, and it takes a reconsideration request to lift
What's happening to you

A structural byproduct, handled algorithmically

  • Nobody created 46,000 filter URLs to game anything — WooCommerce created them when shoppers used the filters
  • No manipulative intent, so no manual action
  • Google marks them "duplicate without user-selected canonical", or crawls them without indexing
  • Ranking signals spread thin across dozens of near-identical URLs

Your real pages don't rank as well as they should — not because Google punished them, but because it can't find them clearly among the noise. That's fixable. A penalty is a much harder conversation.

Two GSC warnings, one root cause

Two labels in Search Console. One thing generating both of them.

Coverage report Duplicate without user-selected canonical

Google found several URLs serving similar content and no canonical saying which is authoritative. It had to guess — and it often guesses wrong, indexing a filter URL instead of your real category page.

Coverage report Crawled — currently not indexed

Google crawled these, judged them not worth indexing, and kept revisiting anyway. It recognised they had no unique value but never stopped spending requests on them.

The one cause

Archive URLs that exist as crawlable pages with no canonical signal.

Filter combinations, tag archives, brand pages and paginated archives, all reachable and all silent about their relationship to the real page. One implementation — noindex, canonical, then a robots.txt block — clears both categories.

Filters are for your customers

Your filter pages stay exactly where they are. Customers need them. Google doesn't.

This is the part that surprises most store owners: nothing about the shopping experience changes. We are not removing your brand filter, your rating filter, your colour filter, or any navigation your customers rely on. The same URL behaves differently depending on who is looking at it.

What your customer sees

/product-category/shoes/?filter_store=nike,adidas&rating_filter=4

  • The filtered product grid appears exactly as it does today
  • Filters, sorting and pagination all work normally
  • The URL is shareable and bookmarkable
  • Nothing about the store changes
What Google sees

/product-category/shoes/?filter_store=nike,adidas&rating_filter=4

  • noindex — read it, do not add this to the index
  • canonical → /product-category/shoes/ — the real page is over here
  • Signals and attention reassigned to the page that deserves them
  • Once it has dropped, robots.txt stops the crawling entirely

The order matters. The robots.txt block goes on last. Block the URLs first and Google can never recrawl them to see the noindex, so they sit in the index instead of leaving it.

What actually causes it

Six archive types, and what each one stops Google being able to work out.

Filter combination URLs

Every active filter selection creates a unique URL serving the same product grid, filtered differently. The filter page fix in detail.

What Google can't tell

Which of the dozens of combinations is the page that should rank.

Tag archive pages

Every product tag gets a browsable archive. Two hundred tags means two hundred grids differing only by label.

What Google can't tell

Whether a tag archive is a real destination or a label on an existing one.

Brand archive pages

Brand filter plugins create an archive per brand, then multiply it by pagination and other filters.

What Google can't tell

Whether the brand archive or the category page owns those products.

Paginated archive duplicates

Page 2, page 3 and page 4 of any archive serve the same template with a different slice of products.

What Google can't tell

Which page in the series is the canonical entry point.

Add-to-cart parameter URLs

?add-to-cart=193 on any URL creates another crawlable copy of that page. Google's crawl team filed a bug over this.

What Google can't tell

That the parameter is a cart action rather than a different page.

Collection and duplicate paths

Some themes serve the same grid at two paths — /product-category/shoes/ and /collection/shoes/.

What Google can't tell

Which of the two paths is the original.

These same six types are also what drains your crawl budget — the cost there is measured in wasted requests rather than lost clarity. See the crawl budget side

Why word count padding doesn't fix this

Adding 300 words to your category page will not resolve thin content caused by archive URLs.

The signal Google is responding to is not on your category page. It is on the filter combination URL — the page neither you nor Google intended to rank, which exists as a crawlable duplicate of your category page.

When you add content to /product-category/shoes/ you improve the real page. But the duplicate still exists, still has no canonical pointing home, and Google still finds both. The thin content signal does not go away — it just coexists with a slightly better real page.

Improve your category page copy and fix the archive URLs. Both are worth doing. But if your rankings are suppressed by archive URL thin content, better words on the real page alone will not move the needle.

/product-category/shoes/+300 words

The real page is better. Worth doing.

/shoes/?filter_store=nikeunchanged

Still a crawlable duplicate. Still no canonical. Still generating the signal.

How WooScrub finds it

See exactly which archive URLs are generating your thin content signal.

WooScrub runs a diagnostic crawl of your store and cross-references it against your Google Search Console data, so the archive types you have are mapped directly onto the Search Console categories they are landing in.

After the diagnosis, the fix is implemented manually — reviewed against your theme, filter plugin, permalink structure and server configuration. The solution covers future archive URLs automatically.

See how the fix works

What this is not

Not a generic thin content checker. It does not measure word count, readability or content quality. It finds the structural archive URLs generating thin content signals on your specific store — the ones you have never seen, because they are created in the background while the store runs normally.

Archive types mapped to Search Console
Filter combinationsDuplicate without user-selected canonical
12,480
Paginated archivesCrawled — currently not indexed
4,120
Tag archivesDuplicate without user-selected canonical
1,904
Brand archivesCrawled — currently not indexed
1,236
Add-to-cart parametersCrawled — currently not indexed
842
Duplicate pathsDuplicate without user-selected canonical
318

Every figure comes from your store and your Search Console account — counted by type, not estimated from a benchmark.

Before changing anything

Find out which archive pages are causing your thin content signal.

The audit scans your store and cross-references your Search Console data. You see your exact thin URL count, which categories they are appearing in, and your damage score — all before spending anything.

Free. No credit card. Your actual thin content picture, not a generic estimate.