Filter combination URLs
Every active filter selection creates a unique URL serving the same product grid, filtered differently. The filter page fix in detail.
Which of the dozens of combinations is the page that should rank.
WooCommerce thin content fix
Thin content in WooCommerce is almost never about description length. It is about archive pages — filter combinations, tag archives, brand pages and parameter URLs — that Google encounters with no signal telling it what they are or which page is the real one. The result is not a penalty. It is confusion. And confusion has a fix.
Free audit. Real Search Console data. Your exact thin content picture in minutes.
/product-category/shoes/No canonical/product-category/shoes/?filter_store=nikeNo canonical/product-category/shoes/?filter_store=nike,adidasNo canonical/product-category/shoes/?filter_store=nike,adidas&rating_filter=4No canonicalFour addresses. One page. Same grid, same template, same headings — and nothing on any of them saying which is the original.
The myth and the reality
When store owners see "thin content" mentioned in SEO forums, the instinct is to write longer product descriptions. Add more words. Hit 300. Hit 500. Hire a copywriter. That is the wrong fix for the wrong problem, and it costs time and money while the real issue continues untouched.
Thin content as Google understands it is not primarily about word count. It is about pages that exist with no unique value — pages serving the same content as other pages on your site, with nothing differentiating them or telling Google which one matters.
Your products almost certainly have unique names, attributes, images and real differentiating content. Your archive pages are the problem.
Google calls this thin content — not because your words are short, but because it cannot find a reason to treat these pages as meaningfully different from each other or from your real category pages.
120 words → 500 words. The page was never the problem, so nothing moves.
/shoes/?filter_store=nike
/shoes/?rating_filter=4
/shoes/page/2/
/product-tag/sale/
Untouched by any amount of copywriting.
Google isn't penalising you
Your real pages don't rank as well as they should — not because Google punished them, but because it can't find them clearly among the noise. That's fixable. A penalty is a much harder conversation.
Two GSC warnings, one root cause
Google found several URLs serving similar content and no canonical saying which is authoritative. It had to guess — and it often guesses wrong, indexing a filter URL instead of your real category page.
Google crawled these, judged them not worth indexing, and kept revisiting anyway. It recognised they had no unique value but never stopped spending requests on them.
Archive URLs that exist as crawlable pages with no canonical signal.
Filter combinations, tag archives, brand pages and paginated archives, all reachable and all silent about their relationship to the real page. One implementation — noindex, canonical, then a robots.txt block — clears both categories.Filters are for your customers
This is the part that surprises most store owners: nothing about the shopping experience changes. We are not removing your brand filter, your rating filter, your colour filter, or any navigation your customers rely on. The same URL behaves differently depending on who is looking at it.
/product-category/shoes/?filter_store=nike,adidas&rating_filter=4
/product-category/shoes/?filter_store=nike,adidas&rating_filter=4
The order matters. The robots.txt block goes on last. Block the URLs first and Google can never recrawl them to see the noindex, so they sit in the index instead of leaving it.
What actually causes it
Every active filter selection creates a unique URL serving the same product grid, filtered differently. The filter page fix in detail.
Which of the dozens of combinations is the page that should rank.
Every product tag gets a browsable archive. Two hundred tags means two hundred grids differing only by label.
Whether a tag archive is a real destination or a label on an existing one.
Brand filter plugins create an archive per brand, then multiply it by pagination and other filters.
Whether the brand archive or the category page owns those products.
Page 2, page 3 and page 4 of any archive serve the same template with a different slice of products.
Which page in the series is the canonical entry point.
?add-to-cart=193 on any URL creates another crawlable copy of that page. Google's crawl team filed a bug over this.
That the parameter is a cart action rather than a different page.
Some themes serve the same grid at two paths — /product-category/shoes/ and /collection/shoes/.
Which of the two paths is the original.
These same six types are also what drains your crawl budget — the cost there is measured in wasted requests rather than lost clarity. See the crawl budget side
Why word count padding doesn't fix this
The signal Google is responding to is not on your category page. It is on the filter combination URL — the page neither you nor Google intended to rank, which exists as a crawlable duplicate of your category page.
When you add content to /product-category/shoes/ you improve the real page. But the duplicate still exists, still has no canonical pointing home, and Google still finds both. The thin content signal does not go away — it just coexists with a slightly better real page.
Improve your category page copy and fix the archive URLs. Both are worth doing. But if your rankings are suppressed by archive URL thin content, better words on the real page alone will not move the needle.
/product-category/shoes/+300 wordsThe real page is better. Worth doing.
/shoes/?filter_store=nikeunchangedStill a crawlable duplicate. Still no canonical. Still generating the signal.
How WooScrub finds it
WooScrub runs a diagnostic crawl of your store and cross-references it against your Google Search Console data, so the archive types you have are mapped directly onto the Search Console categories they are landing in.
After the diagnosis, the fix is implemented manually — reviewed against your theme, filter plugin, permalink structure and server configuration. The solution covers future archive URLs automatically.
Not a generic thin content checker. It does not measure word count, readability or content quality. It finds the structural archive URLs generating thin content signals on your specific store — the ones you have never seen, because they are created in the background while the store runs normally.
Every figure comes from your store and your Search Console account — counted by type, not estimated from a benchmark.
Before changing anything
The audit scans your store and cross-references your Search Console data. You see your exact thin URL count, which categories they are appearing in, and your damage score — all before spending anything.
Free. No credit card. Your actual thin content picture, not a generic estimate.