Case study
400,000 thin URLs. One WooCommerce store. Here's what the damage looked like.
This is a real WooCommerce store. The store is not named and the URLs below have had their identifiers replaced — the structure, the counts and the dates are unchanged. This is what happens when filter, brand, rating and pagination combinations are left unchecked for years.
The store
What we were looking at.
- Store type
- Large WooCommerce catalogue with multiple product categories, plus brand, rating and colour filters and pagination on every archive.
- Store age
- Multi-year, established store carrying historic SEO signals worth protecting.
- Problem discovered
- A routine Search Console check revealed a catastrophic indexing situation nobody had been alerted to.
- Status
- Fix applied. Recovery in progress — Google is still working through the removals.
The damage
Four categories of waste, all generated automatically.
Excluded by noindex
28,000+ and processingWhat it shows
Filter combination URLs, paginated archives and add-to-cart parameter URLs — all auto-generated by WooCommerce, all crawled, now being excluded.
What it means
Google spent crawl budget visiting these repeatedly before the noindex fix was applied. 372,000 URLs have already been cleaned. 28,000 remain in the exclusion queue, being processed gradually.
These URLs were never meant to exist as indexable pages. WooCommerce created every one of them automatically.
The first two rows are the same product grid reachable under two different paths. The last row is a cart action Google crawled as a page.
Duplicate without user-selected canonical
20,000+What it shows
Archive pages serving identical product grids, with no canonical tag telling Google which URL is the real one.
What it means
Google could not determine which URL was authoritative — the thin content signal in its clearest form. Instead of ranking the real category page, it ranked nothing, or the wrong URL entirely.
Two of these differ only in the order of the filter values. To Google they are separate pages. To a shopper they are the same grid.
rating_filter=5,4 and rating_filter=4,5 return identical content at different addresses.
Crawled — currently not indexed
8,000+What it shows
URLs Google crawled, found thin or duplicate, and refused to index — but kept revisiting on every crawl cycle.
What it means
Pure crawl budget waste. Google kept checking pages it had already decided were not worth indexing, leaving less budget for real product and category pages.
Pages 2 and 3 of the same filtered archive, each crawled as an independent page.
Not found (404)
8,000+What it shows
URLs deleted from the store that Google is still reaching from internal links and old sitemaps.
What it means
Every 404 Google hits wastes a crawl request. Eight thousand dead URLs being revisited repeatedly means real pages wait longer to be discovered and indexed.
The 404s were caused by premature deletion of URLs, not by the thin URL problem. It is a separate issue WooScrub does not create and in most cases cannot reverse — once a URL is deleted without a redirect, its historical value is gone.
Deleted products and categories still being requested months later.
The numbers
Where it started, and where it is now.
Google does not clean up overnight. Recovery is gradual — but it is happening. The noindex directives are being processed, real pages are being re-evaluated, and crawl budget is being reclaimed.
What caused this
Nobody did anything wrong.
WooCommerce generates archive URLs automatically. Every time a shopper applies a filter, selects a brand, adjusts a rating or combines several at once, a new URL exists. For a store with:
The combinations multiply exponentially. Multiply that across the whole catalogue and the numbers reach six figures without a single deliberate act.
This is not a bug. It is how WooCommerce works by default.
Why the number gets so large
Brand filters are multi-select. A shopper does not pick one brand — they pick any combination of them, and every combination is its own URL.
Brand filters alone account for most of it. Nothing else has to go wrong for a catalogue this size to produce six figures of thin URLs.
The fix
What was done — and what it could not do.
What we implemented
- Every thin archive URL type identified — filters, ratings, brands, colours, parameters and pagination combinations
- Custom noindex and canonical implementation, specific to this store's theme and plugin setup
- Robots.txt updated afterwards, once the existing URLs had been dropped, to stop new ones being crawled at all
- Fix built to cover future URLs automatically — new filter combinations are handled the moment WooCommerce creates them
- Fix location chosen to suit the store's specific technical stack
What the fix did not do
- Restore the 8,000+ deleted URLs — those were already gone before we arrived
- Clean Search Console instantly — Google processes removals over weeks and months
- Guarantee ranking recovery on any particular timeline
Recovery status
Where each category stands today.
| Category | Before | Current status |
|---|---|---|
| Indexed thin URLs | 400,000+ | Being removed — 372,000 already gone |
| Excluded by noindex | 0 | 28,000+ and processing |
| Canonical conflicts | 20,000+ | Fix applied, Google processing |
| Crawled — not indexed | 8,000+ | Fix applied, crawl budget recovering |
| Not found (404) | 8,000+ | Ongoing — separate from the thin URL fix |
This case study will be updated as recovery progresses. Search Console data reflects the state at the time of writing — September 2026.
Your store may have the same problem — at any scale.
The free audit shows you exactly what Google sees: archive type counts, Search Console visibility data and a damage score. No credit card. Takes minutes.