WooCommerce crawl budget

WooCommerce is wasting Google's crawl budget — and your real pages are paying the price

Google gives every website a crawl budget: a limited number of pages it will visit in a given period. WooCommerce stores burn through it on filter combinations, tag archives, brand pages and parameter URLs that should never have been crawlable in the first place. While Google is busy crawling junk, your real product and category pages wait.

Free audit. Connect Search Console. See your numbers in minutes.

Your WooCommerce store Googlebot · limited budget
What the store puts in front of Googlebot
?filter_store=brand-1 ?filter_color=red&size=m /add-to-cart/sale/ ?add-to-cart=27 /shop/page/7/ ?rating_filter=5 ?filter_price=100-200 + thousands more
What the crawl budget actually gets spent on
Thin archive URLs — crawled first Your real pages — still waiting

Even Google said so

This isn't a theory. Google's own crawl team filed a bug against WooCommerce.

In a report covered by Search Engine Journal, Google's crawl team identified a specific WooCommerce behaviour as a crawl waste problem: add-to-cart URLs being generated and crawled at scale. These are transactional parameters with zero SEO value — but WooCommerce creates them for every product, and Google follows them.

The bug was filed against WooCommerce core. Not a plugin. Not a theme. The platform itself.

This matters because it confirms something store owners have been discovering in their Search Console data for years: WooCommerce's default URL architecture creates a crawl budget problem by design. It generates URLs Google cannot ignore — across every product, every category, every filter and every archive. Automatically, indefinitely, and at scale.

If Google's crawl team called it out, your store almost certainly has it.

Search Engine Journal

Google files bug against WooCommerce for add-to-cart URL crawl waste

"Google's crawl team identified add-to-cart URLs as a crawl waste problem."
Read the full article

https://your-store.com/product/?add-to-cart=193

Filed against WooCommerce core — not a plugin, not a theme.

What crawl budget actually means

Google doesn't crawl everything. It makes choices — and WooCommerce is making bad ones for you.

Crawl budget is the number of URLs Googlebot will crawl on your site within a given timeframe. For most small and medium WooCommerce stores that budget is limited. Google decides how to spend it on two factors: crawl rate, meaning how fast your server responds, and crawl demand, meaning how important Google thinks your pages are.

When your store generates thousands of low-value URLs, Google spends its budget crawling them — repeatedly — instead of discovering and indexing your real pages. The same URLs are also what generates your thin content signal. The consequences are direct and measurable:

  • New products take weeks to appear — Googlebot spent its last visit on a filter URL instead of your new product page.

  • Category pages get deranked or deindexed — Google found dozens of near-identical archive URLs and could not determine which was the real page.

  • Historic ranking signals dilute — The same content exists on multiple URLs, splitting whatever authority your real page had earned.

  • Crawl stats look active while the index grows slowly — Search Console shows thousands of crawl requests while your real page count barely moves.

This is not a performance problem. It is not a content quality problem. It is a URL architecture problem — and WooCommerce creates it by default.

How Google spends your crawl budget

Crawl budget

Wasted on thin URLs

70%

Filters, tags, brands, parameters, pagination

Your real pages

30%

Products, categories, important content

Google Search Console — crawl stats (example)

60K40K20K0 Aug 1Aug 8Aug 15Aug 22Aug 29
Thin URLs 42,310 requests Real pages 12,480 requests

High crawl activity doesn't mean good news. It usually means Google is crawling the wrong pages.

What is burning your crawl budget

Six WooCommerce URL types that consume crawl budget without returning any SEO value.

1. Filter combination URLs

Every time a shopper selects a brand filter, a rating filter, a colour filter, or any combination of them, WooCommerce creates a new URL. The combinations reach into the tens of thousands.

/product-category/category-a/?rating_filter=5&filter_store=brand-1,brand-2

2. Paginated archive duplicates

Every category, tag, brand and filter archive has a paginated version — /page/2/, /page/3/ and onward. Each is a near-identical copy of the base archive.

/shop/page/3/?filter_store=brand-2,brand-5,brand-7

3. Add-to-cart parameter URLs

When a shopper clicks Add to Cart without JavaScript, WooCommerce appends ?add-to-cart={product-id} to the URL. These are crawlable, indexable URLs with zero SEO value.

/shop/?filter_store=brand-4&add-to-cart=3003

4. Tag archive pages

WooCommerce creates a browsable archive for every product tag. If your store has 200 tags, that is 200 archive pages — each with no unique content and no reason to outrank your real categories.

/product-tag/sale/

5. Brand archive pages

Each brand gets its own archive URL. Combined with pagination and other filter combinations, brand archives multiply rapidly — every combination another address to crawl.

/brand/brand-1/page/2/

6. Collection and duplicate paths

Some themes and plugins create duplicate path structures for the same content. Google crawls both. Neither ranks well, because they compete with each other.

/collection/category-a/ = /product-category/category-a/

What you see in Google Search Console

The symptoms are visible in GSC. Here's how to read them.

If your store has a crawl budget problem, Search Console will show it — but never as a single clear warning. The signals are spread across several reports, and each one on its own looks like a minor housekeeping note.

Read together, they describe a store where Googlebot is working hard and achieving very little.

Where the problem shows up
Coverage — "Crawled, currently not indexed"high

Google crawled these, decided they were not worth indexing, and kept visiting anyway. Budget spent on pages that return nothing.

Coverage — "Duplicate without user-selected canonical"high

Google found multiple URLs serving similar content and no canonical telling it which is real. It chose for you — often wrong.

Coverage — "Excluded by noindex tag"high

URLs excluded by a partial fix already in place. A large number here means the problem existed at scale before that fix.

Crawl stats — requests vs indexed pagesratio

Thousands of daily requests against a small indexed count for your catalogue size means Googlebot is working outside the index.

Pages — real pages missingabsence

Products and categories published weeks ago that still have not appeared. Often crawl budget exhaustion, not content quality.

None of these is labelled "crawl budget problem". That is why it goes unnoticed for years.

Why the standard advice fails

"Just noindex your filter pages" is the right idea applied the wrong way.

The standard recommendation for WooCommerce crawl budget problems is: noindex your filter pages. Apply it in Yoast, RankMath or a dedicated plugin. Done.

The problem is that the advice gets applied blindly — without first knowing which specific filter combinations exist on your store, how many Google has already indexed, whether any of them are currently ranking, and whether your theme and plugin setup means a blanket rule will even fire correctly. Why noindexing filter pages backfires covers that in full.

There is a documented case in the SEO community of a store owner who applied blanket noindex to all filter pages and watched traffic drop immediately after. Not because noindex was wrong — but because some filter pages were already ranking, and the rule removed them from the index with nothing to replace them.

The order matters
The fix is correct. The diagnosis has to come first.

Applying robots.txt blocks or noindex rules without a site-specific diagnosis is the fastest way to remove pages that were working — while the ones causing the real problem carry on untouched.

What has to be known before any rule is written

  • Which filter combinations your store actually generates
  • How many of them Google has already indexed
  • Which of them are currently ranking and earning traffic
  • Whether your theme and plugins will let the rule fire at all

How WooScrub diagnoses it

See exactly how much crawl budget WooCommerce is wasting on your store.

WooScrub runs a diagnostic crawl of your store and cross-references it against your Google Search Console data. The result is a store-specific damage report, not a generic site audit.

After the diagnosis, if you want the fix implemented, we do it manually — specific to your theme and plugin setup, with a solution that covers future URLs automatically.

See how the fix works

What the report shows

  • Every archive URL type your store generates — filter combinations, paginated archives, tag archives, brand pages, add-to-cart parameters
  • A count for each type, not just a total
  • A Search Console cross-reference: indexed, crawled but not indexed, excluded
  • A crawl budget impact estimate — how much is going to thin archive URLs rather than real pages
  • A damage score with a severity band — low, medium, high or critical

Find out

Exactly how much crawl budget WooCommerce is wasting on your store.

The audit is free. It connects to your Search Console account and shows you your real numbers — not a generic estimate.

Free. No credit card. Your crawl budget breakdown in minutes.

Frequently asked questions

How do I check my crawl budget in Google Search Console?

There's no report called "crawl budget". Open Settings → Crawl stats for total requests over the last 90 days, broken down by response, file type and purpose, then compare that against the Pages report. What you're looking for is a mismatch: thousands of daily requests against an indexed count that barely moves, or a large "Crawled — currently not indexed" group. That gap is crawl budget going somewhere it isn't earning anything.

How many filter URLs is too many for Google?

There's no threshold Google publishes, and no number that's safe in isolation — what matters is the ratio to your real pages. A store with 200 products and 3,000 filter URLs is in worse shape than one with 50,000 products and 20,000 of them, because the second store has far higher crawl demand to begin with. The practical test is whether your real pages get discovered promptly. If new products take weeks to appear, your archive URLs are already crowding them out, whatever the absolute count.

Does crawl budget matter for small WooCommerce stores?

It matters more, not less. Google allocates crawl budget partly on crawl demand — how important it judges your site to be — so a small store starts with a smaller allowance. Filter combinations don't scale down with catalogue size: 15 brands and 5 rating levels generate the same combinations whether you sell 80 products or 8,000. A small store can easily have more thin archive URLs than real pages, which means most of an already-modest budget goes to pages that can't sell anything.

How long does it take to recover crawl budget after fixing thin URLs?

Faster than rankings recover. Once the robots.txt block is in place, Google stops requesting those URLs within days to weeks — visible in Crawl stats as total requests fall and the mix shifts toward real pages. Deindexing the URLs already in the index is the slow part, 3–6 months, because Google has to recrawl each one to see the noindex before dropping it. Ranking recovery follows that, and continues after it.