Get in touch
  • WordPress
  • SEO
  • v1.0.0

Crawler Shield for WooCommerce — protects the catalogue from AI crawlers, meta-externalagent and filter crawling

A WooCommerce plugin that keeps crawlers from walking the endless filter combinations of the catalogue and taking the shop down. Bot groups with their own rules (search engines, Meta ad previews, AI crawlers, SEO tools, scripts), limits per bot and per visitor address, refusal before WooCommerce starts building the page, robots.txt, noindex for filter pages, llms.txt and a report. The base version is free; Pro checks that Googlebot and AI bots are real, adds data-centre networks, a browser check, stored copies for bots and nginx/Apache rules.

Download Free
Pro · one-year licence The more sites, the less each one costs
Compatibility
WordPress 6.2+ · WooCommerce 7.0+ (tested on WordPress 7.1 and WooCommerce 11.1) · PHP 7.4+ · nginx + PHP-FPM and nginx → Apache (as on HestiaCP) · Redis, APCu or a database table for counters
Licence
one site, annual
Trial
7 days, no card, one per site
Updated
September 2026
Or all 144 Pro modules with All Access — 4 990 ₴ / year

What Free has,
and what Pro adds

The free version works with no time limit. Pro adds the rest of the features in the table.

Feature Free Pro
Bot groups: search engines, link previews (Facebook and Instagram ads, Telegram, Viber), AI crawlers, SEO tools, scripts and headless browsers, other bots ✔ ✔
A rule for each group: allow, limit the rate, "pages yes, filter URLs no", block ✔ ✔
Filter URL detection: WooCommerce filters and sorting, price, rating, add-to-cart, HUSKY (WOOF), YITH, BeRocket, WBW, CatCode Smart Filter + your own parameters and paths ✔ ✔
Per-bot limit — all its addresses together, separately for filter URLs; 429 with Retry-After ✔ ✔
Limit of filter URLs a minute from one visitor address; logged-in users are never limited ✔ ✔
Refusal of URLs with an implausible number of filter values ✔ ✔
Refusal before WooCommerce and the theme build the page: a few bytes instead of a catalogue page ✔ ✔
Meta ad previews (facebookexternalhit) are only slowed down, never cut off ✔ ✔
Correct visitor address behind a local proxy and behind Cloudflare, or a header of your choice ✔ ✔
Counters in Redis / Memcached, APCu or one small table ✔ ✔
robots.txt: Disallow for filter parameters and Disallow: / for blocked groups ✔ ✔
noindex, follow on filter pages (one switch turns it off) ✔ ✔
/llms.txt — a short map of the shop for AI assistants ✔ ✔
Monitor mode: counts what would happen without refusing anything ✔ ✔
"Never limit" and "always refuse" lists for IPs, networks and User-Agents ✔ ✔
Request tester in the admin and in WP-CLI: what the shield would do and why ✔ ✔
Report history 7 days up to 90 days
Check of Googlebot, Bing, Apple, OpenAI, Anthropic, Perplexity against official address lists — fakes go to their own group — ✔
Data-centre networks with a browser User-Agent (Alibaba Cloud preset) — ✔
Per-/24 limit (/48 for IPv6) against scrapers spread over hundreds of addresses — ✔
Browser check before filter URLs: automatically during a wave or always, its own limit for every checked browser — ✔
Stored copies of pages for bots: repeated requests with new fbclid/utm tags served from disk — ✔
nginx and Apache rules generated from your settings — refusal before PHP starts — ✔
Detailed report: top networks, addresses and URLs for 24 hours — ✔
E-mail about a crawl wave — ✔
Popular products with prices in llms.txt — ✔

What it looks like

WooCommerce storefront with a colour filter — nothing changes for the shopper
The shopper sees an ordinary shop: colour filter, sorting, cart. A Storefront test shop with 2,000 products; they have no photos, hence the WooCommerce placeholder.
Crawler Shield "Rules" tab: shield mode and a rule for each bot group
The “Rules” tab (admin in Ukrainian): shield mode and a rule for each group — search engines allowed, link previews only limited, AI crawlers and SEO tools get “pages yes, filter URLs no”. The two bottom rows are Pro groups.
Crawler Shield 24-hour report: how many requests of each bot were served, slowed down and refused
The 24-hour report from a test run: meta-externalagent got 3,622 refusals on filter URLs while search bots got their pages. The “Stored copy” and “Browser check” columns come with Pro.

Key points

What the shopper sees

Nothing new. The shopper clicks filters, sorts, adds to the cart and checks out — in our test 25 quick filter clicks in a row, classic and block checkout went through without a single refusal and without check pages.

What the admin gets

Settings are in WooCommerce → Crawler Shield, five tabs: “Report”, “Rules”, “robots.txt and llms.txt”, “Test a request” and “Server rules” (Pro).

Pro: when bots pretend

Some crawlers call themselves Googlebot, some — Chrome.

The trial starts when you ask

A fresh install is the free version, nothing switches on by itself. The “Try Pro for 7 days” button → e-mail → the key right in the window and by e-mail.

One key — one site

Moving the store? Unbind the licence on the old domain and activate it on the new one yourself. If our server is unreachable, Pro keeps working for 14 more days. For studios — keys for 5 or 25 sites, or unlimited.

Technical requirements

  • WordPress 6.2+ and WooCommerce 7.0+. End-to-end run on WordPress 7.1 and WooCommerce 11.1.
  • PHP 7.4+. End-to-end run on PHP 8.3 (nginx + PHP-FPM) and PHP 8.2 (Apache), syntax also on 7.4.
  • Server. nginx + PHP-FPM, Apache, nginx in front of Apache (HestiaCP, VestaCP). Behind Cloudflare the address is detected automatically; behind another CDN choose the header in “Server” → “Visitor address”.
  • Counters. Redis or Memcached as object cache, APCu or one database table — the plugin takes the fastest one available. Redis and APCu are tested, Memcached is not.
  • robots.txt. Lines are added to the virtual robots.txt of WordPress. If there is a physical robots.txt file in the site root, copy the lines from the tab by hand.
  • Filter plugins. Parameter names of HUSKY, YITH, BeRocket and WBW were checked against their code; we have not tested it with those plugins installed or with page caching (WP Super Cache, LiteSpeed) — start in “Monitor” mode.

Version history

v1.0.0 Current September 2026

First public release: bot groups with “allow / limit / pages yes, filters no / block” rules, filter URL detection for WooCommerce and popular filter plugins, per-bot and per-address limits,…

Frequently bought with Crawler Shield for WooCommerce

4 modules in one order — 40% cheaper than separately

“Store starter” bundle Crawler Shield Nova Poshta Premium LiqPay Telegram notifications and Viber/SMS for customers Need more — all 144 modules in All Access for 4 990 ₴.
3 590 ₴ / year instead of 5 960 UAH bought separately
Buy the bundle →

Full module description

Crawler Shield for WooCommerce is a plugin that keeps crawlers from taking the shop down on filter pages. A catalogue with filters has millions of URLs: every combination of colour, size, brand, price and sorting is a page of its own, and each one is a heavy database query. People open a handful of them. Crawlers — meta-externalagent, GPTBot, ClaudeBot, Bytespider, SEO tools and scrapers — open all of them, and a shop on ordinary hosting runs out of PHP workers: pages take tens of seconds, and a host may suspend the site for “abnormal load”.

The plugin decides who gets what before WooCommerce and the theme start building the page and answers a refusal with a few bytes instead of a catalogue page. Search engines and your ad previews go through, AI crawlers and SEO tools see pages but not filter combinations, scripts are rate limited. There is also an OpenCart version.

What the shopper sees

Nothing new. The shopper clicks filters, sorts, adds to the cart and checks out — in our test 25 quick filter clicks in a row, classic and block checkout went through without a single refusal and without check pages. The shield does not count ordinary pages for ordinary visitors at all and never limits logged-in shoppers.

There is only one limit for people: 40 filter URLs a minute from one address — more than a person manages to click. If many shoppers share one address (a mobile carrier, an office), raise it or add the address to “Never limit these addresses”. During a wave Pro also shows unknown visitors a browser check page for a second — a real browser passes by itself and then has its own limit.

What the shop gets

  • A catalogue that bots do not bring down. In a local load test (php-fpm with 5 workers, 2,000 products) the shop refused 600 meta-externalagent requests to filter URLs at 477 requests a second versus 63 pages a second without the shield, while a shopper got a filter page in 0.14–0.19 s.
  • Rules for each bot group. Search engines — allow, link previews — limit, AI crawlers and SEO tools — “pages yes, filter URLs no”, scripts and other bots — limit. Any group can be switched to “Block”.
  • Limits that fit a shop. A bot is counted as a bot, not as an address: facebookexternalhit from a hundred addresses is one client, 60 requests a minute, 10 of them filter URLs. After that — 429 with Retry-After, not a broken page.
  • Meta ads keep working. Link previews are never cut off by default — Meta checks ad landing pages with them.
  • Less junk in the index. Disallow lines for filter parameters and noindex, follow on filter pages — categories and products stay open.

What the admin gets

Settings are in WooCommerce → Crawler Shield, five tabs: “Report”, “Rules”, “robots.txt and llms.txt”, “Test a request” and “Server rules” (Pro). The “Rules” tab has the shield mode, a rule for each group and the limits. The defaults suit most shops; if unsure, choose “Monitor” for a day: the shield only counts what it would do and refuses nothing.

Further down the same tab are the recognised filter parameters and fields for your own: parameters (color, pa_*) and paths, if a filter builds pretty URLs like /filter/. Without them pretty-URL filters are ordinary pages to the shield. There are also “never limit” and “always refuse” lists for addresses, networks and User-Agents, and the “Server” block: where to take the visitor address from (automatically, CF-Connecting-IP, X-Forwarded-For or X-Real-IP) and where to keep the counters.

The “Report” shows every bot for 24 hours and by day: how many requests were served, slowed down (429) and refused (403). This is where you see who really loads the shop.

The “robots.txt and llms.txt” tab has switches for the filter Disallow lines, closing the site to “Block” groups and noindex, follow on filter pages. noindex and Disallow are on by default: if you deliberately rank filter pages with parameters in the URL, untick both here. /llms.txt is set up on the same tab: the “About the shop” text and extra pages.

“Test a request” takes a URL, a User-Agent and an address and shows what the shield would do and why, without counting anything. The same from the console: wp catcode-crsh test "/shop/?filter_color=red" --ua="meta-externalagent/1.1".

Pro: when bots pretend

Some crawlers call themselves Googlebot, some — Chrome. Once a day Pro downloads the official address lists of Google, Bing, Apple, OpenAI, Anthropic and Perplexity: the real Googlebot from a Google network goes through, a “Googlebot” from someone else’s address lands in the “Fake bots” group and is refused. Data-centre networks (an Alibaba Cloud preset — in September 2026 scrapers with a browser User-Agent walked the filters of our demo shop from there) get “pages yes, filter URLs no”. The per-/24 limit catches a scraper spread over hundreds of addresses.

When unknown visitors together open more than 300 filter URLs a minute, Pro turns on a browser check for 10 minutes: a real browser passes in a second, a script does not. Ad previews with new fbclid and utm tags are served from stored copies instead of being built every time (a copy lives up to 6 hours — configurable; bots may see a price that old). The “Server rules” tab generates blocks for nginx and Apache so the heaviest refusals are done by the web server before PHP.

How it works — step by step

  1. A request reaches the site. The shield runs on the earliest WordPress hook, before WooCommerce and the theme initialise.
  2. The client is identified: a bot group by User-Agent (in Pro also an address check against the official lists) or a visitor. The address is taken correctly behind a local proxy and Cloudflare.
  3. The page type is identified: a filter URL (parameters of WooCommerce and popular filter plugins, your parameters and paths) or an ordinary page.
  4. The group rule and the limits apply. The counter lives in the object cache, APCu or one table, one query.
  5. A refusal is 403 or 429 with Retry-After, no-store and noindex, a few bytes. Everything else is the page as usual.
  6. The decision goes into the report: who, how many, what was served, what was refused.

Under the hood

  • An early decision. The shield runs on plugins_loaded at priority -10000: the WordPress core is loaded, but WooCommerce, the theme and catalogue queries are not yet. A refusal is cheap but not free; the full effect comes from the Pro server rules.
  • Counters. Object cache → APCu → a table with one INSERT … ON DUPLICATE KEY UPDATE per request. Statistics are one more small upsert.
  • Visitor address. CF-Connecting-IP is accepted only from Cloudflare networks, behind a local proxy — X-Real-IP or the last X-Forwarded-For. A forged header from the internet does not change the address.
  • robots.txt uses the format from Google’s faceted navigation documentation (Disallow: /*?*parameter=) and is appended after the WordPress and WooCommerce rules. robots.txt always stays readable for blocked bots — so they can read Disallow.
  • Pro address lists are the official JSON files of Google, Bing, Apple, OpenAI, Anthropic and Perplexity, refreshed once a day. The Pro journal of refusals with IP addresses is kept for 7 days.
  • Uninstall removes the tables, settings, cron jobs and the stored-copies folder — robots.txt and /llms.txt go back to the defaults.

How to install

  1. Download the free version catcode-crawler-shield-for-woocommerce-free-1.0.0.zip with the button at the top of the page (or the Pro archive from your account after purchase): Plugins → Add New → Upload Plugin → Install → Activate.
  2. WooCommerce → Crawler Shield → “Rules”. Unsure about the rules — choose “Monitor” for a day and look at the “Report”.
  3. Behind Cloudflare — do nothing. Behind another proxy or CDN — choose the header in “Server” → “Visitor address”.
  4. Filters with pretty URLs — add their path to “Filter paths”. Ranking filter pages in search — switch off noindex and Disallow on the “robots.txt and llms.txt” tab.
  5. For Pro: paste the licence key from the e-mail after purchase or click “Try Pro for 7 days” in the “Pro licence” section.

Questions
about the module

Didn't find the answer? Message us on Telegram and we'll reply within a business day.

@catcode_support Setup, compatibility, activation
Will it hurt my Google rankings?

Search engines are allowed by default. Filter parameters are closed in robots.txt with lines like "Disallow: /*?*min_price=" — the way Google's documentation on faceted navigation recommends — and filter pages get noindex, follow. Categories and products stay open. If you deliberately rank filter pages with parameters in the URL, switch both off on the "robots.txt and llms.txt" tab.

Will my Facebook and Instagram ads keep working?

Yes. Link previews (facebookexternalhit) are never blocked by default — only limited on filter URLs: 10 a minute, then 429 with Retry-After. The ad landing page, including with fbclid and utm, is served as usual. We have not watched Meta's ad account react to the limit live — if your ads point to filter URLs, raise the limit or switch the group to "Allow".

My site is behind Cloudflare. What do I set?

Nothing: the plugin recognises requests from Cloudflare networks itself and takes the address from CF-Connecting-IP. Behind another CDN or proxy it does not know, choose the header in "Rules" → "Server" → "Visitor address", otherwise every visitor gets the proxy address and they share one limit. We imitated Cloudflare rather than connecting the real one.

Can real shoppers hit the limit?

They can if many people share one address: a mobile carrier (CGNAT) or an office network. The limit is 40 filter URLs a minute from one address; ordinary pages and logged-in shoppers are not counted. AJAX filters (HUSKY, YITH, WBW) may make several requests per click — then raise the limit. In Pro the browser check gives every browser its own limit.

Does it replace a firewall?

No, and it does not promise 100% protection. It solves one problem general firewalls do not: bots walking the endless filter combinations of a shop catalogue. A scraper calling itself Chrome is an ordinary visitor with a per-address limit in Free; data-centre networks, the per-/24 limit and the browser check come with Pro.

My site has a physical robots.txt file.

Then WordPress does not build robots.txt and the plugin cannot add its lines. The "robots.txt and llms.txt" tab shows them for copying into your file.

Does the plugin send data anywhere?

Free — nowhere, everything is counted on your server. Pro downloads the official address lists of Google, Bing, Apple, OpenAI, Anthropic and Perplexity once a day and checks the licence at catcode.com.ua. The Pro journal with IP addresses of refused requests is kept on your server for 7 days — under GDPR that is personal data, mention it in your privacy policy.

Does it work with page caching?

Yes. Pages served from a full-page cache never reach WordPress — they cost nothing anyway and the plugin does not see them. Everything that reaches WordPress goes through the shield. Pro stored copies for bots are separate from your cache. We have not tested it together with WP Super Cache or LiteSpeed Cache.

How much does it cost?

The base version is free. Pro is 990 UAH per year, and a longer term is cheaper: 2 years — 1,880 UAH, 5 years — 3,710 UAH. No automatic charges. A purchased licence keeps the Pro features for good; the term covers updates and support.

Not quite what you are looking for?

We build custom modules for WordPress, WooCommerce, OpenCart and Shopify. Tell us about the task and we'll prepare an estimate.

Order a custom module