=== Crawler Shield for WooCommerce ===
Contributors: catcodestudio
Tags: woocommerce, bots, crawler, rate limit, ai crawlers
Requires at least: 6.2
Tested up to: 7.1
Requires PHP: 7.4
Stable tag: 1.0.0
License: GPLv2 or later
License URI: https://www.gnu.org/licenses/gpl-2.0.html

Stops Meta, AI and SEO crawlers from overloading the shop on endless catalogue filter combinations. Rate limits, robots.txt, llms.txt, a report.

== Description ==

A catalogue with filters has millions of URLs: every combination of colour, size, brand, price and sorting is a page of its own, and each one is a heavy database query. People open a handful of them. Crawlers open all of them — meta-externalagent, GPTBot, Bytespider, SEO tools and scrapers — and a shop on ordinary hosting runs out of PHP workers: pages slow down to many seconds, and a host may suspend the site for "abnormal load".

Crawler Shield decides who gets what before WooCommerce and the theme do any work, and answers a refused request with a few bytes instead of a catalogue page.

**Bot groups with their own rules**

* **Search engines** (Google, Bing, Apple…) — allowed by default.
* **Link previews** (facebookexternalhit, Telegram, Viber, WhatsApp…) — only slowed down, never cut off: Meta checks ad landing pages with them.
* **AI crawlers** (meta-externalagent, GPTBot, ClaudeBot, PerplexityBot, Bytespider, Amazonbot, CCBot…) — pages yes, filter combinations no.
* **SEO tools** (Ahrefs, Semrush, Majestic, DataForSEO…) — pages yes, filter combinations no.
* **Scripts and headless browsers** (curl, python-requests, HeadlessChrome…) — rate limited.
* **Other bots** — anything that calls itself a bot, crawler or spider.

For each group: allow, limit the rate, "pages yes, filter URLs no", or block.

**Filter URLs recognised out of the box**

WooCommerce layered navigation (filter_…, query_type_…, min_price, max_price, rating_filter), sorting, add-to-cart links, HUSKY (WOOF), YITH Ajax Product Filter, BeRocket, WBW Product Filter and CatCode Smart Filter. Add your own parameters or pretty-URL paths in one line each.

**Limits that fit a shop**

* A bot's budget is counted for the bot, not per address: facebookexternalhit from a hundred addresses is still one client.
* Visitors are counted only on filter URLs and only per address; logged-in users are never limited.
* A URL with more filter values than a person would ever tick is refused.
* Refused requests get 429 with Retry-After or 403, no-store and noindex.
* Counters live in the fastest store the server has: Redis/Memcached object cache, APCu, or one small table updated with a single query.

**robots.txt, noindex and llms.txt**

* Disallow lines for filter parameters and for the bots you block, added to the robots.txt WordPress builds.
* noindex, follow on filter pages: combinations drop out of search results.
* /llms.txt — a short Markdown map of the shop for AI assistants: categories, shop pages, sitemap.

**Report**

Requests of every bot over the last 24 hours and by day: served, slowed down, refused. Monitor mode counts what the shield would do without refusing anything — switch the rules on once the report looks right.

**Test a request**

Paste a URL, a User-Agent and an address and see what the shield would do and why. The same from WP-CLI: `wp catcode-crsh test "/shop/?filter_color=red" --ua="meta-externalagent/1.1"`.

**Where the data goes**

Nowhere. Everything happens on your server; the plugin makes no external requests.

== Installation ==

1. Install and activate the plugin. WooCommerce must be active.
2. WooCommerce → Crawler Shield → Rules. The defaults suit most shops; if unsure, choose Monitor for a day and look at the report.
3. If the site is behind Cloudflare or another proxy, set "Visitor address" in Rules → Server.

== Frequently Asked Questions ==

= Will it hurt my Google rankings? =

Search engines are allowed by default. Filter parameters are closed in robots.txt with lines like "Disallow: /*?*color=" — the way Google's documentation on faceted navigation recommends — and filter pages get noindex, follow. Categories and products stay open. If your filter pages are meant to rank, switch both off in the robots.txt tab.

= Will my Facebook and Instagram ads keep working? =

Yes. Link preview crawlers are never blocked by default, only slowed down; the ad review and previews keep working.

= My site has a robots.txt file =

Then WordPress does not build one and the shield cannot add its lines. The robots.txt tab shows them for copying.

= Does it replace a firewall? =

No. It solves one problem a general firewall does not: crawlers walking the endless filter combinations of a shop catalogue.

= Does it work with page caching? =

Yes. Pages served by a full-page cache never reach WordPress, so they cost nothing and the shield does not see them; it handles every request that does reach WordPress.

== Changelog ==

= 1.0.0 =
* First release: bot groups with allow / limit / no-filter-URLs / block, filter URL detection for WooCommerce and popular filter plugins, per-bot and per-address rate limits, robots.txt rules, noindex on filter pages, llms.txt, report, monitor mode, request tester, WP-CLI.
