# As a condition of accessing this website, you agree to abide by the following # content signals (also sent as Content-Signal HTTP header on HTML responses): # # (a) If a Content-Signal = yes, you may collect content for the corresponding use. # (b) If a Content-Signal = no, you may not collect content for the corresponding use. # (c) Absence of a Content-Signal grants no permission and reserves no rights. # # Signals: # search — building a search index, returning hyperlinks and short excerpts. # ai-input — using content as real-time input to AI models (RAG, generative answers). # ai-train — training or fine-tuning AI models. # # Policy: search=yes, ai-input=yes, ai-train=no. # RESTRICTIONS EXPRESSED ABOVE ARE EXPRESS RESERVATIONS OF RIGHTS UNDER ARTICLE 4 # OF EU DIRECTIVE 2019/790 ON COPYRIGHT IN THE DIGITAL SINGLE MARKET. # ============================================================================ # Default policy — all unspecified user agents # ============================================================================ User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no # Filter / sort / view params (canonical handles duplicates) Disallow: /*?sort= Disallow: /*?attr= Disallow: /*&sort= Disallow: /*&attr= Disallow: /*?filter= Disallow: /*&filter= Disallow: /*?view= Disallow: /*&view= Disallow: /*?limit= Disallow: /*&limit= # ?page= NOT blocked — Google needs pagination to see canonical → page 1 # Internal search results — zero index value, najcięższe query (ts_rank full-text). # Bingbot crawlował /szukam masowo → 7 workerów na 100% CPU, load ~11 (2026-07-01). # # UWAGA: /szukam robi 301 na /search (hooks.server.ts), więc blokada samego /szukam # była bezskuteczna. Pomiar 25.07.2026 (motor-x.pl): /szukam = 3 req, /search = 13 806 # (10 306 Googlebot + 2 727 bingbot + 773 ludzi). Blokujemy CEL redirectu, nie źródło. # # Incydent 25.07.2026: parasite SEO. Spamerzy rozsiali w sieci ~4,4 tys. linków # /search?q=+, Googlebot je crawlował (31 740 req/dobę, ASN 15169), # licząc na wzmiankę ich domeny w treści zaufanego serwisu. Skutek uboczny: drogie # full-text query saturowały pool pgbouncera → fala 504. 22 spam-URL-e weszły do indexu GSC. # Reguła WAF blokuje query z osadzoną domeną; ten Disallow usuwa problem u źródła # (bot nie kolejkuje URL-i). /*/search obejmuje warianty językowe na .com (/fr/search, /es/search…). Disallow: /szukam Disallow: /search Disallow: /*/search # Private sections Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ # ============================================================================ # AI search / citation bots — ALLOW (citation = traffic) # Private paths still disallowed. # ============================================================================ User-agent: ChatGPT-User Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / User-agent: OAI-SearchBot Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / # OAI-AdsBot — weryfikacja stron docelowych i feedu ChatGPT Ads (MXX-2901). # OpenAI wymaga, by strony docelowe reklam nie blokowały OAI-AdsBot ani OAI-SearchBot. User-agent: OAI-AdsBot Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / User-agent: PerplexityBot Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / User-agent: Perplexity-User Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / User-agent: Claude-SearchBot Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / User-agent: Claude-User Disallow: /customer/ Disallow: /checkout/ Disallow: /cart Disallow: /api/ Disallow: /search Disallow: /*/search Allow: / # ============================================================================ # AI training bots — BLOCK (content extraction without return value) # ============================================================================ User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: anthropic-ai Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: meta-externalfetcher Disallow: / User-agent: meta-webindexer Disallow: / User-agent: MistralAI-User Disallow: / User-agent: ProRataInc Disallow: / User-agent: DuckAssistBot Disallow: / User-agent: archive.org_bot Disallow: / User-agent: CloudflareBrowserRenderingCrawler Disallow: / # ============================================================================ # Unwanted search crawlers — BLOCK (no meaningful traffic for EU moto shop, # heavy long-tail crawling of dead URLs). Also enforced at Cloudflare WAF. # ============================================================================ User-agent: PetalBot Disallow: / User-agent: AspiegelBot Disallow: / # Yandex (rosyjska wyszukiwarka) — zero ruchu dla EU sklepu moto, ciężki crawl # long-tail martwych URL-i. `User-agent: Yandex` obejmuje wszystkie roboty Yandex # (YandexBot, YandexImages, ...) per ich dokumentacja. User-agent: Yandex Disallow: / # Backlink / SEO-index crawlery — nie wyszukiwarki, zero ruchu, ciężki crawl # long-tail martwych URL-i (Yii2 legacy). Łącznie biły ~400k 404/dobę. # DotBot (Moz Link Explorer) User-agent: dotbot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: IbouBot Disallow: / User-agent: Barkrowler Disallow: / # DataForSEO — SEO-data SaaS/API. Crawl służy budowie bazy on-page/treści/linków # dostępnej dla każdego klienta (w tym konkurencji audytującej motor-x). Nasz # rank tracking idzie ich SERP API server-side i NIE zależy od tego crawla. User-agent: DataForSeoBot Disallow: / Sitemap: https://www.motor-x.it/sitemap.xml