User-agent: * Allow: / Sitemap: https://invorec.com/sitemap-index.xml # ───────────────────────────────────────────────────────────────────── # AI crawler policy: allow crawlers that make citation in AI answers # possible; block crawlers that only collect training data. # Note: robots.txt compliance is voluntary — some crawlers ignore it. # ───────────────────────────────────────────────────────────────────── # --- AI search and retrieval: allowed --- # These make citation in AI answers possible. User-agent: OAI-SearchBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: DuckAssistBot Allow: / User-agent: Amazonbot Allow: / # --- User-triggered fetches: allowed --- # Fired when a person explicitly asks an assistant to read a page. # Blocking these only breaks a request the user made themselves. User-agent: ChatGPT-User Allow: / User-agent: Claude-User Allow: / User-agent: Perplexity-User Allow: / User-agent: MistralAI-User Allow: / User-agent: Meta-ExternalFetcher Allow: / # --- Training-only collectors: blocked --- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: cohere-ai Disallow: / User-agent: AI2Bot Disallow: / User-agent: Bytespider Disallow: / # --- Google and Apple control tokens --- # Google-Extended is not a crawler; it governs Gemini training AND # grounding, and cannot be split. Allowed deliberately so the site stays # eligible for citation in Gemini answers. It does not affect Googlebot # indexing or ranking either way. User-agent: Google-Extended Allow: / # Applebot-Extended governs Apple Intelligence training only. User-agent: Applebot-Extended Disallow: /