# Global rules User-agent: * Allow: / Allow: /*.jpg$ Allow: /*.jpeg$ Allow: /*.gif$ Allow: /*.png$ Allow: /*.webp$ Allow: /*.svg$ Allow: /*.pdf$ Allow: /*.doc$ Allow: /*.docx$ Allow: /*.xls$ Allow: /*.xlsx$ Allow: /llms.txt Allow: /llms-full.txt Allow: /pricing.md Allow: /index.md Allow: /openapi.json Allow: /api/errors.md Allow: /ask Allow: /agents.md Allow: /.well-known/ # Product-internal file path patterns used as code examples in technical # articles (e.g., /workspace/repoA, /scratch/scratch.md, /research-notes). # These never resolve to real routes — all real marketing pages live under # /{lang}/... (en, zh, de, es, fr, ja, ko, pt). Blocking the bare top-level # forms keeps Google's URL-discovery-from-text from speculatively crawling # phantom paths and 404'ing them. Disallow: /workspace Disallow: /scratch Disallow: /research Disallow: /specs Disallow: /drafts Disallow: /reviews Disallow: /plans Disallow: /db Disallow: /runs Disallow: /policies Disallow: /output Disallow: /finance Disallow: /drive Disallow: /notes Disallow: /handoff Disallow: /codebase Disallow: /repos Disallow: /quit Disallow: /executions Disallow: /n8n Disallow: /agent-state # Use trailing slash so we only block the /agents/* path prefix, NOT # the standalone /agents.md agent-readable resource (which is Allowed above). # Bing/Bingbot's robots parser does longest-match prefix matching and is # stricter than Googlebot about trailing slashes, so be explicit here. Disallow: /agents/ # Use trailing slash to preserve the real /context-base marketing route Disallow: /context/ # Old docs paths — these routes don't exist in the current codebase Disallow: /en/reference Disallow: /zh/reference Disallow: /en/connect Disallow: /zh/connect Disallow: /en/distribute Disallow: /zh/distribute Disallow: /en/quickstart Disallow: /zh/quickstart Disallow: /en/concepts Disallow: /zh/concepts Disallow: /en/audit Disallow: /zh/audit Disallow: /en/auth-for-agents Disallow: /zh/auth-for-agents Disallow: /en/editor Disallow: /zh/editor Disallow: /en/conflict Disallow: /zh/conflict Disallow: /doc/ # Old joinWaitingList feature — removed, all pages 404 Disallow: /joinWaitingList Disallow: /en/joinWaitingList Disallow: /zh/joinWaitingList Disallow: /de/joinWaitingList Disallow: /fr/joinWaitingList Disallow: /es/joinWaitingList Disallow: /pt/joinWaitingList Disallow: /ja/joinWaitingList Disallow: /ko/joinWaitingList # Cloudflare internal URL Disallow: /cdn-cgi/ # Phantom marketing pages — routes don't exist Disallow: /signup Disallow: /demo Disallow: /get-started Disallow: /case-studies Disallow: /solutions Disallow: /templates Disallow: /login Disallow: /userLoveWall Disallow: /zh/userLoveWall # Host Host: https://www.puppyone.ai # Sitemaps # Listing both the index and per-language sitemaps. Bing/Yandex/Baidu treat # direct sitemap URLs more aggressively than nested entries inside # an index, so we hint both forms here. Sitemap: https://www.puppyone.ai/sitemap_index.xml Sitemap: https://www.puppyone.ai/sitemap-en.xml Sitemap: https://www.puppyone.ai/sitemap-zh.xml Sitemap: https://www.puppyone.ai/sitemap-ja.xml Sitemap: https://www.puppyone.ai/sitemap-ko.xml Sitemap: https://www.puppyone.ai/sitemap-de.xml Sitemap: https://www.puppyone.ai/sitemap-fr.xml Sitemap: https://www.puppyone.ai/sitemap-es.xml Sitemap: https://www.puppyone.ai/sitemap-pt.xml Sitemap: https://www.puppyone.ai/sitemap-docs-en.xml Sitemap: https://www.puppyone.ai/sitemap-docs-zh.xml # ============================================================================= # Traditional search engines — image / mobile crawlers # ============================================================================= User-agent: Googlebot-Image Allow: /*.jpg$ Allow: /*.jpeg$ Allow: /*.gif$ Allow: /*.png$ Allow: /*.webp$ Allow: /*.svg$ User-agent: Bingbot-Image Allow: /*.jpg$ Allow: /*.jpeg$ Allow: /*.gif$ Allow: /*.png$ Allow: /*.webp$ Allow: /*.svg$ User-agent: Googlebot-Mobile Allow: / User-agent: Bingbot-Mobile Allow: / User-agent: Bingbot Allow: / User-agent: adidxbot Allow: / # Brave Search — independent index (not Google/Bing-based). Brave's own # crawler is Bravebot (a.k.a. "Bravest"); their policy is that if Googlebot # can crawl a page, Bravebot can too. Explicit allow is a signaling courtesy # even though their docs note Bravebot does not fully obey robots.txt. User-agent: Bravebot Allow: / User-agent: Bravest Allow: / # ============================================================================= # AI inference / search crawlers — ALLOW # These power real-time AI answers with citations and drive referral traffic. # ============================================================================= # Anthropic — on-demand user browse + Claude search User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / # OpenAI — ChatGPT search + user-triggered fetches User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / # Perplexity — AI search index + user-triggered fetches User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Apple — Siri / Spotlight (NOT the AI-training Applebot-Extended) User-agent: Applebot Allow: / # You.com User-agent: YouBot Allow: / # ============================================================================= # AI training crawlers — DISALLOW # These scrape content for model training without driving traffic back. # Opting out preserves copyright while keeping AI-search visibility above. # ============================================================================= # Anthropic training User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / # OpenAI training User-agent: GPTBot Disallow: / # Google Gemini training (does NOT affect Google Search indexing) User-agent: Google-Extended Disallow: / # Apple Foundation Models training User-agent: Applebot-Extended Disallow: / # Common Crawl — feeds most public training datasets User-agent: CCBot Disallow: / # Meta Llama training User-agent: Meta-ExternalAgent Disallow: / User-agent: FacebookBot Disallow: / # ByteDance (Doubao / other LLMs) training User-agent: Bytespider Disallow: / # Commercial web-intelligence scraper reselling data for AI training User-agent: Diffbot Disallow: / User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: /