# robots.txt for alacraft.day # Updated: 2026-08-28 # ----------------------------------------------- # All crawlers general rules # ----------------------------------------------- User-agent: * # Block private/system paths - no SEO value Disallow: /craftmaster-panel/ Disallow: /admin/ Disallow: /api/ # ...except the Log Checker's public service endpoints. llms.txt tells AI # assistants that the machine-readable API contract lives at # /api/v1/logs/openapi.json, and a blanket Disallow: /api/ forbade the very # crawlers we invite below from ever fetching it. These three carry no user # content: the OpenAPI document, the published size/retention limits, and the # list of redaction filters. Per-log endpoints stay blocked — they serve other # people's crash reports. # Longest matching rule wins in Google and Bing, so these override the line above. Allow: /api/v1/logs/openapi.json Allow: /api/v1/logs/limits Allow: /api/v1/logs/filters # NOTE: /*/auth/*, /*/profile, /*/logout, /*/register, /*/email/confirm/ # are intentionally NOT listed here. These pages use # instead. # A URL blocked in robots.txt cannot be crawled, so Google never reads its # noindex meta. Use either noindex OR Disallow — never both (SEO-checklist BLOCK 1.3). # Block duplicate content via URL parameters Disallow: /*?sort= Disallow: /*?filter= Disallow: /*?order= Disallow: /*?lang= # Allow all language directories and main pages explicitly Allow: /de/ Allow: /en/ Allow: /es/ Allow: /fr/ Allow: /ja/ Allow: /ru/ Allow: /uk/ Allow: /zh/ # ----------------------------------------------- # AI / LLM crawlers policy: OPEN # ----------------------------------------------- # alacraft.day explicitly welcomes all AI/LLM crawlers for indexing, # citation, and model training. No per-bot restrictions — all crawlers # follow the "User-agent: *" rules above. # # Recognized AI/LLM user-agents (informational, not enforced — they all # inherit the wildcard policy): # - OpenAI: GPTBot, ChatGPT-User, OAI-SearchBot # - Anthropic: ClaudeBot, anthropic-ai, Claude-Web # - Google: Googlebot, Google-Extended (Gemini training data) # - Perplexity: PerplexityBot # - Common Crawl: CCBot (dataset used by many LLMs) # - Apple: Applebot, Applebot-Extended # - Microsoft: Bingbot, MSNBot # - Meta: meta-externalagent, FacebookBot # - Yandex: YandexBot # - DuckDuckGo: DuckDuckBot # # If we ever need to block a specific bot, add a dedicated # "User-agent: BotName" section below with its own rules — # it will override the wildcard for that bot only. # ----------------------------------------------- # Sitemap # ----------------------------------------------- Sitemap: https://alacraft.day/sitemap.xml