# robots.txt for autobiz.digital # Last updated: 2026-08-03 # # Crawler policy is maintained, not set-and-forget. Every decision below carries # the date it was made and the reason, so a future edit is an informed one. # # Standing position: for a one-person business selling a named, accountable # service, EXPOSURE BEATS PRIVACY. Being quotable by an answer engine is worth # more than withholding text that is already public on the page. The exception # is a crawler that costs bandwidth without ever citing anyone. User-agent: * Allow: / Disallow: /api/ Disallow: /assets/ Disallow: /*.json$ # Raw RSC payloads emitted next to every page by the static export. Needed by the # client router, but they are not pages — they render as machine text. Disallow: /index.txt Disallow: /*/index.txt # --------------------------------------------------------------------------- # Classic search # --------------------------------------------------------------------------- User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Slurp Allow: / # --------------------------------------------------------------------------- # Social / link unfurling # --------------------------------------------------------------------------- User-agent: LinkedInBot Allow: / User-agent: facebookexternalhit Allow: / User-agent: Twitterbot Allow: / # --------------------------------------------------------------------------- # AI answer engines — the surfaces that cite a source back to a reader. # These are the highest-value crawlers on this list: a citation here reaches a # buyer who has already asked a purchase question. # --------------------------------------------------------------------------- # OpenAI User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / # Anthropic. Claude-Web was listed here until 2026-08-03 — that agent name is # retired and matched nothing. The two below are the live ones. User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Apple — Siri and Spotlight User-agent: Applebot Allow: / # Amazon — Rufus and Alexa answers User-agent: Amazonbot Allow: / User-agent: cohere-ai Allow: / User-agent: YouBot Allow: / User-agent: AI2Bot Allow: / User-agent: Meta-ExternalAgent Allow: / # --------------------------------------------------------------------------- # AI training crawlers — allowed 2026-08-03. # A model that read this site during training can describe the offer without # retrieving anything. For a business with no brand recognition to protect, that # is the whole game; the copy here is public either way. # --------------------------------------------------------------------------- User-agent: GPTBot Allow: / User-agent: ClaudeBot Allow: / User-agent: anthropic-ai Allow: / # Google-Extended is dual-purpose: it gates BOTH Gemini training AND inclusion # in AI Overviews. Blocking it removes the site from Google's AI answers. User-agent: Google-Extended Allow: / # Common Crawl. Unblocked 2026-08-03 — it is the shared corpus behind GPT, # Claude, Llama and Mistral, so one allow reaches many models at once. It was # blocked before this date, which was the single largest self-inflicted gap in # this file. User-agent: CCBot Allow: / # --------------------------------------------------------------------------- # Blocked # --------------------------------------------------------------------------- # Aggressive crawl rate, and not a citation surface for a Hebrew B2B service. # Pure bandwidth cost with no path back to a reader. User-agent: Bytespider Disallow: / # --------------------------------------------------------------------------- # Machine-readable context # --------------------------------------------------------------------------- Sitemap: https://autobiz.digital/sitemap.xml Host: https://autobiz.digital # Plain-language summary of what this business is, for LLMs: /llms.txt # Human-readable version of the same, with recommendation criteria: /for-ai/