# robots.txt for The Movie Browser # Production domain: themoviebrowser.com # # Jul 2 2026 — REVERSAL of the Jun 11 "allow AI answer engines" policy. The # Jun 11 assumption was that CloudFront serves AI crawlers cheaply from the # edge cache. ClickHouse proved that false: ClaudeBot (~1.4M views/wk) + GPTBot # (~0.5M/wk) crawl ~1.5M UNIQUE long-tail catalog URLs — every hit is an edge # cache MISS that falls through Origin Shield → an origin render → a puppeteer # ratings scrape. Those two bots were ~98% of all bot load and the dominant # driver of the Lambda ($52/mo) + Origin Shield ($10/mo) + CloudFront-request # bill, while the REAL search indexers (Googlebot, Bingbot) crawled just 6 and # 5 times in the same week. So AI training/bulk crawlers now get NO index value # for real money → block them. Search engines we actually want stay allowed. # ClaudeBot + GPTBot both honor robots.txt; the proxy also 429s them for # immediate enforcement (see src/proxy.ts BLOCKED_BOT_TYPES + .claude/rules/cdn.md). # --------------------------------------------------------------------------- # TIER 2 — SEARCH INDEXERS WE WANT: allow site, block API + private paths # --------------------------------------------------------------------------- User-agent: Googlebot User-agent: Bingbot User-agent: Applebot # OpenAI's user-facing fetchers (a human asking ChatGPT to open a URL / ChatGPT # search) are kept — low volume, real intent. The bulk GPTBot crawler is blocked # below. Same rationale for keeping these but blocking the training crawlers. User-agent: OAI-SearchBot User-agent: ChatGPT-User Disallow: /api/ Disallow: /ratings Disallow: /watchlist Disallow: /watched # DISCUSSION SUB-PAGES — see the note in the DEFAULT group below. These are # ~1.2M near-empty shells and were 41% of Googlebot's crawl on Aug 3 2026. Disallow: /*/discussions Disallow: /*/discuss/ Allow: / # --------------------------------------------------------------------------- # TIER 1 — AI TRAINING/BULK CRAWLERS + PURE SCRAPERS (no search-index value): # block entirely (ClaudeBot + GPTBot also 429'd at proxy/Caddy edge) # --------------------------------------------------------------------------- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / User-agent: Google-Extended Disallow: / User-agent: PerplexityBot Disallow: / User-agent: CCBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: meta-externalagent Disallow: / User-agent: cohere-ai Disallow: / User-agent: Bytespider Disallow: / User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: PetalBot Disallow: / # --------------------------------------------------------------------------- # DEFAULT — everyone else: index public pages, never the API/SSE or private # user content. # --------------------------------------------------------------------------- User-agent: * # API + SSE enrich endpoints (GET /api/[mediaType]/[id]/enrich) — never crawl Disallow: /api/ # Private user content - not for indexing Disallow: /ratings Disallow: /watchlist Disallow: /watched # DISCUSSION SUB-PAGES (added Aug 3 2026) — a crawl-budget AND origin-CPU brake. # Every title has a `/discussions` page and every episode a `/discuss/sXeY` page: # ~1.2M+ URLs (1.07M movies + 141k series + their episodes). Site-wide there is # currently exactly **1 published comment**, so essentially all of them are empty # shells. Measured the day Googlebot was unblocked: they were 41% of Googlebot's # crawl and 29% of ALL origin-reaching requests, and because each one is a full # SSR render that writes an ISR entry, they were churning ~32k of 322k cache # entries per HOUR against a cache pinned at its size cap — evicting the popular # pages real users hit. `noindex` alone does NOT fix that (Google must still crawl # a page to see it), so the crawl has to be blocked here. # The pages keep working for direct visitors and are still `noindex` when empty in # `generateMetadata`, which self-corrects as threads gain content. # LIFT THESE TWO LINES once discussions carry real content — they are a response # to an empty feature, not a permanent policy. # NOTE: the cross-catalog hub `/discussions` is deliberately NOT matched # (`/*/discussions` needs a second slash), so the one aggregate page stays crawlable. Disallow: /*/discussions Disallow: /*/discuss/ # Allow everything else Allow: / # Sitemap location Sitemap: https://themoviebrowser.com/sitemap.xml # --------------------------------------------------------------------------- # LLM-FRIENDLY MARKDOWN LAYER (llmstxt.org) # --------------------------------------------------------------------------- # Every page has a clean markdown twin at `.md`, indexed by /llms.txt. # There is no standardised robots directive for llms.txt, so this is a comment # for the humans and agents that read robots.txt looking for it — it is NOT a # parsed instruction and no crawler acts on it. # # Agents: prefer the .md twins. They are cheaper for us to serve (a single # read-only Postgres query, no SSR, no ratings scrape) and cleaner for you to # parse, which is why they are exempt from the scraper shed that 429s bulk # crawlers on the HTML paths. Please use them instead of the HTML. # # LLM index: https://themoviebrowser.com/llms.txt # Example: https://themoviebrowser.com/movie/27205/inception.md # Search: https://themoviebrowser.com/search.md?q=QUERY # # The individual .md twins are deliberately NOT enumerated in a sitemap: they # are noindex (X-Robots-Tag) alternates of pages already in sitemap.xml, so # listing ~600k of them would add crawl cost and zero index coverage.