# Long Life Clinic - robots.txt # https://longlifeclinic.com # Updated: 2026-01-16 # ===================================================== # GOOD BOTS - Full access # ===================================================== User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: /images/ User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Slurp Allow: / User-agent: facebookexternalhit Allow: / User-agent: LinkedInBot Allow: / User-agent: Twitterbot Allow: / User-agent: WhatsApp Allow: / User-agent: Applebot Allow: / User-agent: Bravebot Allow: / # ===================================================== # GENERAL CRAWLERS - Allow with restrictions # ===================================================== User-agent: * Allow: / Allow: /images/ # Note: Googlebot ignores Crawl-delay Crawl-delay: 1 # Block non-essential paths Disallow: /es/index.html Disallow: /en/index.html # /blog/ now allowed for indexing (migrated from blog.longlifeclinic.com) # Block query strings (prevent duplicate content) Disallow: /*?* # Block common WordPress/CMS paths (legacy) Disallow: /wp-admin/ Disallow: /wp-includes/ Disallow: /wp-content/ Disallow: /wp-json/ Disallow: /xmlrpc.php Disallow: /feed/ Disallow: /*.php$ # ===================================================== # MALICIOUS BOTS - Block completely # ===================================================== # ===================================================== # AI CRAWLERS — split by PURPOSE, not blanket-blocked (2026-08-14) # # These are two different jobs and they were previously blocked together: # TRAINING crawlers → harvest content for model training. Blocking = opt out. # USER/SEARCH agents → fetch a page live when someone asks the assistant a # question, and are what allow the site to be CITED. # # GA4 shows "AI Assistant" as the 3rd-largest channel (36 sessions / 90d, ahead # of Social and Referral) — earned DESPITE blocking the agents that feed it. # Blocking ChatGPT-User / OAI-SearchBot suppresses exactly that traffic while # doing nothing extra for training opt-out, which the GPTBot/CCBot rules already # cover. So: training stays blocked, citation is allowed. # Revisit if AI-referred sessions do not grow over the next quarter. # ===================================================== # --- Training crawlers: ALLOWED (changed 2026-08-14) --- # Reversed the previous blanket opt-out. Rationale, for whoever revisits this: # • This is a MARKETING site. Its content exists to be found and cited. There is # no paywall and no content-licensing revenue to protect — the reason news # publishers block (their product IS the content) does not apply here. # • Checked what comparable clinics do: biodrip.es, sensor23.com (direct Marbella # competitors), neko.health, fountainlife.com and clevelandclinic.org ALL allow # AI crawlers — none of them even mention GPTBot/CCBot/ClaudeBot in robots.txt. # Blocking meant competitors get represented in AI answers and this clinic did not. # • GA4 already shows "AI Assistant" as the 3rd-largest channel (36 sessions/90d). # To reverse: change these five back to `Disallow: /`. Nothing else depends on it. User-agent: GPTBot Allow: / User-agent: CCBot Allow: / User-agent: ClaudeBot Allow: / User-agent: anthropic-ai Allow: / User-agent: Claude-Web Allow: / User-agent: Google-Extended Allow: / # --- User-initiated / search agents: ALLOWED so the clinic can be cited --- User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: meta-externalagent Allow: / User-agent: Applebot-Extended Allow: / User-agent: MistralAI-User Allow: / User-agent: DuckAssistBot Allow: / User-agent: cohere-ai Allow: / User-agent: YouBot Allow: / # NOTE: `User-agent: *` above is `Allow: /`, so ANY assistant not named here is # already permitted. The named entries exist to make the decision explicit, not to # create permission. Only Bytespider and the scraper/SEO bots below stay blocked. User-agent: Bytespider Disallow: / User-agent: Amazonbot Allow: / # Known Bad Bots / Scrapers User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: DotBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: SeznamBot Disallow: / User-agent: PetalBot Disallow: / User-agent: Sogou Disallow: / User-agent: Yandex Disallow: / User-agent: Baiduspider Disallow: / User-agent: MegaIndex.ru Disallow: / User-agent: serpstatbot Disallow: / User-agent: SeekportBot Disallow: / User-agent: ZoominfoBot Disallow: / # Email Harvesters User-agent: EmailCollector Disallow: / User-agent: EmailSiphon Disallow: / User-agent: WebEmailExtractor Disallow: / # Content Scrapers User-agent: HTTrack Disallow: / User-agent: WebCopier Disallow: / User-agent: Offline Explorer Disallow: / User-agent: WebZIP Disallow: / User-agent: Teleport Disallow: / # Aggressive Crawlers User-agent: Screaming Frog Disallow: / User-agent: Xenu Disallow: / # ===================================================== # SITEMAP # ===================================================== Sitemap: https://longlifeclinic.com/sitemap.xml Sitemap: https://longlifeclinic.com/blog-sitemap.xml # ===================================================== # BLOCK OLD BACKUP CONTENT # ===================================================== Disallow: /backup-old-wordpress-static-blog/ Disallow: /blog-static/