# llms.txt — AI Crawling & Training Policy for Utah Tech Labs # Generated: 2025-09-02 # This file tells AI crawlers how they may access and use our website content. # It complements robots.txt and sitemap.xml # Reference sitemap for full URL discovery Sitemap: https://www.utahtechlabs.com/sitemap.xml # ======================= # OPENAI / CHATGPT # ======================= User-agent: GPTBot Crawl: allow Train: allow User-agent: ChatGPT-User Crawl: allow Train: allow User-agent: OAI-SearchBot Crawl: allow Train: allow # ======================= # GOOGLE / GEMINI # ======================= # Google-Extended controls content usage for Gemini and Vertex AI User-agent: Google-Extended Crawl: allow Train: allow # GoogleOther is used for AI research crawling User-agent: GoogleOther Crawl: allow Train: allow # Standard Googlebot (SEO-focused, not LLM-specific) User-agent: Googlebot Crawl: allow # ======================= # MICROSOFT / COPILOT / BING # ======================= User-agent: Microsoft-Extended Crawl: allow Train: allow User-agent: bingbot Crawl: allow User-agent: BingPreview Crawl: allow # ======================= # ANTHROPIC / CLAUDE # ======================= User-agent: ClaudeBot Crawl: allow Train: allow User-agent: Claude-Web Crawl: allow Train: allow # ======================= # PERPLEXITY AI # ======================= User-agent: PerplexityBot Crawl: allow Train: allow User-agent: Perplexity-User Crawl: allow Train: allow # ======================= # APPLE AI / APPLE INTELLIGENCE # ======================= User-agent: Applebot-Extended Crawl: allow Train: allow User-agent: Applebot Crawl: allow # ======================= # META / LLAMA / THREADS AI # ======================= User-agent: Meta-ExternalAgent Crawl: allow Train: allow # ======================= # MISTRAL AI # ======================= User-agent: MistralAI Crawl: allow Train: allow # ======================= # COHERE AI # ======================= User-agent: cohere-ai Crawl: allow Train: allow # ======================= # COMMON CRAWL (USED BY MANY AI COMPANIES) # ======================= User-agent: CCBot Crawl: allow Train: allow # ======================= # YOU.COM AI # ======================= User-agent: YouBot Crawl: allow Train: allow # ======================= # AMAZON AI # ======================= User-agent: Amazonbot Crawl: allow Train: allow # ======================= # DEFAULT RULE (CATCH-ALL) # ======================= User-agent: * Crawl: allow Train: allow # ======================= # OPTIONAL — BLOCK PRIVATE DATA # Uncomment these lines if you want to restrict LLMs from accessing sensitive areas # User-agent: * # Disallow: /admin/ # Disallow: /dashboard/ # Disallow: /internal/ # Disallow: /api/ # Disallow: /confidential/