Launch offer: your first press release for €9.95 (normally €49.95) Claim offer →
Pressonify.ai
Log in Create PR
AI Crawler Audit

AI Crawler Audit: Check Access for Every Major AI Crawler

Complete guide to auditing your site for AI crawler discoverability. Learn how to configure robots.txt, llms.txt and XML sitemaps, and verify that the major AI crawlers can access your content.

5 Related Articles
5 min read
Last Updated: September 22, 2026
7 Sections

Why Audit for AI Crawlers?

Traditional SEO audits check for search engine crawlers (Googlebot, Bingbot). But in 2026, AI crawlers are equally important:

  • OAI-SearchBot (OpenAI) - Surfaces sites in ChatGPT search results
  • GPTBot (OpenAI) - Collects content for model training
  • ChatGPT-User (OpenAI) - Fetches pages when a ChatGPT user asks it to
  • ClaudeBot / Claude-SearchBot / Claude-User (Anthropic) - training, search indexing and user-requested fetches for Claude
  • PerplexityBot - Indexes pages for Perplexity answers
  • Google-Extended - robots.txt token that controls use of your content for Gemini
  • Applebot-Extended - Controls use of your content for Apple Intelligence
  • Amazonbot - Amazon Alexa and AI services
  • Bingbot - Bing search, which also grounds Microsoft Copilot
  • Bytespider - TikTok/ByteDance AI
  • Meta-ExternalAgent - Meta AI (Facebook, Instagram)

Blocking the wrong one of these can remove you from a major AI platform. An AI Crawler Audit ensures you're maximizing visibility in the Citation Economy. Every release on the Pressonify platform is published on pages open to all of these crawlers.

robots.txt Configuration

Step 1 of any AI crawler audit: check your robots.txt file (yoursite.com/robots.txt). You should explicitly allow major AI crawlers:

# robots.txt - AI Crawler Configuration

# Allow major AI crawlers
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: Bingbot
Allow: /

User-agent: Bytespider
Allow: /

User-agent: Meta-ExternalAgent
Allow: /

# Optional: Protect sensitive areas
User-agent: *
Disallow: /admin/
Disallow: /private/

Critical mistake: Many sites have User-agent: * with Disallow: / which blocks ALL crawlers including AI. If you see this, you're invisible to AI systems. Fix immediately. Learn more in our LLMO guide.

llms.txt Verification

Step 2: Verify you have a properly formatted llms.txt file (yoursite.com/llms.txt). This is the 'AI context' file that provides AI crawlers with structured information about your site.

Checklist:

  • ✅ File exists at domain root (/llms.txt)
  • ✅ File is publicly accessible (200 status code, no authentication)
  • ✅ File is concise (use llms-full.txt for extended content)
  • ✅ Shows when it was last updated
  • ✅ Contains site description and expertise areas
  • ✅ Lists 10-20 key pages with descriptions
  • ✅ Includes contact information
  • ✅ Kept current as key pages change

Use our free llms.txt generator to create a compliant file in 60 seconds. View our live llms.txt example for reference. Full implementation details in our llms.txt guide.

Sitemap Check

Step 3: Verify your sitemap. AI search crawlers use the same XML sitemaps as search engines, so check that:

  • Every public page is listed, including new releases
  • lastmod dates are accurate and change only when the content does
  • No dead ends: no redirected, noindexed or 404 URLs
  • robots.txt points to it with a Sitemap line

Google ignores changefreq and priority, so an entry only needs a URL and an honest lastmod:

<url>
  <loc>https://pressonify.ai/learn/geo</loc>
  <lastmod>2026-01-03</lastmod>
</url>

Include your sitemap in robots.txt:

Sitemap: https://yoursite.com/sitemap.xml

AI Discovery Protocol (ADP) Endpoints

Step 4: Audit for full AI Discovery Protocol (ADP) compliance. Check for these endpoints:

  • ✅ /ai-discovery.json - Main ADP manifest
  • ✅ /robots.txt - AI crawler permissions
  • ✅ /llms.txt - Compact site context
  • ✅ /llms-full.txt - Extended content (optional)
  • ✅ /sitemap.xml - Standard sitemap
  • ✅ /feed.json - JSON Feed v1.1
  • ✅ /rss.xml - Traditional RSS feed
  • ✅ /knowledge-graph.json - Schema.org entity catalog
  • ✅ /.well-known/security.txt - Security contact

Use our AI Visibility Checker to check the essentials automatically.

HTTP Header Verification

Step 5: Check that your AI-related endpoints include proper HTTP headers:

  • ETag: Cache validation for efficient re-crawling
  • Content-Digest: SHA-256 integrity verification (RFC 9530)
  • X-Update-Frequency: Signals to AI crawlers (hourly/daily/weekly)
  • X-LLM-Optimized: Indicates AI-optimized content
  • Access-Control-Allow-Origin: CORS for AI tools (* for public content)
  • Cache-Control: Appropriate caching directives

Test headers using:

curl -I https://yoursite.com/llms.txt

Look for:

HTTP/2 200
ETag: W/"abc123"
Content-Digest: sha-256=xyz789
X-Update-Frequency: weekly
Access-Control-Allow-Origin: *

Pressonify's own discovery endpoints, such as /llms.txt, send these headers, and every release it publishes is listed in them. Learn more about technical implementation in our Schema.org for AI guide.

Action Plan: Fixing Common Issues

Based on your audit, prioritize fixes:

🔴 Critical (fix immediately):

  • robots.txt blocking AI crawlers with Disallow: /
  • Key pages returning errors or hidden behind logins
  • 404 errors on referenced endpoints in ai.json

🟡 High Priority (fix this week):

  • Outdated llms.txt (lastModified > 6 months old)
  • No Schema.org markup on key pages

🟢 Medium Priority (fix this month):

  • Missing llms.txt (a helpful extra for AI agents)
  • Missing HTTP headers (ETag, Content-Digest)
  • No RSS or JSON Feed for new content

Start with Critical fixes to get AI crawlers accessing your site, then layer in higher-level optimizations. Track progress with our AI Visibility Checker.

Frequently Asked Questions

For most businesses seeking visibility, allow all major AI crawlers. Only block if you have proprietary content, paywalled resources, or specific ethical concerns about AI training on your data.
Quarterly for most sites. New AI crawlers emerge regularly, so periodic audits ensure you're not accidentally blocking new platforms. Update your robots.txt when new major crawlers launch.
Update your robots.txt to allow them, then submit your sitemap to accelerate re-crawling. Re-crawl timing varies by crawler, from days to weeks.
No. Crawlable pages, a sitemap and a sensible robots.txt matter most. llms.txt and the other ADP endpoints add extra, cheap discovery paths but aren't required.

Audit Your AI Crawler Access

Run our free AI Visibility Checker to see which AI crawlers can reach your site and whether your llms.txt is in place.