Why Audit for AI Crawlers?
Traditional SEO audits check for search engine crawlers (Googlebot, Bingbot). But in 2026, AI crawlers are equally important:
- OAI-SearchBot (OpenAI) - Surfaces sites in ChatGPT search results
- GPTBot (OpenAI) - Collects content for model training
- ChatGPT-User (OpenAI) - Fetches pages when a ChatGPT user asks it to
- ClaudeBot / Claude-SearchBot / Claude-User (Anthropic) - training, search indexing and user-requested fetches for Claude
- PerplexityBot - Indexes pages for Perplexity answers
- Google-Extended - robots.txt token that controls use of your content for Gemini
- Applebot-Extended - Controls use of your content for Apple Intelligence
- Amazonbot - Amazon Alexa and AI services
- Bingbot - Bing search, which also grounds Microsoft Copilot
- Bytespider - TikTok/ByteDance AI
- Meta-ExternalAgent - Meta AI (Facebook, Instagram)
Blocking the wrong one of these can remove you from a major AI platform. An AI Crawler Audit ensures you're maximizing visibility in the Citation Economy. Every release on the Pressonify platform is published on pages open to all of these crawlers.
robots.txt Configuration
Step 1 of any AI crawler audit: check your robots.txt file (yoursite.com/robots.txt). You should explicitly allow major AI crawlers:
# robots.txt - AI Crawler Configuration
# Allow major AI crawlers
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Amazonbot
Allow: /
User-agent: Bingbot
Allow: /
User-agent: Bytespider
Allow: /
User-agent: Meta-ExternalAgent
Allow: /
# Optional: Protect sensitive areas
User-agent: *
Disallow: /admin/
Disallow: /private/
Critical mistake: Many sites have User-agent: * with Disallow: / which blocks ALL crawlers including AI. If you see this, you're invisible to AI systems. Fix immediately. Learn more in our LLMO guide.
llms.txt Verification
Step 2: Verify you have a properly formatted llms.txt file (yoursite.com/llms.txt). This is the 'AI context' file that provides AI crawlers with structured information about your site.
Checklist:
- ✅ File exists at domain root (/llms.txt)
- ✅ File is publicly accessible (200 status code, no authentication)
- ✅ File is concise (use llms-full.txt for extended content)
- ✅ Shows when it was last updated
- ✅ Contains site description and expertise areas
- ✅ Lists 10-20 key pages with descriptions
- ✅ Includes contact information
- ✅ Kept current as key pages change
Use our free llms.txt generator to create a compliant file in 60 seconds. View our live llms.txt example for reference. Full implementation details in our llms.txt guide.
Sitemap Check
Step 3: Verify your sitemap. AI search crawlers use the same XML sitemaps as search engines, so check that:
- Every public page is listed, including new releases
- lastmod dates are accurate and change only when the content does
- No dead ends: no redirected, noindexed or 404 URLs
- robots.txt points to it with a Sitemap line
Google ignores changefreq and priority, so an entry only needs a URL and an honest lastmod:
<url>
<loc>https://pressonify.ai/learn/geo</loc>
<lastmod>2026-01-03</lastmod>
</url>
Include your sitemap in robots.txt:
Sitemap: https://yoursite.com/sitemap.xml
AI Discovery Protocol (ADP) Endpoints
Step 4: Audit for full AI Discovery Protocol (ADP) compliance. Check for these endpoints:
- ✅
/ai-discovery.json- Main ADP manifest - ✅
/robots.txt- AI crawler permissions - ✅
/llms.txt- Compact site context - ✅
/llms-full.txt- Extended content (optional) - ✅
/sitemap.xml- Standard sitemap - ✅
/feed.json- JSON Feed v1.1 - ✅
/rss.xml- Traditional RSS feed - ✅
/knowledge-graph.json- Schema.org entity catalog - ✅
/.well-known/security.txt- Security contact
Use our AI Visibility Checker to check the essentials automatically.
HTTP Header Verification
Step 5: Check that your AI-related endpoints include proper HTTP headers:
- ETag: Cache validation for efficient re-crawling
- Content-Digest: SHA-256 integrity verification (RFC 9530)
- X-Update-Frequency: Signals to AI crawlers (hourly/daily/weekly)
- X-LLM-Optimized: Indicates AI-optimized content
- Access-Control-Allow-Origin: CORS for AI tools (* for public content)
- Cache-Control: Appropriate caching directives
Test headers using:
curl -I https://yoursite.com/llms.txt
Look for:
HTTP/2 200
ETag: W/"abc123"
Content-Digest: sha-256=xyz789
X-Update-Frequency: weekly
Access-Control-Allow-Origin: *
Pressonify's own discovery endpoints, such as /llms.txt, send these headers, and every release it publishes is listed in them. Learn more about technical implementation in our Schema.org for AI guide.
Action Plan: Fixing Common Issues
Based on your audit, prioritize fixes:
🔴 Critical (fix immediately):
- robots.txt blocking AI crawlers with
Disallow: / - Key pages returning errors or hidden behind logins
- 404 errors on referenced endpoints in ai.json
🟡 High Priority (fix this week):
- Outdated llms.txt (lastModified > 6 months old)
- No Schema.org markup on key pages
🟢 Medium Priority (fix this month):
- Missing llms.txt (a helpful extra for AI agents)
- Missing HTTP headers (ETag, Content-Digest)
- No RSS or JSON Feed for new content
Start with Critical fixes to get AI crawlers accessing your site, then layer in higher-level optimizations. Track progress with our AI Visibility Checker.