AI Bot Log Analyzer: Which AI Bots Crawl Your Website?
Upload your server log file and see in under a minute which AI bots read your content, which URLs they request and how many visitors click through from ChatGPT, Claude or Perplexity. No signup, no cost.
Upload log file
Supports Apache Combined, Nginx, Common Log Format, IIS W3C. Also .gz compressed files.
What happens to your file
- The file is parsed in memory only and is never written to disk.
- No individual IP addresses are taken from the log. They only enter the result as a count of distinct addresses per bot.
- No sharing with third parties, no transfer to AI providers. The analysis runs entirely on our server in the EU.
- Only the aggregated result is stored and deleted after 90 days. A share link expires after 30 days.
You can truncate or strip the IP column before uploading, the bot analysis stays complete. Details in our privacy policy, section 3.4. A data processing agreement under Art. 28 GDPR is available on request.
AI bot database 49 Bots
Every AI bot user agent this tool recognises. As of April 2026.
| Bot | Operator | Category | Purpose | User-Agent Pattern |
|---|---|---|---|---|
| Brave-Search | Brave | AI Search | Brave Search AI answers | brave-search |
| ChatGPT-User | OpenAI | AI Search | Real-time web browsing for ChatGPT users | chatgpt-user |
| Claude-SearchBot | Anthropic | AI Search | Claude search indexing | claude-searchbot |
| Claude-User | Anthropic | AI Search | Real-time web access for Claude users | claude-user, claude user |
| CopilotBot | Microsoft | AI Search | Microsoft Copilot grounding | copilotbot |
| DeepSeek-Bot | DeepSeek | AI Search | DeepSeek search | deepseek |
| Google-CloudVertexBot | AI Search | Vertex AI grounding | google-cloudvertexbot |
|
| GoogleOther | AI Search | Google AI / research crawling | googleother |
|
| Grok | xAI | AI Search | Grok web access | grok, xai-grok |
| KagiBot | Kagi | AI Search | Kagi search AI features | kagibot |
| OAI-SearchBot | OpenAI | AI Search | ChatGPT Search / SearchGPT indexing | oai-searchbot, openai searchbot |
| Perplexity-User | Perplexity | AI Search | User-triggered Perplexity search | perplexity-user, perplexity user |
| PerplexityBot | Perplexity | AI Search | Perplexity answer engine retrieval | perplexitybot |
| YouBot | You.com | AI Search | You.com AI search | youbot |
| Ai2Bot | Allen AI | AI Training | Research model training | ai2bot |
| Amazonbot | Amazon | AI Training | Alexa AI / model training | amazonbot |
| Bytespider | ByteDance | AI Training | TikTok / Douyin AI training | bytespider, bytedance |
| CCBot | Common Crawl | AI Training | Open dataset used by most LLMs | ccbot |
| Claude-Web | Anthropic | AI Training | Web content for Claude training | claude-web |
| ClaudeBot | Anthropic | AI Training | Model training data collection | claudebot, anthropic-ai |
| Diffbot | Diffbot | AI Training | Knowledge graph for AI applications | diffbot |
| FacebookBot | Meta | AI Training | Meta AI crawling | facebookbot, facebook crawler |
| GPTBot | OpenAI | AI Training | Model training data collection | gptbot |
| Meta-ExternalAgent | Meta | AI Training | LLaMA training + AI features | meta-externalagent, meta externalagent |
| Meta-Fetcher | Meta | AI Training | Meta AI content fetching | meta-fetcher, meta fetcher |
| Omgilibot | Webz.io | AI Training | AI training data | omgili, omgilibot |
| Timpibot | Timpi | AI Training | Decentralized AI search | timpibot |
| cohere-ai | Cohere | AI Training | Model training | cohere-ai |
| iaskspider | iAsk.ai | AI Training | AI search training | iaskspider |
| AdsBot-Google | Search Engine | Google Ads quality check | adsbot-google |
|
| BaiduSpider | Baidu | Search Engine | Baidu Search indexing | baiduspider |
| Bingbot | Microsoft | Search Engine | Bing Search indexing | bingbot |
| DuckDuckBot | DuckDuckGo | Search Engine | DuckDuckGo indexing | duckduckbot |
| Googlebot | Search Engine | Google Search indexing | googlebot |
|
| YandexBot | Yandex | Search Engine | Yandex Search indexing | yandexbot |
| AhrefsBot | Ahrefs | SEO Tool | Ahrefs backlink crawling | ahrefsbot |
| DataForSEOBot | DataForSEO | SEO Tool | DataForSEO crawling | dataforseobot |
| DotBot | Moz | SEO Tool | Moz crawling | dotbot |
| MJ12bot | Majestic | SEO Tool | Majestic backlink crawling | mj12bot |
| PetalBot | Huawei | SEO Tool | Petal Search crawling | petalbot |
| ScreamingFrog | Screaming Frog | SEO Tool | Screaming Frog crawler | screaming frog |
| SemrushBot | Semrush | SEO Tool | Semrush crawling | semrushbot |
| Discordbot | Discord | Discord link preview | discordbot |
|
| FacebookExternalHit | Meta | Facebook link preview | facebookexternalhit |
|
| LinkedInBot | LinkedIn link preview | linkedinbot |
||
| Slackbot | Slack | Slack link unfurling | slackbot |
|
| TelegramBot | Telegram | Telegram link preview | telegrambot |
|
| TwitterBot | X/Twitter | Twitter card rendering | twitterbot |
|
| Meta | WhatsApp link preview | whatsapp |
robots.txt directives, not crawlers
These tokens are addressed in robots.txt only. They issue no HTTP requests of their own and therefore never appear as a user agent in a log file. They are deliberately kept out of the bot database above so the tool does not count them as separate requests.
| Token | Operator | Actually crawls | Controls |
|---|---|---|---|
| Applebot-Extended | Apple | Applebot |
Use of content for Apple Intelligence and Siri |
| Google-Extended | Googlebot |
Use of content in Gemini and AI Overviews |
How the AI bot analysis works
Every request to your website leaves a line in the server log file: timestamp, requested URL, status code and the caller's user agent. That user agent is exactly where AI bots identify themselves. Google Analytics never shows them, because bots do not execute JavaScript. The log file is therefore the only reliable source for whether AI systems actually read your content.
- Export the log file Download access.log from your server or hosting panel. Compressed and rotated files work too.
- Upload Processing happens in memory. Your file is never stored on the server, only the aggregated metrics are kept.
- Review the results You get the share of AI traffic, the split by operator, your most-read URLs, the trend over time and an AI Visibility Score from 0 to 100.
AI search, AI training and LLM referrals: the difference matters
Not every AI request means the same thing for your business. The tool separates three kinds, and their order of importance surprises most people.
AI search: someone is asking right now
Bots such as ChatGPT-User, Perplexity-User or Claude-User fetch your page at the moment a human asks a question. Every one of these requests proves your content was used to support a specific answer. It is the strongest AI visibility signal a log file can give you.
AI training: material for tomorrow
GPTBot, ClaudeBot, CCBot and Bytespider collect content for future model versions. The effect only shows with the next model, but it lasts. Block these bots in robots.txt and you gradually disappear from what the models know.
LLM referrals: real visitors
When someone clicks a link to you inside ChatGPT or Perplexity, the chat domain appears in the referrer. These are not bots but people with concrete buying intent. They are the most direct revenue contribution of AI visibility, and standard analytics usually files them under Direct.
Where to find your log file
Depending on your setup the access log lives in a different place. This overview covers the most common environments. Two to four weeks of data already gives a reliable picture.
| Environment | Path or menu entry |
|---|---|
| Apache | /var/log/apache2/access.log |
| Nginx | /var/log/nginx/access.log |
| Plesk | Websites & Domains, then Logs |
| cPanel | Metrics, then Raw Access |
| IIS | C:\inetpub\logs\LogFiles |
| Cloudflare | Logpush, Enterprise plan and up |
| Managed hosting | File manager or SFTP, logs folder |
Note: some hosts serve rotated logs with a .gz extension that your browser already decompresses on download. This tool detects that from the file content and handles both variants.
Frequently asked questions
Which AI bots does the tool detect?
The tool matches every user agent against 49 stored bot signatures, among them ChatGPT-User and GPTBot from OpenAI, ClaudeBot and Claude-User from Anthropic, PerplexityBot, OAI-SearchBot, Google-CloudVertexBot, Meta-ExternalAgent, Bytespider and CCBot. It also detects clicks coming out of AI chats via the referrer domain, for example chatgpt.com or claude.ai.
Why are Google-Extended and Applebot-Extended missing from the bot database?
Because neither is a crawler. Google-Extended and Applebot-Extended are purely robots.txt control tokens: they issue no HTTP requests of their own and therefore never appear as a user agent in a log file. The actual crawling is still done by Googlebot and Applebot respectively. The tokens only determine whether the already crawled content may be used for Gemini, AI Overviews or Apple Intelligence. Counting them as separate requests would make the result wrong. They are listed separately under robots.txt directives.
Where do I find my server log files?
On Apache and Nginx the logs usually sit at /var/log/apache2/access.log or /var/log/nginx/access.log. In Plesk you find them under Websites & Domains, Logs; in cPanel under Metrics, Raw Access. Cloudflare offers Logpush from the Enterprise plan. On managed hosting, download the file through the file manager or via SFTP.
What is the difference between AI search bots and AI training crawlers?
AI search bots such as ChatGPT-User or Perplexity-User fetch your page the moment someone asks a question. They prove your content serves as a source for a specific answer. AI training crawlers such as GPTBot or CCBot instead collect material for future model versions. For your visibility today, the first group counts far more.
Are LLM referrals the same as AI bot requests?
No. An AI bot request is a crawler reading your page. An LLM referral is a real person clicking a link to your site inside ChatGPT, Claude or Perplexity. The tool reports both separately, because referrals can generate revenue directly while crawls only create the precondition for it.
Is my log data stored?
The uploaded file is processed in memory only and never written to the server. Only the aggregated metrics of the analysis are stored, meaning hit counts per bot and the most frequent URLs. Individual IP addresses from your log are not retained.
Which log formats are supported?
Apache Combined, the Nginx default format, Common Log Format and IIS W3C Extended. Gzip files are decompressed automatically, detected by file content rather than extension. That also covers rotated logs your host already decompressed on download.
Do I need an account?
No. The analysis is free and works without signing up. You only need an email address if you want the report sent to you as a PDF.
Want better AI visibility?
You now know which AI bots use your page. Let's optimize your AI visibility together.
Request SEO consulting