AI crawler log analysis

Which Pages Do AI Crawlers Actually Read?

Your access log already knows. Upload it and KinetixSEO compares every page in your sitemap with the requests GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and other AI crawlers really made, crawler by crawler, so you see which pages each one read in that window and which it skipped.

Create a free account to upload a log

Log analysis needs a free account, because the report is saved for you to come back to.

What the report shows

Robots.txt says who may come in. The log says who did.

Per crawler, never blended

GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot and the other AI crawlers in your log each get their own row: how many of your sitemap pages it fetched in the window, and which it skipped. GPTBot reading a page while ClaudeBot never does is exactly the kind of difference a single "fetched by AI" number hides.

Skipped pages grouped by section

A list of four thousand URLs is not a finding. The report groups the pages a crawler skipped by section of your site, least-read first, so a pattern such as "no AI crawler touched /product/ this week" stands out.

Verified crawlers, not user-agent strings

Requests are checked with forward-confirmed reverse DNS wherever the vendor documents it. Requests that fail are left out, and crawlers that cannot be verified are labelled as based on the user-agent alone.

A firewall note when a crawler is blocked

For a site you track, a crawler with low coverage is set against our latest reachability check. When that check found the crawler blocked, the report says the likely cause is a firewall rule, not your content.

The same upload also shows how many people clicked through to your site from AI answers, and where crawlers waste crawl budget on errors and parameterised URLs.

How it works

  1. 1

    Download your access log

    The Apache or Nginx access log in the standard combined format, as a .log or .txt file up to 50 MB.

  2. 2

    Upload it to KinetixSEO

    Pick the tracked site the log belongs to, or type its domain. The raw file is deleted the moment it has been parsed.

  3. 3

    Read the report

    We read your site's robots.txt and sitemap and compare every page it declares with the crawler requests in the log.

What this doesn't do

A fetched page is not a cited page: being read by a crawler says nothing about whether an AI answer quotes it. A page missing from your log was not fetched in that window, which is not the same as never, and every figure in the report carries the window it measured. When a crawler made too few requests to judge, the report says "too little data" instead of printing a percentage. It reads logs you upload; it does not collect traffic from your site continuously.

FAQ

AI crawler log analysis — FAQ

How do I find out which pages AI crawlers actually read? +

Look at your server's access log, not at robots.txt. robots.txt only says which crawlers are allowed; the access log records which pages GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and other AI crawlers really requested. Upload that log to KinetixSEO and the report compares it with the pages your sitemap declares, crawler by crawler, so you see which pages each one fetched in the log's window and which it skipped, grouped by section of your site.

Does a fetched page mean an AI engine will cite it? +

No. A crawler reading a page is a precondition for being cited, not a promise of it. The report tells you which pages AI crawlers read; whether an AI answer then quotes a page depends on the question being asked and on the other sources the engine found. Citation tracking measures that second half separately.

How do you know a request really came from GPTBot? +

A user-agent string is easy to fake, so the report checks the requesting address with forward-confirmed reverse DNS for every crawler whose vendor documents that method. Requests that fail the check are left out. Crawlers whose vendor publishes no verification method are still counted, and the report labels them as based on the user-agent alone, so you can see which figures rest on which evidence.

What if my log only covers a few days? +

A page missing from a short log was not fetched in those days, which is not the same as never. The report states the exact window it measured on every figure. When a crawler made too few requests for the size of your sitemap to say which pages it skips, it shows "too little data" instead of a coverage percentage, and a short window gets its own note.

What happens to my access log? +

The raw log file is deleted as soon as it has been parsed. Only the aggregated report is kept, for the retention window of your plan, and it never stores an IP address. You can delete a report yourself at any time. To make the comparison, the report reads your site's own robots.txt and sitemap.

See which pages AI crawlers read

Upload one access log and get the report per crawler, with the window and the evidence behind every figure.