list_ai_crawler_pages

list_ai_crawler_pages MCP tool: the pages of your site visited by AI crawlers, with answer reads, search indexing, training crawls, citations in AI answers, Mentionable-generated flag and last visit.

Updated 2026-09-29

list_ai_crawler_pages lists the pages of the project's site that AI crawlers visited over a date range. Each row gives the visits, split into answer reads, search indexing and training crawls, the number of times the page was cited as a source in answers to the tracked prompts, whether the page was generated with Mentionable, and the day of the last visit.

The visits come from Cloudflare AI Crawl Control, synced every night through yesterday once the project is connected to Cloudflare on its AI crawlers page. History starts 7 days before the connection.

When to use

Use it to find out which pages the AI assistants actually read: "which of my pages does ChatGPT fetch to answer questions?", "are my generated pages crawled?", "which pages get crawled a lot but never cited?".

Crossing crawls with citations is the point of the tool. A page with many answer reads and no citation is read but not credited. The tool reads stored data and costs no credits.

Input

Field Type Default Description
projectId string (CUID) required Project to query.
cursor string Opaque pagination cursor from a prior pageInfo.nextCursor.
limit integer 20 1 to 100.
filters.dateRange { from?, to? } last 30 full days UTC days, bounds included. to defaults to yesterday, the latest synced day.
filters.crawlers string[] Crawler ids, e.g. chatgpt-user, gptbot, claudebot.
filters.operators string[] OpenAI, Anthropic, Perplexity, Mistral, DuckDuckGo, Apple, Amazon, Meta, ByteDance, Common Crawl.
filters.kinds string[] user (answer reads), search (search indexing), training (training crawls).
filters.pathContains string Case-insensitive substring of the URL path, up to 200 characters.
filters.generated boolean true keeps only pages published with Mentionable, false leaves them out.
filters.statusClasses string[] 2xx, 3xx, 4xx, 5xx. Counts only the visits that got a code of these classes, e.g. ["4xx", "5xx"] for pages in error. Pages without such visits are left out.
sortBy enum visits_desc visits_desc, answer_reads_desc, citations_desc, last_visit_desc.

The crawler ids, operators and kinds are listed on get_ai_crawler_overview. Crawler, operator and kind filters add up and narrow the visits counted on each page.

Response

data holds one row per page, pageInfo the usual pagination envelope, and summary the range actually read and the last synced day.

{
  "data": [
    {
      "id": "/blog/geo-audit-checklist",
      "path": "/blog/geo-audit-checklist",
      "visits": 318,
      "answerReads": 64,
      "searchIndexing": 121,
      "trainingCrawls": 133,
      "citations": 9,
      "generated": true,
      "statuses": [{ "status": 200, "visits": 318 }],
      "lastSeenAt": "2026-09-28"
    },
    {
      "id": "/pricing",
      "path": "/pricing",
      "visits": 204,
      "answerReads": 41,
      "searchIndexing": 70,
      "trainingCrawls": 93,
      "citations": 0,
      "generated": false,
      "statuses": [{ "status": 200, "visits": 198 }, { "status": 404, "visits": 6 }],
      "lastSeenAt": "2026-09-27"
    }
  ],
  "pageInfo": { "hasMore": true, "nextCursor": "20", "totalCount": 146 },
  "summary": {
    "connected": true,
    "dateRange": { "from": "2026-08-30", "to": "2026-09-28" },
    "syncedThrough": "2026-09-28"
  }
}
  • id and path: the URL path, without the domain and without a trailing slash. Pass it as is to get_ai_crawler_page.
  • visits: visits over the range, with the crawler filters applied. answerReads, searchIndexing and trainingCrawls split them by kind.
  • citations: times a URL with this path, on your domain or one of its subdomains, was cited as a source in an answer to a tracked prompt over the range. All engines are counted, whatever the crawler filters.
  • generated: true when a page generated with Mentionable was published at this path.
  • statuses: HTTP codes served to crawlers on this page with their visits, most frequent first. A 4xx or 5xx means the crawler could not read the page.
  • lastSeenAt: the last day in the range with a visit (YYYY-MM-DD).

When the project is not connected to Cloudflare, the tool returns an empty data with summary: { "connected": false }. An invalid range (from after to) or filters that match no crawler return an empty data with summary: { "connected": true }.

Tips and patterns

  • Filter statusClasses: ["4xx", "5xx"] to list the pages AI crawlers could not read, then fix them first: an unreadable page cannot be cited.
  • sortBy: "answer_reads_desc" ranks the pages assistants fetch live to answer. They are the pages worth keeping fresh and fast.
  • Look for rows with a high answerReads and citations: 0: the engine reads the page but does not credit it. Run a Page Audit on it with run_page_audit.
  • filters: { generated: true } checks whether the pages you published with Mentionable are picked up, then get_ai_crawler_page gives the timeline from publication to first citation.
  • filters: { pathContains: "/blog" } narrows the list to one section of the site.
  • sortBy: "last_visit_desc" surfaces pages crawled in the last days, handy right after a publication.

Related tools

You can see where you stand today.

Free trial. Start tracking your AI visibility, no credit card.