list_ai_crawler_pages
list_ai_crawler_pages MCP tool: the pages of your site visited by AI crawlers, with answer reads, search indexing, training crawls, citations in AI answers, Mentionable-generated flag and last visit.
Updated 2026-09-29
list_ai_crawler_pages lists the pages of the project's site that AI crawlers visited over a date range. Each row gives the visits, split into answer reads, search indexing and training crawls, the number of times the page was cited as a source in answers to the tracked prompts, whether the page was generated with Mentionable, and the day of the last visit.
The visits come from Cloudflare AI Crawl Control, synced every night through yesterday once the project is connected to Cloudflare on its AI crawlers page. History starts 7 days before the connection.
When to use
Use it to find out which pages the AI assistants actually read: "which of my pages does ChatGPT fetch to answer questions?", "are my generated pages crawled?", "which pages get crawled a lot but never cited?".
Crossing crawls with citations is the point of the tool. A page with many answer reads and no citation is read but not credited. The tool reads stored data and costs no credits.
Input
| Field | Type | Default | Description |
|---|---|---|---|
projectId |
string (CUID) | required | Project to query. |
cursor |
string | Opaque pagination cursor from a prior pageInfo.nextCursor. |
|
limit |
integer | 20 | 1 to 100. |
filters.dateRange |
{ from?, to? } |
last 30 full days | UTC days, bounds included. to defaults to yesterday, the latest synced day. |
filters.crawlers |
string[] | Crawler ids, e.g. chatgpt-user, gptbot, claudebot. |
|
filters.operators |
string[] | OpenAI, Anthropic, Perplexity, Mistral, DuckDuckGo, Apple, Amazon, Meta, ByteDance, Common Crawl. |
|
filters.kinds |
string[] | user (answer reads), search (search indexing), training (training crawls). |
|
filters.pathContains |
string | Case-insensitive substring of the URL path, up to 200 characters. | |
filters.generated |
boolean | true keeps only pages published with Mentionable, false leaves them out. |
|
filters.statusClasses |
string[] | 2xx, 3xx, 4xx, 5xx. Counts only the visits that got a code of these classes, e.g. ["4xx", "5xx"] for pages in error. Pages without such visits are left out. |
|
sortBy |
enum | visits_desc |
visits_desc, answer_reads_desc, citations_desc, last_visit_desc. |
The crawler ids, operators and kinds are listed on get_ai_crawler_overview. Crawler, operator and kind filters add up and narrow the visits counted on each page.
Response
data holds one row per page, pageInfo the usual pagination envelope, and summary the range actually read and the last synced day.
{
"data": [
{
"id": "/blog/geo-audit-checklist",
"path": "/blog/geo-audit-checklist",
"visits": 318,
"answerReads": 64,
"searchIndexing": 121,
"trainingCrawls": 133,
"citations": 9,
"generated": true,
"statuses": [{ "status": 200, "visits": 318 }],
"lastSeenAt": "2026-09-28"
},
{
"id": "/pricing",
"path": "/pricing",
"visits": 204,
"answerReads": 41,
"searchIndexing": 70,
"trainingCrawls": 93,
"citations": 0,
"generated": false,
"statuses": [{ "status": 200, "visits": 198 }, { "status": 404, "visits": 6 }],
"lastSeenAt": "2026-09-27"
}
],
"pageInfo": { "hasMore": true, "nextCursor": "20", "totalCount": 146 },
"summary": {
"connected": true,
"dateRange": { "from": "2026-08-30", "to": "2026-09-28" },
"syncedThrough": "2026-09-28"
}
}
idandpath: the URL path, without the domain and without a trailing slash. Pass it as is toget_ai_crawler_page.visits: visits over the range, with the crawler filters applied.answerReads,searchIndexingandtrainingCrawlssplit them by kind.citations: times a URL with this path, on your domain or one of its subdomains, was cited as a source in an answer to a tracked prompt over the range. All engines are counted, whatever the crawler filters.generated: true when a page generated with Mentionable was published at this path.statuses: HTTP codes served to crawlers on this page with their visits, most frequent first. A4xxor5xxmeans the crawler could not read the page.lastSeenAt: the last day in the range with a visit (YYYY-MM-DD).
When the project is not connected to Cloudflare, the tool returns an empty data with summary: { "connected": false }. An invalid range (from after to) or filters that match no crawler return an empty data with summary: { "connected": true }.
Tips and patterns
- Filter
statusClasses: ["4xx", "5xx"]to list the pages AI crawlers could not read, then fix them first: an unreadable page cannot be cited. sortBy: "answer_reads_desc"ranks the pages assistants fetch live to answer. They are the pages worth keeping fresh and fast.- Look for rows with a high
answerReadsandcitations: 0: the engine reads the page but does not credit it. Run a Page Audit on it withrun_page_audit. filters: { generated: true }checks whether the pages you published with Mentionable are picked up, thenget_ai_crawler_pagegives the timeline from publication to first citation.filters: { pathContains: "/blog" }narrows the list to one section of the site.sortBy: "last_visit_desc"surfaces pages crawled in the last days, handy right after a publication.
Related tools
get_ai_crawler_overview: site-wide totals, trends and errorsget_ai_crawler_page: one page in detaillist_ai_crawler_visits: the raw visit rows behind these totalslist_llm_sources: every domain the LLMs cite or search