get_ai_crawler_page

get_ai_crawler_page MCP tool: one page's AI crawler activity over a date range. Visits by crawler, HTTP statuses, daily series by kind, citations by LLM and, for pages generated with Mentionable, publication to first AI visit to first citation.

Updated 2026-09-29

get_ai_crawler_page returns the AI crawler activity of a single page of the project's site over a date range: visits by crawler, the HTTP statuses served, a daily series split by kind, the citations of the page by LLM, and, when the page was generated with Mentionable, the timeline from publication to the first AI visit to the first citation.

The visits come from Cloudflare AI Crawl Control, synced every night through yesterday once the project is connected to Cloudflare on its AI crawlers page. History starts 7 days before the connection.

When to use

Use it once you know which page matters: "when did ChatGPT first read the article I published?", "which engines cite this page?", "why did this page stop being crawled?". Get the path from list_ai_crawler_pages, or pass any path of the site.

The publication block answers the question every content team asks after publishing: how long until an AI crawler reads the page, and how long until an engine cites it. The tool reads stored data and costs no credits.

Input

Field Type Default Description
projectId string (CUID) required Project to query.
path string required URL path starting with /, up to 500 characters, e.g. /blog/guide. A trailing slash is removed.
filters.dateRange { from?, to? } last 30 full days UTC days, bounds included. to defaults to yesterday, the latest synced day.
filters.crawlers string[] Crawler ids, e.g. chatgpt-user, gptbot, claudebot.
filters.operators string[] OpenAI, Anthropic, Perplexity, Mistral, DuckDuckGo, Apple, Amazon, Meta, ByteDance, Common Crawl.
filters.kinds string[] user (answer reads), search (search indexing), training (training crawls).

The crawler ids, operators and kinds are listed on get_ai_crawler_overview. The crawler filters narrow the visits only: citations and the publication timeline stay the same whatever the filters.

Response

A get_* tool: the payload sits next to success: true. All dates are UTC days (YYYY-MM-DD).

{
  "success": true,
  "path": "/blog/geo-audit-checklist",
  "dateRange": { "from": "2026-08-30", "to": "2026-09-28" },
  "visits": 318,
  "byCrawler": [
    { "crawler": "gptbot", "label": "GPTBot", "operator": "OpenAI", "kind": "training", "visits": 98 },
    { "crawler": "chatgpt-user", "label": "ChatGPT-User", "operator": "OpenAI", "kind": "user", "visits": 52 },
    { "crawler": "perplexitybot", "label": "PerplexityBot", "operator": "Perplexity", "kind": "search", "visits": 47 }
  ],
  "statuses": [
    { "status": 200, "visits": 311 },
    { "status": 304, "visits": 7 }
  ],
  "daily": [
    { "date": "2026-09-10", "user": 0, "search": 2, "training": 5 },
    { "date": "2026-09-11", "user": 3, "search": 6, "training": 4 }
  ],
  "citations": {
    "total": 9,
    "byLlm": [
      { "llm": "PERPLEXITY", "count": 5 },
      { "llm": "CHATGPT", "count": 4 }
    ]
  },
  "publication": {
    "publishedAt": "2026-09-10",
    "firstVisit": {
      "date": "2026-09-10",
      "crawler": "oai-searchbot",
      "label": "OAI-SearchBot",
      "operator": "OpenAI",
      "kind": "search"
    },
    "firstCitation": { "date": "2026-09-16", "llm": "PERPLEXITY" }
  },
  "events": [
    { "date": "2026-09-10", "type": "page", "title": "GEO audit checklist for coaches", "path": "/blog/geo-audit-checklist" },
    { "date": "2026-09-15", "type": "action", "title": "Rewrote the pricing page FAQ", "path": null }
  ]
}
  • path: the normalized path that was read.
  • visits, byCrawler: visits to the page over the range, total and per crawler, most visits first.
  • statuses: the HTTP statuses served to AI crawlers on this page, most visits first.
  • daily: one point per day of the range, days without visits at zero, split into user (answer reads), search (search indexing) and training (training crawls).
  • citations: times the page was cited as a source in answers to the tracked prompts over the range, total and per LLM (LLM enum).
  • publication: null unless a page generated with Mentionable was published at this path. publishedAt is the publication day, firstVisit the first day an AI crawler visited the page from that day on, with the crawler seen, and firstCitation the first citation after publication, up to the end of the range. Both are null until they happen.
  • events: the actions logged in the project journal over the range, plus the Mentionable pages published at this path.

A path with no visit is not an error: visits is 0 and the lists are empty. Errors: { "success": false, "error": "cloudflare_not_connected" } when the project has no Cloudflare connection, invalid_date_range when from is after to.

Tips and patterns

  • After a publication, call it with filters.dateRange.from set to the publication day: publication.firstVisit and publication.firstCitation give the two delays to report.
  • firstVisit with a search crawler means the page entered the engine's index. A first user visit means an assistant already fetched it to answer.
  • Any status of 400 or more in statuses deserves a fix: the crawler got an error instead of your content.
  • Compare citations.byLlm with byCrawler: an engine that crawls the page but never cites it is the one to work on.
  • filters: { operators: ["OpenAI"] } isolates ChatGPT's crawlers without changing the citations block.

Related tools

You can see where you stand today.

Free trial. Start tracking your AI visibility, no credit card.