Google’s John Mueller has a simple message for site owners who want AI training crawlers to find their content: make your sitemap easy to find, or lean into RSS. On the October 1 episode of Google’s Search Off the Record podcast, Mueller and Martin Splitt talked about whether sitemaps still matter — and Mueller shared what he sees in his own server logs.
The AI crawler discovery gap
Unlike Google Search, most AI training crawlers do not give site owners a submission console. There is no Search Console equivalent, Mueller said. That means the usual move — create a private sitemap with an unusual filename, leave it out of robots.txt, and submit it to Google — does not automatically help AI systems.
For sites that actually want their content picked up by AI crawlers, Mueller offered two practical paths: keep the generic sitemap.xml filename, or focus on RSS feeds. Feeds are naturally discoverable, he noted, because they are usually linked from a page’s HTML head.
He added that he has seen AI crawlers hit his sitemap and RSS files in his own server logs. He did not name the crawlers, and said he isn’t sure whether AI companies document this behavior.
Private sitemaps still work — for Google
If privacy is the goal, Mueller explained, site owners can use an unusual filename, leave the sitemap out of robots.txt, and submit it directly to Google. The tradeoff is that other systems won’t find it; Bing, for example, would need its own submission.
He also reminded site owners that a robots.txt Sitemap line is independent of the user-agent line. It isn’t tied to any single crawler’s rules, so a publicly listed sitemap remains a universal signpost.
Don’t wait on llms.txt
Asked whether llms.txt could replace an XML sitemap, Mueller was blunt. Google’s systems can’t use the Markdown file as a sitemap because it lacks the strict format, he said, comparing it to an HTML sitemap.
“I think the hope is bigger than the reality.”
He acknowledged search systems might read Markdown files someday, but “currently none of this happens.” In August, Mueller said the only crawlers on his test sites claiming to accept Markdown were SEO tools — not the AI engines marketers are chasing.
Why a valid sitemap can show ‘Couldn’t fetch’
Splitt asked why Search Console sometimes reports “Couldn’t fetch” on a valid, public sitemap linked from robots.txt. Mueller pointed to two causes outside the file itself: host load and crawl demand.
“And the crawl demand is very often based on the perceived quality of a website.”
When Google sees no need to crawl much more from a site, it can skip the sitemap. This is not purely technical; better content raises crawl demand. Search Console’s help page lists low crawl demand among possible reasons, alongside robots.txt blocks, manual actions and wrong URLs.
What marketers should do now
For teams treating AI visibility as a growth channel, this is a discoverability exercise, not a submission one. AI crawlers won’t ask politely for your files.
- Use sitemap.xml as your default filename and list it in robots.txt.
- Maintain an RSS feed and link it in your HTML head.
- Check server logs for AI crawler hits on sitemap and feed URLs.
- Test llms.txt if you like, but do not make it your primary discovery mechanism.
If your sitemap shows “Couldn’t fetch,” diagnose host load and content quality before blaming the file format. A crawlable site with fresh, high-quality pages earns more demand.
Source: Search Engine Journal



