Your XML validates. It loads. It sits in robots.txt. And Google still says “couldn’t fetch.” That’s not necessarily a bug. In a Search Off the Record episode published October 1, 2026, Google’s John Mueller and Martin Splitt walked through why Google may ignore a perfectly healthy sitemap.
Two reasons Google skips a valid sitemap
Mueller said there are primarily two reasons. First, host load: Google throttles how much it requests from a server. If systems are “too busy with other things,” a missed fetch is reported as couldn’t fetch. Second, crawl demand: Google may decide there’s little need to crawl more from a site. That judgment, he said, is very often based on the perceived quality of the website.
“So it’s not purely a technical thing,” Mueller said. The episode description put it plainly: the cause is often host load throttling or low crawl demand linked to perceived site quality.
For agencies, that matters. A clean sitemap offers no protection if Google has deprioritized the site. The error message doesn’t tell you whether the problem is your server, a Google-side throttle, or a quality verdict.
What Google actually reads
Two legacy fields are gone. Priority was useless because every SEO said everything was number one. Change frequency died because server-side pages are always fresh, so the signal wasn’t useful. What remains is the URL and the lastmod date — and lastmod is on probation.
If dates look reasonable, Google uses them. If every URL carries today’s date, Google simply ignores the dates. Mueller stressed it’s not a spam penalty; the signals are just set aside.
The episode also confirmed the format’s limits: 50,000 URLs and 50MB uncompressed per file. Compression doesn’t raise the ceiling. You can submit many files, link them from robots.txt, or group them through a sitemap index.
The e-commerce case, feeds and AI crawlers
Mueller’s clearest example is e-commerce. If a price changes, waiting for normal crawling can be late. A sitemap entry points Google straight to the changed product page. Publishers get a separate news sitemap and are only supposed to list the last 1,000 pages that changed.
For teams maximizing AI visibility, discovery is not as simple as submitting a file. AI crawlers usually have no console for sitemap submission. Mueller said he has seen them read generic sitemap files and feeds in his server logs, but he offered it as observation rather than documentation.
- Keep small-site sitemaps on: no harm, and CMSs generate them anyway.
- Use accurate lastmod dates; don’t stamp every URL with today’s date.
- Submit an RSS feed as a short-form sitemap for recently changed pages.
- Don’t rely on LLMs.txt: Google’s search systems don’t process it as a sitemap.
An obscure filename is a real trade-off: it stays private from competitors, but Bing and AI crawlers are less likely to find it.
Bottom line
A sitemap is a freshness and canonical signal, not a quality shortcut. Fix what Google can verify: crawlability, content quality, and accurate dates. If Google still says couldn’t fetch, treat it as a prompt to look at demand, not just XML.
Source: PPC Land



