Breaking
ChatGPT ads: up to 34% invalid clicks on some accountsJohn Mueller: Use Default Sitemap or RSS for AI CrawlersAI Agent Traffic Jumps 1,700% as Google Fights SERP ScrapingGoogle's Aug 17 change ends 4,000% Shopping ROASTikTok Opens 400,000-App Ad Network to US AdvertisersChatGPT ads: up to 34% invalid clicks on some accountsJohn Mueller: Use Default Sitemap or RSS for AI CrawlersAI Agent Traffic Jumps 1,700% as Google Fights SERP ScrapingGoogle's Aug 17 change ends 4,000% Shopping ROASTikTok Opens 400,000-App Ad Network to US Advertisers
SEO

62% of Top Sites With Sitemaps Fail at Least One Check

A lintlab crawl of Chrome UX Report top 1,000 origins found most sitemaps carry flawed URLs or dates, and only 66 serve llms.txt.

62% of Top Sites With Sitemaps Fail at Least One Check

A new crawl of the web’s most visited origins found that basic SEO plumbing is often imperfect. lintlab, which sells website-checking tools, scanned the global August 2026 Chrome UX Report top-1,000 origins. Of the 916 sites that responded as websites, 400 had a sitemap — and 248 of those carried at least one measurable issue.

Clean XML, messy signals

The faults were rarely in the XML itself. Malformed XML appeared on only four sites, and missing or wrong namespaces on 15. The problems were in the URLs and dates. Redirects were the most common issue, on 92 sites. Almost as many, 88 sites, had at least one multi-URL sitemap where every entry shared the same lastmod value.

That flat timestamp usually records when the file was generated, not when a page actually changed. Google has said it uses lastmod only when it is consistently and verifiably accurate. Bing has gone further, calling sitemaps and accurate lastmod values key inputs for deciding what to recrawl as AI-assisted search grows.

The denominator matters

Across the 916 responding origins, 43.7% had a sitemap where lintlab looked. But 165 origins were blocked, 77 were kept out by robots rules, and 18 hit fetch errors. A narrower denominator — the 656 origins the crawler could examine — puts the share at 61%.

Blocked is not missing. Large sites now routinely challenge third-party crawlers, so any scan is likely a floor, not a full census. Cloudflare began blocking training and agent crawlers by default on ad-carrying pages for newly onboarded domains in September 2026.

What to check in your own sitemap

  • Replace redirecting sitemap URLs with their final destinations.
  • Use lastmod values that reflect actual page updates, not a single generation date.
  • Keep listed URLs on the same host and scheme as the origin.
  • Remove duplicate entries before they consume the 50,000-URL limit.

llms.txt: adoption is still thin

In a second pass, lintlab requested /llms.txt on the same 1,000 origins. Only 66 returned HTTP 200 with a text or Markdown content type. Another 112 returned 200 with other content, usually an HTML “not found” page. lintlab itself cautions that publishing the file does not show any AI system reads it.

That matches the wider evidence. Ahrefs logs across 137,000 domains showed 97% of llms.txt files received no requests in May 2026. Google has also said the file neither helps nor hurts rankings.

The practical takeaway is less about llms.txt and more about the old plumbing. Sitemaps remain the main way to signal new and updated pages to Google, and recrawl timing can determine whether the price Googlebot last saw matches the one at checkout.

Source: PPC Land

Leave a Reply