Blog Crawling and robots.txt
Crawling and robots.txt.
Crawl budget, robots.txt, noindex, and everything that decides what Google reaches. 12 posts, newest first.
What is crawl budget, and which pages actually waste it?Google moved its crawl budget guide and now says do not use noindex to save crawling. The fetch cost of every Search Console status, row by row.
No information is available for this page: what Google meansGoogle shows this instead of a snippet when it indexed a URL it was never allowed to crawl. The cause, what URL Inspection says, and the two fixes.
Robots noarchive: Google ignores it, Bing uses it for AIGoogle retired the noarchive rule with the cached link. Bing turned it into a generative AI opt-out. What the tag does on your site in 2026.
Blocked by robots.txt in Search Console: fix it or leave it alone?Blocked by robots.txt in Search Console is often your own rule working correctly. How I decide which blocked URLs matter and which to leave.
Excluded by 'noindex' tag in Search Console: where the tag is hidingExcluded by noindex tag, now shown as URL marked noindex, means Google obeyed a directive you may not know you set. The four places it hides.
X-Robots-Tag noindex: how to keep PDFs and images out of GoogleHow the X-Robots-Tag HTTP header applies noindex to PDFs, images and other non-HTML files, why robots.txt cannot, and how to test it.
Indexed, though blocked by robots.txt: what it meansGoogle indexed a URL it was never allowed to fetch. Why robots.txt cannot keep a page out of Google, and the two fixes depending on what you wanted.
Does Google use lastmod in a sitemap?Google ignores priority and changefreq outright, and uses lastmod only when it is consistently and verifiably accurate. What that condition actually means.
Googlebot does not parse JSON: what that changesGary Illyes confirmed Google's crawlers only download JSON, JSON-LD included. Parsing happens at indexing. What that changes when your schema breaks.
Failed: robots.txt unreachable in Search ConsoleWhat failed: robots.txt unreachable means, the 12 hour and 30 day clock Google runs behind it, and the fix order that gets crawling back the same day.
How many URLs can a sitemap have?A sitemap can hold 50,000 URLs or 50MB uncompressed, whichever comes first. The sitemap index math, and why I split mine far earlier than that.
Noindex vs robots.txt: which one removes a pageNoindex removes a page from Google, robots.txt only stops crawling, and combining them breaks both. Which one to use for each removal job.Nothing in that category yet.
Or browse every post, and the evergreen answers in the guides.