Harvest robots.txt, recurse gzipped sitemaps, pull 14 well-known files plus ads.txt, classify disclosed paths into 13 categories with uniqueness scoring, and auto-generate probe variants. Batch up to 3 domains.
Robots & Sitemap Harvester fetches a target’s robots.txt and every linked sitemap, plus the public discovery files at /.well-known/ and the ad-network manifests at /ads.txt and /app-ads.txt. Ironically, the very robots file meant to hide directories from crawlers often points testers straight at admin panels, staging areas and hidden endpoints.
The tool parses robots.txt per user-agent so a path blocked only for Googlebot still surfaces, recurses sitemap index files into their child sitemaps with gzip support, extracts <lastmod> for freshness sorting, and groups thousands of sitemap URLs into endpoint patterns. The disclosed paths are classified into 13 high-signal categories and scored by uniqueness so common patterns (like /wp-admin/) sink below site-specific finds.
robots.txt, sitemaps, /.well-known/ files and ads.txt are public files intended to be read by any client. Acting on what you find still requires authorisation for the target.
No. The tool only reads the public discovery files and classifies what they reveal. The flagged paths are unverified leads. Use the Copy probe list / Copy probe variants buttons and HTTP ProbeMaster to check which are actually live.
Common disclosed paths like `/wp-admin/`, `/api/` and `/cdn-cgi/` are present on most sites and add noise to the report. Paths matching a curated common-list are flagged `common` and sorted below site-specific finds, so the gold floats to the top.
When a sitemap entry includes a `