SEO & SERP 14 min read

Site Crawler Tools: Map Website Structure and Find Errors

What a site crawler tool shows you and where it stops: desktop and cloud crawlers compared, with August 2026 prices, URL limits, and the storage setting that decides whether a large crawl finishes.

ST
Scraping.Pro Team
Data collection for business needs
Published: 11 August 2025

Google's own documentation names two thresholds for crawl budget: a million or more unique pages changing about weekly, or ten thousand or more changing daily. Below those lines the topic does not apply to you. Search Console draws it lower still, telling anyone with fewer than a thousand pages not to open the Crawl Stats report at all. Almost nobody running a crawler on their own site is near either number.

So the usual reason given for crawling is the wrong one. The real reason is duller: nobody in the building knows how many URLs the site has. A site crawler tool settles that by walking internal links the way a search engine bot does, recording every URL it meets, the response code, the nesting depth, and the links pointing in and out. What comes back is a census, and it is almost always bigger than the page count in the CMS.

This piece is part of our series on web scraping for SEO. Prices and versions below were read off each vendor's own pricing or release page on 13 August 2026.

What you are actually looking for

Google's guidance on large sites, last updated 22 July 2026, calls its thresholds "a rough estimate" and "not exact thresholds". The third case it lists is the one that catches ordinary sites: a large share of URLs sitting in Search Console under Discovered - currently not indexed. That is not a bandwidth problem. That is Google looking at your URLs and deciding they are not worth fetching. A structure crawl gives you the evidence to change its mind.

  • How many real pages exist, and which template made them. A shop with 4,000 products can present 40,000 crawlable URLs once filters, sort orders, pagination and session parameters combine, and every one is a real fetch for a real bot. Sorting by URL pattern rather than by section finds the generator in a minute.
  • Which URLs return 404 or 5xx, and where redirect chains end up after two or three hops.
  • Where duplicates formed, by URL and by rendered content.
  • How deep the pages that earn money sit, counted in clicks from the homepage rather than slashes.
  • Orphan pages. Pages that exist, get traffic, and have no internal link pointing at them. The hardest thing on this list to find, and the section on scale below explains why a default crawl misses them.

What the crawler records

For each URL the tool stores the response code, the address, the depth, incoming and outgoing internal links, and the page metadata: title, description, headings, canonical, meta robots, hreflang. Crawling the structure and scraping title and header tags therefore happen in one pass, which is the whole argument for a crawler over a one-off script. When the metadata table is the point rather than a side effect, the walkthrough is in bulk-checking title tags and headings.

Modern crawlers also draw the result. A force-directed crawl graph, a directory tree, a depth histogram: not decoration at ten thousand URLs, because a bloated section or a dead-end branch shows up in the picture and hides in the spreadsheet.

The crawl is not what Google sees

An earlier version of this article said the output "shows you the site the way Google sees it". That is wrong, and it is the most repeated claim on this topic.

Google processes pages in three phases: crawling, rendering, indexing. Every page returning 200 goes into a render queue, and Google's own wording is that it "may stay on this queue for a few seconds, but it can take longer than that". Rendering runs on an evergreen Chromium, and the rendered HTML, not the source HTML, is what gets indexed. Your desktop crawler does none of this on Google's schedule, from Google's IP ranges, with Google's cached copy of your robots.txt.

A crawl is a hypothesis about your site. Logs are the record of what happened. Search Console's Crawl Stats report keeps 90 days of Google's actual requests, split by response code, file type and host status, including failed robots.txt fetches and DNS lookups. Server logs cover every bot. Screaming Frog's Log File Analyser is at version 7.1 and costs £99 a year; the free build stops at 1,000 log events.

Desktop crawlers

Screaming Frog SEO Spider is the default and deserves to be. Version 24.3 shipped on 29 June 2026 per the vendor's release history. It follows internal links, collects metadata, response codes and link relationships, visualizes the tree several ways, executes JavaScript, and pulls arbitrary fields by XPath through Custom Extraction. Crawl comparison between two runs is built in, which is what you want after a migration.

Correct one thing you will read elsewhere: the free build is not crippled. On the vendor's pricing page the free and paid feature columns are identical, JavaScript rendering, Custom Extraction, scheduling, crawl comparison and the Analytics and Search Console integrations included. The only difference is the 500 URL ceiling. A license is £199 a year, £169 at twenty or more seats. Version 24.0, released 19 May 2026, added an MCP server so a crawl can be driven from an AI assistant. The release notes warn that it is "not a replacement for an experienced SEO professional".

Sitebulb turns the crawl into prioritized, plain-English audit hints instead of a data dump, and its internal-link visualizations are the best here. Two constraints matter. Sitebulb Desktop runs on Windows or Mac only, so it cannot live on a cheap headless box. And Lite is hard-capped: "With Lite, you are restricted to crawling 10,000 URLs in total." Pro defaults to 500,000 URLs per audit, raisable to 2 million, extra seats £7 a month. Desktop prices sit behind a currency toggle that needs JavaScript, so the only figure the pricing page prints as text is the cloud one: Sitebulb Cloud starts at £95 a month and runs "in excess of £20k/year for big teams".

Netpeak Spider exports the full URL-based structure, shows depth, takes a proxy list and is quick on large catalogs. The 2026 news is commercial: no free tier any more. The product page offers a 3-day trial, "Then $20 per month". Windows and macOS, same as Sitebulb.

Cloud crawlers

Running a website crawler online buys scheduled recrawls, shareable reports, history, and no laptop pinned for six hours.

Ahrefs is the cheapest way in, because Ahrefs Webmaster Tools is free for sites you can verify and now carries the broader name Ahrefs Free. Site Audit there gets 5,000 crawl credits a month per project, and credits burn only on HTML pages returning 200, so broken URLs and images do not eat the budget. Paid plans run $129, $249 and $449 a month for 100,000, 500,000 and 1.5 million credits.

Semrush has renamed and repriced its plans, so older comparisons name tiers that no longer exist. The current ones are SEO at $139, Starter at $199, Pro+ at $299 and Advanced at $549 a month. Site Audit page allowances appear on neither the pricing page nor the feature page.

Lumar, the former DeepCrawl, Oncrawl and JetOctopus are the enterprise end, built for very large sites and log-file analysis. Lumar and Oncrawl are quote-only: both pricing pages lead to a form rather than a number. JetOctopus publishes: its pricing page lists a Standard plan at €383 a month billed annually for 1 million crawl pages, 500,000 of them with JavaScript, and 1 million monthly log lines. Its 7-day trial caps at 10,000 URLs.

Prices and limits, read 13 August 2026

Tool Free tier Paid Crawl ceiling Platforms
Screaming Frog SEO Spider 24.3 500 URLs, all features £199/yr 5M in database mode, ~10M with 16 GB Win, macOS, Linux
Sitebulb Desktop 14-day trial behind a JS toggle Lite 10,000; Pro 500,000, up to 2M Win, macOS only
Sitebulb Cloud no from £95/mo 10M URLs per audit browser
Netpeak Spider 3-day trial only $20/mo not published Win, macOS
Ahrefs Site Audit 5,000 credits/mo per project $129–$449/mo 100k to 1.5M credits/mo browser
Semrush Site Audit limited free account $139–$549/mo not published browser
JetOctopus 7-day trial, 10,000 URLs from €383/mo 1M pages, 1M log lines browser
katana 1.6.1 fully open source, Apache-2.0 none none Win, macOS, Linux

Free tools that will do the job

Screaming Frog's free build is the best free site crawler that exists, provided your site is under 500 URLs. Nothing is switched off. For a brochure site or a documentation tree that is the whole answer.

katana is a different animal: a Go crawler from ProjectDiscovery, Apache-2.0, 17,000 stars, release 1.6.1 on 5 May 2026. No SEO reporting at all, and in exchange no URL ceiling, a headless Chrome mode, endpoint extraction from JavaScript files, a known-files mode that reads robots.txt and sitemap.xml, and JSON output you can pipe. When the deliverable is a list of every URL rather than a report on titles, this is the faster route. Default depth is 3; raise it.

Greenflare is the cautionary one. A free GPL-3.0 SEO crawler for Linux, Mac and Windows, it does titles, canonicals, meta robots, status codes and XPath extraction, which reads like exactly the tool this section should recommend. Its last tagged release is 0.98.1, February 2021. The repository is not archived and nothing is broken, but nothing is moving. Broader coverage of the free end is in free tools for scraper developers.

What breaks at scale

Storage mode decides everything. Screaming Frog stores a crawl in RAM by default and allocates just 2 GB on a 64-bit machine out of the box. Its own guide to crawling large websites puts memory mode at "websites under 500k URLs" and notes that 8 GB of RAM "will generally allow you to crawl a couple of hundred thousand URLs". Switch to database storage on an internal SSD and the numbers change shape: 4 GB gets 2 to 3 million URLs, 8 GB about 5 million, and "a machine with a 500gb SSD and 16gb of RAM, should allow you to crawl up to 10 million URLs approximately". People who report that Screaming Frog cannot handle big sites are running memory mode on a laptop.

Time is the other wall. The default is 5 concurrent threads, deliberately, so as not to flatten the server. At a steady 5 URLs a second, 500,000 URLs is roughly 28 hours. Turn on JavaScript rendering and every page becomes a browser page load, so the same crawl stretches past a working week.

Orphan pages need a second pass, and this is where most people quietly get it wrong. Screaming Frog's orphan pages tutorial is explicit: the three Orphan URLs filters "can only be viewed at the end of a crawl", after Crawl Analysis is run by hand. Connecting Analytics and Search Console is not enough either. The Crawl New URLs Discovered In Google Analytics option and its Search Console twin have to be ticked before the crawl starts. Skip both steps and the orphan count comes back zero, which looks like good news.

API quotas cap the verification step. Search Console's usage limits put the URL Inspection API at 2,000 queries per day and 600 per minute per site. Checking indexation for a 500,000-URL catalog that way would take 250 days.

Sitemaps have their own ceiling. The protocol allows "no more than 50,000 URLs" and 50 MB uncompressed per file, and the same limits apply to an index file. Using sitemaps as a second seed source is covered in crawling a sitemap before scraping.

Where a crawl stops working

Anything behind a login is invisible unless you configure authentication, and account pages should stay out of the crawl anyway.

Faceted navigation is infinite. Three filters with ten values each plus a sort parameter generate more URL combinations than the site has products, and a crawler with no exclusion rules walks them until the disk fills. Exclude by parameter before the first run, not after.

JavaScript-only navigation hides the graph. If links are click handlers rather than <a href>, a non-rendering crawl sees a homepage and stops. Rendering finds the pages and still misses anything needing a scroll or a tab click.

Personalized and geo-varied output makes the crawl a single sample. One IP, one language header, one currency. A site serving different catalogs by country needs one crawl per country. Past that point the job stops being an audit and becomes an extraction problem, which is where a managed extraction service earns its keep.

Crawling someone else's site

Pointing a crawler at a competitor is common, usually to export the category tree of a large catalog. XPath against breadcrumb markup rebuilds their taxonomy, which shows you the sections you never built. Dumping the result into a document store keeps the nesting intact; the trade-offs are in the piece on NoSQL storage. It is a close cousin of keyword research scraping: a rival's structure is a map of query groups. If the question is who ranks rather than what exists, that is SERP scraping, a different job with different tooling.

Three things have changed since this kind of advice was first written down.

robots.txt is a standard now, and it is not a lock. RFC 9309 made the Robots Exclusion Protocol a Standards Track document in September 2022, authored mostly at Google. It says plainly that the rules "are not a form of access authorization" and that the protocol "is not a substitute for valid content security measures". It also fixes what a compliant crawler must do: parse at least 500 KiB of the file, and never reuse a cached copy for more than 24 hours. Ignoring it is not hacking. It is also not free of consequence, because the file is evidence that you were told.

The headline court case went the opposite way to the folklore. hiQ Labs is still cited as the company that beat LinkedIn over scraping public profiles. hiQ won the Computer Fraud and Abuse Act argument at the Ninth Circuit in April 2022, then lost on breach of LinkedIn's user agreement, and the case ended in a December 2022 consent judgment: $500,000 and a permanent injunction. Public and permitted turned out to be different words.

The network layer changed under everyone. On 1 July 2025 Cloudflare switched the default for new domains to blocking AI crawlers unless they pay, and launched Pay Per Crawl as the mechanism. Its figures for the crawl-to-referral gap: against what Google search once returned, OpenAI is 750 times harder to earn a visitor from and Anthropic 30,000 times. A large share of the web now sits behind a bot policy that did not exist two years ago, and a crawler with a default user agent meets a challenge page long before it meets the category tree.

So: throttle hard, identify yourself honestly, read the target's robots.txt and terms first, and stay off pages that need an account. A competitor audit of a few thousand URLs is unremarkable. Continuous monitoring of hundreds of catalogs is a different activity with different obligations, and usually a job for a dedicated extraction pipeline rather than a desktop tool left running overnight.

Running the first crawl

  1. Start at the homepage and let it run unconstrained for two minutes. Then stop and look. Those two minutes tell you whether you have a parameter problem before you waste six hours on one.
  2. Set the scope. Limit to a subfolder, exclude session and filter parameters, cap depth. Decide now whether subdomains are in or out.
  3. Switch storage mode above roughly 200,000 URLs and raise the memory allocation. This is the step people skip and then blame the tool.
  4. Enable JavaScript rendering only if you need it. View source on a category page and search for a product name. If it is in the HTML, rendering costs hours and buys nothing.
  5. Connect Search Console and Analytics, tick the two "crawl new URLs discovered" options, crawl, then run Crawl Analysis. Orphan pages appear only after that last step.
  6. Read the reports in this order: response codes, redirect chains, indexability, duplicate titles and descriptions, crawl depth. Fix 5xx first.
  7. Export the tree, crawl again after the fixes, and compare the two runs. The comparison is the deliverable, not the first crawl.

Getting the fixed pages recrawled is a separate step. IndexNow takes up to 10,000 URLs per POST and is honored by Bing, Yandex, Naver, Seznam, Yep and Amazon. Google is not on that list and has never joined it, whatever the plugin descriptions imply.

How often, and what it is not

Monthly is a sane baseline for an active site, plus an on-demand crawl after any migration or template change. Screaming Frog schedules it, and since 24.0 can email the comparison against the previous run.

A crawler and a scraper are not the same tool wearing different hats, though they overlap. A crawler discovers and walks URLs to map a site; a scraper extracts fields from pages it was pointed at. Most crawlers can extract and most scrapers can follow links, so the honest distinction is intent, and the long version is in crawling versus scraping. Crawling a site you own is routine and safe. On a site you do not own, the crawl budget you spend is somebody else's.