Tools & Reviews 30 min read

Web Scraping Tools: Overview of Ready-Made Software

Ready-made web scraping tools checked against vendor pricing pages on 13 August 2026: desktop crawlers, spreadsheet formulas, no-code clouds, scraping APIs and extensions, with derived cost per 1,000 pages from $0.45 to $36.75.

ST
Scraping.Pro Team
Data collection for business needs
Published: 14 June 2025

The Apify Store listed 59,444 ready-made scrapers on the morning of 13 August 2026. Three point-and-click extractors in the Chrome Web Store carry more than two million installs between them. Finding a tool has not been the problem for years. The problem is that a thousand JavaScript-rendered pages behind anti-bot protection cost $0.75 through one service and $36.75 through another on comparable entry plans, and the expensive one is not forty-nine times better.

This article is about ready-made web scraping software and services: what they cost, what they actually do, and where each one stops working. No scraper code, no libraries. If you would rather build your own, start with what programming language you should write a scraper in and the survey of web scraping libraries.

Every price, version and date below was read from the vendor's own pricing page, documentation, changelog or store listing on 13 August 2026. Where a vendor publishes nothing checkable, this article says so instead of filling the gap.


The unit you pay in

Most roundups compare features. Features are the least useful variable here, because nearly every tool in every category can pull a title, a price and a link out of a page. What separates them is the unit you are billed in, and how many of those units a real job burns.

There are four units in this market.

  • A license. Pay once or once a year, then run as hard as your own machine and your own IP address can stand. Screaming Frog, Datacol, A-Parser, ZennoPoster, ScrapeBox, GSA.
  • A row. No-code clouds meter the records they hand you, which is the most honest unit and usually the most expensive one.
  • A credit. Scraping APIs meter requests, weighted by how hard the request was: plain fetch, JavaScript rendering, residential proxy, full anti-bot bypass.
  • Nothing. Spreadsheet formulas and browser extensions cost zero and hit a wall early.

Converted to one comparable number, the spread across the metered products looks like this. Costs are USD per 1,000 pages.

Product and mode Entry tier Best published tier
Zyte API, browser rendering $0.48–$16.08 $0.48–$7.68
Scrapfly, anti-bot mode, datacenter IPs $0.75 $0.45
ScrapingBee, JavaScript rendering $0.98 $0.37
Bright Data Web Scraper API $1.50 $1.30
ZenRows, JavaScript rendering $1.78 $0.46
Firecrawl, one page one credit $3.20 $0.60
Scrapfly, anti-bot plus residential $3.75 $2.27
ScraperAPI, JavaScript rendering $4.90 $0.92
ZenRows, JavaScript plus premium proxy $8.89 $2.28
Web Scraper cloud, per URL credit $10.00 $5.00
ScrapingBee, stealth proxy $14.70 $5.62
Browse.ai, per credit $24.00 $13.80
ScraperAPI, ultra-premium plus rendering $36.75 $6.89

How these were derived. List price divided by the credits the plan includes, multiplied by the credit weight the vendor publishes for that mode. ScrapingBee charges 1 credit for a plain fetch, 5 with rendering, 25 with premium proxy and rendering, 75 for stealth. ScraperAPI charges 10 for render=true, 25 for premium plus render, 75 for ultra_premium plus render. Scrapfly charges 5 for anti-scraping protection or rendering and 25 for residential. ZenRows publishes a fixed table of 1, 5, 10 and 25 and states that the weights "are fixed and never change" across every plan. Zyte publishes dollar rates directly rather than credits, and its range covers site difficulty from simple to advanced. Entry tier means the cheapest paid plan, best published tier means the largest plan with a number attached to it. Enterprise quotes are not in here because they are not published.

Three things fall out of that table.

The cheap end and the expensive end are the same request. Scrapfly's anti-bot mode and ScraperAPI's ultra-premium mode both promise a rendered page from a defended site. One costs $0.75 per thousand and the other $36.75 on entry plans. Part of the gap is real infrastructure, and part of it is that credit weights are a pricing lever with no external audit.

The no-code clouds sit at the top of the range. Browse.ai at $24.00 per 1,000 credits on its monthly Personal plan is roughly thirty times Scrapfly's entry rate for a comparable page. You are paying for the point-and-click builder, the scheduler, the change alerts and the fact that nobody on your team has to touch an API key. That can be worth thirty times. It is not free, and no comparison article seems to say the number out loud.

Failure billing changes the ranking. Scrapfly does not charge for failed requests. ZenRows charges only for successful ones and counts 404 and 410 as usable. ScraperAPI bills 200 and 404 only. ScrapingBee's auto mode costs nothing if every configuration fails. Bright Data bills on success. On a target with a 70% success rate, a vendor that bills every attempt is 43% more expensive than its rate card suggests, and one that bills only on success is not.

Octoparse is the odd one out and worth noting for it: its Standard plan is $69 a month with unlimited data export and 100 tasks, metered on concurrency rather than rows. If you can keep three cloud processes saturated, the per-row cost falls through the floor. If you run two tasks a week, you are paying $69 a month for two tasks.


How the tool market breaks down

Six groups, and they solve genuinely different problems.

  1. Desktop programs installed on your machine. Maximum control, your IP address, your electricity bill, no per-row charge. Screaming Frog, Datacol, A-Parser.
  2. Scraping inside a spreadsheet. Excel Power Query and Google Sheets formulas. Free, instant, and limited in ways that are easy to hit and hard to see coming.
  3. Cloud no-code platforms. Point at elements in a browser preview, the service infers the extraction logic and runs it on a schedule. Octoparse, ParseHub, Web Scraper, Browse.ai.
  4. Automation and SEO harvesters. Scraping is one feature beside browser automation, account creation and link building. ZennoPoster, GSA, ScrapeBox.
  5. APIs and infrastructure. You send a URL and get content back; proxies, rendering and anti-bot handling are the product. Apify, Bright Data, ScrapingBee, ScraperAPI, Zyte, Scrapfly, ZenRows, Firecrawl.
  6. Browser extensions. Grab what is on the screen right now, in two clicks, for nothing.

Three questions sort you into one of them: how much technology you are willing to learn, how many pages you need, and how hard the target fights back. The third question dominates the other two, and it is the one people estimate worst.


Desktop programs

Screaming Frog SEO Spider

Built as a crawler for technical SEO audits. It walks a site the way a search bot does and collects links, response codes, titles, meta tags, redirects and duplicates. For scraping, the part that matters is Custom Extraction, which pulls arbitrary values out of the HTML by XPath, CSS selector or regular expression.

Point it at a list of URLs and it will collect prices, SKUs, product names or any markup element across all of them in one pass. It renders pages through an integrated Chromium engine, so React, Angular and Vue sites are in scope. It pulls in Google Analytics, Search Console and PageSpeed Insights data, and link metrics from Ahrefs, Majestic and Moz.

A correction to the earlier version of this article. We previously quoted the license at around £259 per user per year. The pricing page reads £199 per year on 13 August 2026, with volume discounts at £189 for 5 to 9 seats, £179 for 10 to 19, and £169 above 20. The FAQ states the price has been raised twice since the 2010 launch. Licences are per individual user and cannot be shared.

What the free version gives you is the crawl itself, capped at 500 URLs. The license-only list, from the vendor's own documentation, is: configuration, saving and opening crawls, custom search, custom extraction, JavaScript rendering, Google Analytics integration, Search Console integration, PageSpeed Insights integration and scheduling. Custom extraction being paid-only is the load-bearing detail for anyone here for scraping rather than SEO. The free tier will crawl your 500 URLs and hand you titles and status codes, and nothing you actually wanted.

The release history shows 24.3 on 29 June 2026, following 24.2 on 22 June and 24.1 on 8 June, all logged as bug fixes. Builds ship for Windows, macOS on both Apple Silicon and Intel, and Linux. The product page also advertises something the older write-ups do not mention: the crawler can fire custom AI prompts against OpenAI, Gemini, Ollama and Anthropic models while crawling, and run arbitrary JavaScript snippets in the page the way you would in the Chrome console. A desktop SEO auditor with an LLM call in the crawl loop is a different product from the one most reviews describe.

Who it is for: anyone extracting structured data from sites with predictable markup, especially where the same crawl also has to answer SEO questions. It is not a replacement for a real scraper on a large, well-defended target, and the advanced configuration has a genuine learning curve.

Datacol

A general-purpose visual scraper that has been sold for well over a decade. You point it at a site, pick elements through a wizard, and set up navigation and pagination without writing code. Underneath it is XPath and extraction rules with a builder on top. It ships more than 80 preconfigured parsers for specific platforms: online stores, classified-ad boards, marketplaces, social networks, Google Maps and keyword-driven sources, search results and news. Output goes to CSV, Excel, XML, JSON or a database, and posts directly into WordPress, OpenCart and other CMSes. Plugins add spinning, translation and phone-number recognition.

Two things need correcting. The earlier version of this article quoted a license at about $89. The vendor's international site, web-data-extractor.net, advertises editions from $18 with a demo build available, and carries a 2011 to 2026 copyright. The Russian-language domain the product is better known by would not answer an automated request at all on 13 August 2026, so the price ladder above the entry edition could not be read from outside. Treat $18 as the floor, not the figure you will pay for the configuration you need.

Who it is for: people willing to learn the configuration language and some regular expressions to automate content population or repetitive collection. Non-standard targets often mean buying a configuration or a plugin rather than building one.

A-Parser

A multithreaded harvester aimed at volume and SEO work. The site advertises 110+ built-in parsers, builds for Windows, Linux and macOS, and the ability to write your own parsers in JavaScript at close to arbitrary complexity. Unlike a visual tool it exposes an API, so runs can be driven from your own processes.

Pricing is published and one-time, with a separate update subscription: Lite $179, Pro $299, Enterprise $479, plus updates at $49 for three months, $149 for a year or $399 for life. The Enterprise edition is where API management, multi-core task processing and Redis integration live. No version number or dated changelog is published on the pages that answer, so "is this still being developed" cannot be settled from the outside beyond the vendor's claim that parsers are patched within 24 hours when a target changes.

Who it is for: SEO teams and agencies scraping large volumes who want control and can absorb the learning curve. The entry barrier is higher than any no-code tool in this article.


Scraping inside a spreadsheet

Sometimes a separate program is overkill. Both major spreadsheets will fetch a web page for you, and both will do it badly enough that you need to know exactly where the edge is.

Excel and Power Query

Excel's Power Query sits under Get Data, From Web. It pulls tables and structured data straight off a page, refreshes on a click and transforms on the way in. For anything finer, people drop into VBA with MSXML2.XMLHTTP or WinHTTP and parse the response by hand.

The JavaScript claim needs qualifying. It is standard to say Excel cannot handle dynamic sites. Microsoft's own Web connector documentation shows three different functions with three different engines behind them: Web.Contents does a plain HTTP fetch, Web.Page requires Internet Explorer 10, and Web.BrowserContents requires the Microsoft Edge WebView2 runtime. That last one is a Chromium browser, and it does execute page scripts. So Power Query can read a rendered page, on Windows, with WebView2 installed, using the "add table using examples" path. What it cannot do is scale, authenticate flexibly, or survive the target changing shape. Web.Contents also restricts POST requests to anonymous authentication only.

Pros: nothing to install, and the data lands where the analysis already lives. Cons: one machine, one IP, no proxy story, and the moment your logic outgrows the query editor you are writing VBA, which is programming with worse tools than a real scraper in Python would give you.

Google Sheets

Four formulas turn a cell into a small scraper:

  • IMPORTHTML(url, "table"|"list", index) pulls the nth table or list from a page.
  • IMPORTXML(url, xpath_query, locale) extracts by XPath. Google's own example is IMPORTXML("https://en.wikipedia.org/wiki/Moon_landing", "//a/@href").
  • IMPORTDATA(url) fetches a CSV or TSV.
  • IMPORTFEED(url) reads RSS and Atom.

The official documentation is thin on limits, which is the trap. It describes the syntax and the locale argument and says nothing about JavaScript, refresh behavior or throttling. What you will actually hit, in order: no script execution, so anything rendered client-side comes back empty; Google's fetchers arriving from Google IP ranges that a lot of sites block outright, which returns an error that looks like a broken XPath; a hard ceiling of 10 million cells per spreadsheet; and recalculation storms, because every IMPORT formula refreshes when the sheet opens and periodically after that.

Pros: free, shareable, zero install, and genuinely the right answer for a hundred rows off a static page. Cons: everything above. It is a prototyping tool that happens to be permanent.


Automation and SEO harvesters

A separate category: programs designed to be broader than a scraper. Their job is automating browser actions or SEO processes, and data collection rides along. If the task is not "pull a table" but "log in, click through five steps, collect, post it somewhere," this is the shelf.

ZennoPoster

A visual builder for browser automation from ZennoLab, a company whose own footer dates it to 2008. You assemble logic from blocks, or record yourself doing the job in a browser and let the program turn the recording into a project that then runs across many parallel threads. Device and fingerprint emulation, proxy handling, CAPTCHA integration through the same company's CapMonster, a scheduler, database output, and C# or JavaScript inserts for anything the blocks cannot express.

The earlier version of this article called it a paid product with several license editions. The product page now advertises a free LITE edition: one thread, one machine, slow proxy checking, community support. PRO is $39 a month, $33 a month on a six-month term at $197, or $28 a month annually at $333, with unlimited threads and installation on three machines. No version number or release date is published on the pages that answer, so cadence claims cannot be verified from outside.

Scraping here is one scenario among many: search results, competitor prices, contacts, object IDs. The reason to choose it over a scraper is the combination of collect and act. Anything you can do by hand in a browser, you can record and repeat.

Who it is for: people who need multi-step actions on sites, with extraction as one step in a longer process.

GSA

A family of Windows SEO tools from GSA in Rostock. The flagship, GSA Search Engine Ranker, is an automated link builder with a harvester attached: it searches engines by keyword and footprint to find target sites, then places links on them. Its product page states plainly that it "automatically scraps and identifies new backlink targets by searching google for matching sites" and includes a proxy scraper that pulls from public proxy sources. Nearby sit GSA Proxy Scraper, GSA Content Generator with a page scraper inside it, and GSA Keyword Research.

The product page links a changelog entry dated 3 August 2026, so the software is being worked on. It sells as a one-time payment with a lifetime license and free lifetime updates, and no monthly option. The actual figure lives behind a shop that did not answer an automated request, so this article will not quote one.

Who it is for: SEO practitioners collecting large lists of URLs, proxies or keywords as an input to promotion work. For business data such as prices, catalogs and contacts, it is the wrong shape entirely.

ScrapeBox

The classic SEO harvester, and the one product in this article whose changelog you can read at a glance. The site records v2.1.57 on 29 January 2026 and claims 600 releases since 2009. Its work is mass URL collection from search results by keyword and footprint, proxy harvesting and checking, metric lookups, link extraction, and a long tail of small utilities. The site advertises $97, marked down from $197, as a one-time single-PC license with one free transfer per month and free bug fixes.

Six hundred releases since 2009, with the most recent seven months old, is a stronger signal of life than any marketing page in this article. It is still an SEO tool and not a structured-data scraper.

Who it is for: people working with large URL lists and SEO metrics who need fast bulk harvesting.


Cloud no-code platforms

Open a page inside the service, click the elements you want, and it builds the extraction logic. Running, scheduling and storage happen in the cloud. You pay per row, per task or per concurrent process.

Octoparse

The best-known of these, and the one that changed most since this article was written. The desktop builder now ships for Windows 10 64-bit and above and for macOS, at version 10.1.1, released 28 July 2026. The claim that it is Windows-only, which this article carried and which most competing roundups still carry, is out of date.

Pricing also moved. Octoparse lists Free at $0, Standard at $69 a month, Professional at $249 a month and Enterprise on quote, with 16% off for annual billing. The old $75 to $99 range is gone. The Free tier gives 10 tasks, one device, local extraction only and 50,000 rows of export a month. Standard adds cloud extraction, IP rotation, 100 tasks and up to 3 concurrent cloud processes. Professional raises that to 250 tasks and 20 concurrent processes, which is the number that matters if you are trying to finish a large job overnight rather than over a weekend.

Who it is for: non-programmers who need structured data on a schedule and want somebody else to own the infrastructure.

ParseHub

Same click-the-element model, cross-platform desktop client, and a long-standing reputation for handling SPAs, infinite scroll and dynamic content better than its rivals. That reputation is repeated in every comparison including previous versions of this one, and none of them cite a measurement. We could not verify it either.

What we can report is a sign-of-life problem. ParseHub serves a live site and a live help centre, and neither carries any shutdown or maintenance notice. The pricing and download pages render entirely through JavaScript and returned no readable content to an automated fetch on 13 August 2026, so no plan or price is quoted here. The company blog's most recent posts date to 2023. A live site with a three-year-old blog is not a dead product, but it is not a growing one either, and the cross-platform advantage it was known for evaporated when Octoparse shipped a macOS build.

Who it is for: people already running it successfully. For a new project, verify the current pricing and support responsiveness before you build anything on it.

Web Scraper

A Chrome extension first: you build a sitemap, define fields and pagination, and run the job locally or push it to the cloud. The extension is at version 1.111.13, updated 15 July 2026, with 800,000 users and a 4.1 rating across roughly 1,100 reviews. It is free and it is not crippled.

The cloud is where the money is. Web Scraper lists Project at $50 a month billed annually for 5,000 URL credits and 2 concurrent scrapers, Professional at $100 for 20,000 credits and 3 scrapers, and Scale from $200 for unlimited URLs. Paid plans include datacenter proxies and advertise CAPTCHA and bot-protection bypass. Divide it out and Project costs $10 per 1,000 URLs, which puts it in the upper half of the table at the top of this article.

Who it is for: straightforward extractions where the free extension does the job, with the cloud as an upgrade path when scheduling starts to matter.

Browse.ai

A no-code service built around monitoring rather than one-shot collection. It scrapes, and it also watches a page and fires an alert when a price or an availability flag changes. The setup is AI-assisted rather than pure point-and-click.

Browse.ai publishes Free at 50 credits a month across 2 websites, Personal at $48 a month for 2,000 credits or $19 a month annually for 12,000 credits a year, Professional at $87 a month for 5,000 to 30,000 credits or $69 a month annually for 60,000 a year, and Premium from $500 a month annually. The monthly Personal plan works out at $24 per 1,000 credits. Annual Professional gets that down to $13.80. Both are expensive per row and neither is trying to compete on that axis: you are buying change detection and alerting, which none of the raw APIs give you.

Who it is for: teams tracking a bounded set of pages over time, where knowing that something changed is worth more than the volume.


Services with an API and infrastructure

The pain these solve is not the HTTP request. It is surviving Cloudflare, DataDome and their peers, rotating proxies, rendering JavaScript, and doing all of it at a rate that finishes before the data goes stale. You wire them into your own code.

Apify

A platform built around Actors: serverless programs that take JSON input, do a job and return structured output. The store had 59,444 of them on 13 August 2026, and the useful ones present as a form you fill in and run, which keeps you in no-code territory until you need more. When you do, you write your own on Crawlee or Playwright and deploy it the same way.

Pricing is a monthly platform fee plus consumption. Free is $0 with $5 of monthly store credit, 25 concurrent runs and 5 datacenter proxy IPs. Starter is $29, Scale $199, Business $999, all with pay-as-you-go on top. Compute units cost $0.20 on Free and Starter, $0.16 on Scale and $0.13 on Business. The old $39 entry price is gone. Store Actors carry their own per-result or per-event charges on top of compute, which is why Apify does not appear in the per-1,000-pages table: the number depends entirely on which Actor you run.

Who it is for: teams assembling pipelines who want to mix somebody else's scraper with their own code and one billing relationship.

Bright Data

The upper tier by volume and by price. A very large proxy network under a Web Scraper API that returns structured JSON or CSV without you writing extraction logic, handling proxies, CAPTCHAs, rendering and retries internally. Published pricing is $1.50 per 1,000 records pay-as-you-go, $1.30 per 1,000 on the $499 a month Scale plan which includes 384,000 records, and a free tier of 5,000 records a month with no card required. Billing is on successful records.

Alongside the API sit ready-made datasets: the marketplace advertises 350+ datasets across 250+ domains, 21 billion records, from $0.0025 per record with a $250 minimum order. When the data you want already exists as a product, buying it beats collecting it, which is the same reasoning behind a data-as-a-service arrangement rather than a scraper.

Who it is for: projects where the bottleneck is scale and target defences rather than parsing.

ScrapingBee, ScraperAPI, Zyte, Scrapfly and ZenRows

Scraping APIs in their pure form. You send a URL with parameters for country, rendering and retries; the service routes it through proxies, renders if asked and returns the content. The integration is a wrapper around your existing HTTP client, which is why switching between them is usually an afternoon.

They differ on price and on how the price is expressed:

  • ScrapingBee runs $49 Freelance, $99 Startup, $249 Business and $599 Business+, for 250,000 to 8,000,000 credits. Its documentation is unusually explicit about weights, including an auto mode that tries the cheapest working configuration and charges nothing if every configuration fails. Stealth proxy at 75 credits is the most expensive single mode in this article.
  • ScraperAPI runs $49 to $1,975 across six published tiers. It weights by domain as well as by mode: 1 credit flat, 5 for Amazon, 25 for Google or Bing, 30 for LinkedIn, and 10 extra for any request that has to get through Cloudflare, DataDome or PerimeterX. Charges apply to 200 and 404 responses only.
  • Zyte API is the only one that publishes dollars per 1,000 rather than credits: $0.06 to $1.27 for an HTTP response body and $0.48 to $16.08 for browser rendering, pay-as-you-go, with the range tracking site difficulty. Committing $100, $200 or $500 a month lowers both bands. Smart Proxy Manager now sits under a Legacy heading in the documentation with migration guides pointing at Zyte API.
  • Scrapfly runs $30 Discovery, $100 Pro, $250 Startup, $500 Enterprise, with weights of 1 for a plain fetch, 5 for anti-scraping protection or rendering, 25 for residential and 60 for a full-page screenshot. Failed requests are free, which on a hostile target is worth more than the headline rate. On the numbers it publishes, it is the cheapest verified route to a rendered page from a defended site: $0.75 per 1,000 on the $30 plan, $0.45 on the $500 plan. Discovery has no overflow, so exceeding 200,000 credits stops the job rather than billing you; Pro and above bill overflow at $3.50, $2.00 or $1.20 per 10,000 credits depending on tier.
  • ZenRows runs $16 Build, $57 Launch, $165 Growth, $456 Scale, on 1, 5, 10 and 25 credit weights that the documentation says never change between plans. Residential bandwidth meters separately at 25,000 credits per gigabyte, and browser sessions at 5 credits a minute plus bandwidth.

Who it is for: developers who want one endpoint and no proxy operations.

Firecrawl

An API built for feeding language models. It scrapes, crawls and maps sites and returns Markdown by default, so the output goes straight into a prompt or a vector store without a cleanup stage. The v2 API covers scrape, crawl, map, search, interact and parse, the last of which converts local PDF, DOCX, XLSX and HTML into Markdown or structured JSON. There is a hosted Model Context Protocol server at mcp.firecrawl.dev, so a coding agent can call it without an SDK.

Pricing is one credit per page, which makes the arithmetic honest: Free gives 1,000 credits a month at 2 concurrent requests, Hobby $16 for 5,000 pages, Standard $83 for 100,000, Growth $333 for 500,000, Scale $599 for 1,000,000, all billed yearly. That is $3.20 per 1,000 pages at the bottom and $0.60 at the top.

Who it is for: RAG pipelines, agents and anything where the consumer of the data is a model rather than a database.

Browser infrastructure as a category of its own

The newest shelf, and one no version of this article had. Browserbase rents you real browsers by the hour rather than pages by the credit: Free gives 1 browser hour and 3 concurrent sessions, Developer is $20 a month for 100 hours then $0.12 an hour, Startup is $99 for 500 hours then $0.10. Its Fetch endpoint returns Markdown at $4 per 1,000 calls with proxies, $7 with extraction. The SDK on top, Stagehand, advertises 22,000 GitHub stars and 700,000 weekly downloads and lets you address elements in natural language instead of by selector.

Buying browser time rather than pages inverts the economics. A page that needs one fetch is absurdly expensive this way. A workflow that needs a logged-in session, six clicks and a file download is cheap, and no per-page API will do it at all.


Browser extensions

The fastest route to a one-off answer. Instant Data Scraper and Data Miner find the tables and lists on the page you already have open and export them to CSV or Excel. No logic to build, and almost no control.

Both are alive and neither is what the old write-ups describe. Instant Data Scraper is at 1.6.1, updated 16 July 2026, with over 1,000,000 users and a 4.9 rating from 7,600 ratings. Data Miner is at 5.8.102, updated 18 January 2026, around 300,000 users, rated 3.9 from 703 ratings. The rating gap between them is wider than any feature difference either one advertises.

The platform under them changed, and this is the part that gets missed. Chrome's Manifest V2 to V3 migration finished. Google's support timeline records that MV2 extensions were disabled by default across all channels on 31 March 2025, that Chrome 138 in July 2025 removed the ability to turn them back on, and that the ExtensionManifestV2Availability enterprise policy was deleted in Chrome 139. From 31 August 2026, remaining Manifest V2 extensions are removed from the Chrome Web Store altogether. If a scraping extension you rely on has not shipped an update since 2024, check its listing before you plan around it, because in eighteen days it may not have a listing.

MV3 also changed what an extension can do. Background service workers are terminated when idle, and the webRequest blocking API that some extractors used to intercept and rewrite traffic was replaced by the declarative declarativeNetRequest. A scraping extension that used to sit in the background and shape requests cannot work the same way now.


What breaks between one page and ten thousand

Everything in this article works on one page. That is what the demo shows you. The interesting failures start about three orders of magnitude later.

Time compounds. One wasted second per page across a ten-thousand-page crawl is close to three hours. On Octoparse Standard you get 3 concurrent cloud processes; at 5 seconds a page that job takes 4.6 hours. On Professional's 20 processes the same job takes 42 minutes and costs $180 a month more. That is the entire argument for the higher tier, and it is arithmetic rather than a feature.

Silence is the dangerous failure. A scraper that returns zero rows gets noticed. A scraper that returns ten thousand rows with an empty price field looks exactly like a successful run to a scheduler, an export and a dashboard. Every point-and-click tool in this article infers selectors from one example page, and site layouts drift. Check row counts and null rates per field per run, not just exit codes.

Failure billing is not a footnote. At a 70% success rate on a defended target, an API that bills every attempt costs 43% more than its rate card. Over a million pages at $2 per thousand that is $860 you did not budget. Scrapfly and ZenRows do not bill failures; read the small print for the others before you commit a volume tier.

Free tiers are not scale models. Google Sheets tops out at 10 million cells per spreadsheet and recalculates every IMPORT formula when the file opens. A browser extension needs the tab to stay open. A desktop crawler holds the crawl in memory unless you switch it to database storage. None of these degrade gracefully; they stop.

One IP is one IP. Desktop tools and spreadsheet formulas go out from your address. That is fine for a few hundred requests and it is how you get that address rate-limited on the thousandth. This is the single most common reason a tool that worked in testing fails in production, and no amount of license tier fixes it.

Retries multiply invisibly. Allow three retries against a 30% per-attempt failure rate and 10,000 wanted pages become about 14,200 attempts. Budget by attempts, not by target pages.


Where each of these stops working

Every tool here has a wall. Knowing where it is beats knowing the feature list.

Spreadsheets stop at client-side rendering. They also stop when the target blocks Google's or Microsoft's fetcher ranges, which happens more often than the error message suggests. There is no proxy option, no retry policy and no way to stagger requests.

Browser extensions stop at the tab. No scheduling, no headless run, no orchestration, and from 31 August 2026 no Manifest V2 at all.

Desktop programs stop at your IP address and your machine. They are excellent on cooperative targets with predictable markup and poor on anything that scores requests. Some of them are Windows-only, which decides the question before you get to the features.

No-code clouds stop at three places: markup that changes shape between page types, anything behind a login you have to keep alive, and the per-row price once volume passes a few hundred thousand records. They are built for stable, public, well-structured pages, and that describes a shrinking share of the web.

Scraping APIs stop at parsing. Most of them hand you HTML or Markdown, not fields. Bright Data's Web Scraper API, Zyte's automatic extraction and Firecrawl's JSON mode are the exceptions, and they cost more per page than the raw fetch modes for exactly that reason. An API also cannot help with a flow, only with a request; if the data is three authenticated clicks deep, you need browser infrastructure instead.

All of them stop at content behind a login you are not entitled to use. That is a legal question rather than a technical one, and no rate card answers it.


The ground moved in 2026

Three changes since this article was first written matter more than any product update in it.

Bot management stopped asking whether you are human and started asking who you are. Cloudflare split AI traffic into three categories on 1 July 2026: Search, which indexes content to answer questions later; Agent, which acts in real time on a person's behalf; and Training, which feeds a model. Site owners can allow or block each independently, on every plan including Free. From 15 September 2026, new domains onboarding to Cloudflare get defaults that block Training and Agent traffic on ad-monetized pages while allowing Search. Cloudflare's changelog also warns that multi-purpose crawlers combining Search and Training will be caught by the Training block. Verified-bot status no longer implies default-allow.

For anyone choosing a tool, this reframes the question. A better proxy tier does not help when the block is a category-level policy decision. It also means the behavior you saw when you evaluated a service in July may not be the behavior you get in October, on the same target, with the same plan.

The extension platform closed. Manifest V3 is now the only thing Chrome runs, the enterprise escape hatch is gone, and the Web Store purge lands on 31 August 2026. Anything in the two-clicks-and-export category is now living on the platform Google designed, with the request-interception capabilities that some scrapers depended on removed by design.

Tools grew a model-shaped end. Firecrawl returns Markdown because the consumer is a language model, and runs a hosted MCP server so a coding agent can scrape it without an SDK at all. Browserbase sells browser hours to agents and ships a natural-language SDK on top. Screaming Frog, a desktop SEO crawler from 2010, now fires prompts at OpenAI, Gemini, Ollama and Anthropic models during a crawl. None of that existed when the categories in this article were drawn, and it does not fit them: it is a fifth unit of billing, browser-minutes and tokens, sitting beside licenses, rows and credits.

What has not changed is the underlying trade. You still choose between running the machinery yourself and renting it, and renting it still costs between $0.45 and $36.75 per thousand pages depending on how well you read the credit table.


How to choose

A map by situation, with the prices attached.

  • One-off, static site, under a thousand rows. Google Sheets formulas, Power Query, or a browser extension. Cost: zero. Time to first row: minutes.
  • Recurring collection on cooperative targets, on your own machine. Screaming Frog at £199 a year if SEO questions sit nearby, Datacol from $18, A-Parser from $179. You supply the IP address and the patience.
  • No code, reliable, in the cloud. Octoparse at $69 a month if you can saturate the concurrency, Web Scraper at $50 a month for 5,000 URLs, Browse.ai from $19 a month annually when the job is watching pages rather than harvesting them. ParseHub only if you verify its current state first.
  • Actions, not just data: logins, posting, multi-step flows. ZennoPoster, free at one thread and $39 a month for unlimited. Browserbase at $20 a month if the flow is being driven by an agent rather than a recorded script.
  • SEO harvesting at volume: URLs, keywords, proxies. ScrapeBox at $97 one-time, GSA, A-Parser.
  • Large volume, defended targets, embedded in your own code. Scrapfly at $0.75 per 1,000 rendered pages on entry, Zyte from $0.48, ScrapingBee at $0.98, Bright Data at $1.50 per 1,000 structured records. Price the mode you will actually use, not the headline plan.
  • Data for a language model. Firecrawl at $0.60 to $3.20 per 1,000 pages, or Browserbase Fetch at $4 per 1,000 calls when the page needs a real browser.

The guiding rule survives all the price changes: the harder the target defends itself and the larger the volume, the further the answer moves from spreadsheets toward infrastructure with proxies and rendering behind it. Run the counterfactual before you buy the expensive tier, because a slower crawl on a cheaper mode often beats a fast one on stealth proxies at fifteen times the price.

When no off-the-shelf tool covers the job, the remaining options are building it, which starts at choosing a language and picking libraries, or handing the whole problem to somebody who already runs the infrastructure, which is what a managed extraction service is for. Both are cheaper than fighting a tool that was never shaped for the target.

Prices, versions and dates verified against vendor pricing pages, documentation, changelogs and store listings on 13 August 2026. Rate cards in this market change without announcement.