ScrapeHero publishes $199 a month for one site and $8,000 a month for its top tier on the same page. Grepsr's one-off Starter Pack starts at $350. Zyte sells managed pipelines from $500 a month. PromptCloud publishes no rate card at all and explains why in a line worth reading twice: "A pricing table built around round numbers would either underprice complex work or overprice simple work." Four vendors, one product category, a fortyfold spread in the published numbers.
That spread is not a haggling range. It is what happens when one word covers both a weekend script against a stable site and a hundred-source pipeline with a named engineer and a delivery deadline attached.
A managed web scraping service builds, runs and repairs data-collection pipelines on your behalf and hands you structured data on a schedule. You describe the sites and the fields; the provider absorbs crawling, unblocking, parsing, quality checks and delivery. What follows is the buying decision rather than the vocabulary: what the money buys, what the arithmetic looks like before anyone gets on a call, what breaks between one page and a million, how to buy a pilot that can fail, and which contract clauses decide whether year two is quiet.
Every price, plan and limit below was read from the vendor's own pricing page or documentation on 10 August 2026. They move, sometimes monthly. The earlier version of this guide described a round of quote requests without a single number in it, which was the wrong shape for the advice. Enough of this market now publishes rate cards that the numbers can do the arguing.
What the money actually buys
The build is the cheap part, and it is the part every quote itemizes. Everything expensive happens afterwards.
- Extractor build. Mapping each site to your fields. Days of work per site, once.
- Getting the page at all. Rotating residential proxies, header and TLS fingerprint management, a real browser engine for dynamic content, and CAPTCHA solving where a challenge is unavoidable. The price of this layer moves, because someone on the other side is paid to break it.
- Maintenance. A site redesigns, an A/B test ships a second layout to 10% of visitors, a field moves inside a JSON blob, a login flow gains a step. Your extractor does not error; it returns rows with an empty column. Maintenance is the real recurring cost of web data and the main reason teams hand the work to a managed pipeline. It is also the line item that separates a $199 quote from a $1,500 one, and it is rarely spelled out in either.
- Quality assurance. Deduplication, schema validation, type and range checks, change detection, alerting when a field's fill rate drops. Zyte scores delivery on seven named dimensions: accuracy, completeness, validity, consistency, timeliness, coverage and freshness. Whether or not you buy from them, those seven hold any provider to something.
- Delivery. CSV, JSON or Parquet to a bucket, a warehouse load, a webhook, or an API you poll.
Nobody publishes a breakage rate. No vendor discloses how often its own extractors need repair, and the "sites change every N weeks" figures in circulation trace back to marketing rather than a published measurement. Your own is measurable and takes one column: log the date each extractor last needed a code change. After a quarter you have a number for your sites, and that is the only number that matters when you weigh a maintenance retainer against an engineer's time.
Three shapes of the market, and their August 2026 rate cards
There is no single service shape behind the phrase. The market splits into three, and picking the wrong shape costs more than picking the wrong vendor inside the right one.
- Fully managed, done-for-you. You specify requirements and receive data. This is the data as a service end of the market, and it is right when web data is an input to your business rather than your product, when targets are numerous or hostile, and when you need somebody accountable for a feed arriving at 06:00.
- Self-serve scraping APIs. You write the extraction logic; the vendor sells proxy rotation, browser rendering and unblocking behind one endpoint. Right when you have engineers and want infrastructure instead of people.
- DIY in-house. Open-source stack, your servers, your proxies, your on-call rotation. Cheapest in licence fees, most expensive in attention.
Plenty of teams run two of the three: managed for the messy high-value sources, DIY for a couple of stable ones. That is a sensible split, not a failure to commit.
Published managed prices
| Provider | Published entry price | Notes |
|---|---|---|
| Zyte Data | from $500/month | sample data in one day, production pipeline in ten; standard and custom plans |
| ScrapeHero Business | $199/month per site | 1–5K pages per site, up to 2 sites; setup fee charged on top |
| ScrapeHero On Demand | from $550 per site, per refresh | setup included, no subscription |
| ScrapeHero Enterprise | from $1,500/month | up to 4 sites, unlimited pages, then $650+ per million pages |
| ScrapeHero Enterprise Premium | from $8,000/month | any number of sites, $500+ per million pages, dedicated resources |
| Octoparse managed data service | from $599 | quote-based beyond that |
| Grepsr Starter Pack | from $350 | one-off extraction, priced on records rather than requests |
| Bright Data managed data acquisition | from $1,500/month | listed on its own scraper pricing page |
| PromptCloud | nothing published | quote follows a 30–45 minute scoping call, delivered in 24–48 business hours |
Two things to read out of that table. First, the managed floor is around $200–$600 a month, not the five figures that "enterprise data services" implies. Second, seven of the nine entries open with the word "from", and it is doing real work: source count, anti-bot difficulty, refresh frequency and schema depth all move the number. A quote that arrives without asking about those four was not scoped.
Self-serve API prices
| Service | Entry plan | What it includes | Unit economics |
|---|---|---|---|
| ZenRows | $16 Build | 45,000 credits, 20 concurrent | free tier of 5,000/month; billed on successful requests only |
| Firecrawl | $16 Hobby, billed yearly | 5,000 pages, 5 concurrent | $5 per 1,000 credits pay as you go; markdown output aimed at models |
| ScrapingBee | $49 Freelance | 250,000 credits, 50 concurrent | JavaScript rendering and premium proxies start at the $249 Business plan |
| ScraperAPI | $49 Hobby | 100,000 credits, 20 threads | credits are multiplied by target: Amazon ×5, Google ×25, LinkedIn ×30 |
| Oxylabs Web Scraper API | $49 Micro | up to 98,000 results | from $0.50 per 1,000 results, $0.25 at custom volume |
| Bright Data Web Scraper API | pay as you go | 5,000 free records a month | $1.50 per 1,000 records, $1.30 above $499/month |
| Zyte API | pay as you go, $5 credit | five difficulty tiers, assigned automatically | $0.06–$1.27 per 1,000 raw HTTP, $0.48–$16.08 per 1,000 browser-rendered |
| Apify | $29 Starter | $29 of platform credits, $0.20 per compute unit | store actors priced per event or per result |
| Octoparse | $69 Standard | 100 tasks, 3 concurrent cloud runs | no-code builder; free tier caps exports at 50,000 rows/month |
| Decodo | pay as you go | formerly Smartproxy | Web Scraping API from $0.09 per 1,000; residential proxies $2.75–$4.00/GB |
The headline plan price is the least informative number on that table. Two columns matter more.
Credits are not pages. ScraperAPI charges one credit for an ordinary page, five for Amazon, twenty-five for Google or Bing, thirty for LinkedIn, and adds ten to any page behind Cloudflare, DataDome or PerimeterX. A million credits is a million pages on a plain site and 33,000 on LinkedIn. Zyte does the same thing more openly, sorting every target into one of five tiers automatically and charging Advanced browser-rendered requests at $16.08 per 1,000 against $0.48 for Simple. A factor of thirty-three, inside one product.
Concurrency is your throughput ceiling. Fifty concurrent requests at four seconds each is 45,000 pages an hour. Twenty concurrent is 18,000. If your job is a nightly refresh of 400,000 pages inside a six-hour window, the $16 plan cannot do it at any credit balance.
The arithmetic before anyone gets on a call
Take a concrete job: 250,000 product pages a month across three retail sites, one of them behind commercial bot management, daily refresh, output as JSON to a bucket.
- ScraperAPI: 150,000 ordinary pages plus 100,000 protected pages at 11 credits each is 1.25M credits a month. The $149 Startup plan covers a million, so this lands on Business at $299.
- ZenRows: 250,000 successful requests fits the $165 Growth plan's 1.2M credits with room to spare, assuming one credit per request.
- Bright Data Web Scraper API: 250,000 records at $1.50 per 1,000 is $375 pay as you go.
- Oxylabs: 250,000 results sits just above the $99 Starter allowance of 220,000.
- Zyte API, browser-rendered, Moderate tier: $1.92–$4.02 per 1,000 puts the same job between $480 and $1,005 depending on monthly commitment.
- Managed: Zyte Data from $500 a month, ScrapeHero Enterprise from $1,500.
The API spread for one job is roughly $165 to $1,005 before anybody writes a parser. The managed floor sits inside that spread rather than an order of magnitude above it, and that single fact reframes the decision. You are not choosing between cheap and expensive. You are choosing whether the difference buys an engineer's attention or costs you one.
Put your own number on that attention. At a fully loaded $60 an hour, ten hours a week of maintenance and firefighting is about $2,580 a month. An API stack that saves $700 on invoices while eating those ten hours has lost by more than three to one. Pick a different rate if yours is different. The shape holds until maintenance drops under two hours a week, and then it flips hard.
One more line in the arithmetic: what counts as billable. Zyte's documentation is unambiguous, and it is the sentence to look for in any vendor's terms: "You are only charged for successful responses. Rate-limiting and unsuccessful responses are free." ZenRows also bills successes only, with a wrinkle, since 404 and 410 responses "count as usable results" and are charged. On a catalogue with heavy churn that difference is real money, and it never appears in a comparison table.
What breaks between one page and a million
A script that works once tells you almost nothing about a pipeline that runs every night for two years. Six things go wrong at volume, and only the first is obvious.
A 200 response is not a page. Bot management increasingly returns HTTP 200 with a challenge or an empty shell rather than a 403. Success measured by status code drifts upward while your data quietly empties out. Measure success as "row passed schema validation with the required fields populated" or you are measuring the wrong thing.
Field-level completeness moves independently of row counts. A run can return the same 250,000 rows as yesterday with price missing on a fifth of them, because price is injected by a script that only fires after a check your client failed. Row counts look flat. Revenue decisions do not.
Schema drift is silent by design. Nobody emails you when a site renames a field. The failure mode of a well-built extractor is not a crash, it is a null, and nulls average out into dashboards.
Retries are a time budget, not a reliability feature. Two seconds of extra latency on 4% of a 250,000-page run adds five and a half hours of fetch time, which at twenty concurrent workers is about seventeen minutes on the end of your window. A bad night at the target site turns that into a missed deadline, and the SLA that matters is the one about lateness rather than uptime.
Freshness is not cadence. A daily feed that lands at 06:00 carrying a crawl that began at 20:00 is ten hours old on the rows collected first. Against competitors who reprice intraday, that is the difference between a decision and a guess. Ask when collection starts, not when the file appears.
The last mile fails too. Bucket credentials expire, warehouse schemas get migrated, a webhook returns 500 for six hours and nobody retries. Ask what happens to a failed delivery, how long data is retained on the provider's side, and whether you can re-request a window.
At a million pages a month the cheap failure is a crash and the expensive one is a plausible-looking file. Build your acceptance tests against the second.
Buy a pilot that is allowed to fail
A demo on an easy site tells you nothing. A paid pilot on your hardest target tells you almost everything, and the fee is worth paying because it changes who owes whom.
Design it so it can produce a negative result:
- Pick the worst site, not the representative one. Everyone can do the representative one.
- Build a gold set by hand. Two to four hundred rows collected manually from live pages, with the fields you actually use. A day of tedious work, and the only reference you will ever have. Do not share it with the provider.
- Score per field, not per row. Fill rate and value accuracy for each field separately. A 96% row-level match can hide a price that is wrong 20% of the time on discounted items.
- Include a holdout. Reserve fifty URLs the provider never sees in the brief. Extractors tuned to a supplied list behave differently from extractors that have to generalize.
- Re-run it a week later, unannounced. The first delivery is somebody's hand-tuned best effort. The second is the product.
Then write down what would make you say no, before the data arrives. Acceptance criteria decided after seeing the results are not criteria.
The contract decides year two
Price is the part everyone negotiates and the part that matters least after the first invoice. These clauses hurt later.
An SLA needs four numbers, not one. Uptime, delivery punctuality, field-level completeness, and time to restore after a target site breaks. ScraperAPI advertises a 99.9% uptime guarantee on every plan, a fine example of the genre: the number is public, and what it pays out when missed is not. Ask for the remedy in writing. A service credit against next month's fee is standard and modest; nothing at all is also common.
Who owns the extractors? If the provider built scrapers to your spec, the contract should say whether you leave with the code, a running instance, or nothing.
Who owns the data, and can they resell it? Ask whether your fields, your target list and your derived dataset can appear in the provider's marketplace or in another customer's feed. A target list is commercially sensitive on its own.
Sub-processors. Managed providers subcontract proxies, solving and sometimes parsing. Under GDPR you are entitled to know who they are and to object to changes.
Exit. Notice period, final data dump, format, and whether history stays available for a window after termination. A pipeline you cannot leave is a price increase waiting for a quiet quarter.
Indemnity. If a target site sends a legal letter, whose problem is it? The answer is in the terms rather than in the sales call, and it is worth reading before you sign instead of after the letter.
The compliance you are buying with the data
Legality depends on the data, the source and your use, and no provider's marketing page settles it. A vendor who waves away every legal question is telling you where the liability will land. Our longer treatment is in is web scraping legal; what follows is the part that changes how you choose a supplier.
"We only scrape publicly available data" is not an answer. GDPR does not care that personal data was public. Somebody still needs a lawful basis, and "it was on the open web" is not one. If your use case is model training, the EDPB's Opinion 28/2024, adopted on 18 December 2024, is where the regulator set out its position on personal data inside AI models, and it is worth reading before the contract rather than after.
If it is personal data, your vendor is your processor and you are the controller. That makes GDPR Article 28 your checklist rather than theirs. The contract has to specify subject matter, duration, nature and purpose, the types of personal data and the categories of data subjects. The processor must act "only on documented instructions from the controller", must not add sub-processors without your authorization, must delete or return the data at the end, and must supply "all information necessary to demonstrate compliance". A provider with no data processing agreement leaves that gap with you.
Check the data broker registries. California's Delete Act obliges data brokers to register with the state privacy regulator each January. Since 1 January 2026, California residents can file one deletion request through DROP that reaches every registered broker at once, and brokers had to begin processing those requests on 1 August 2026, nine days before this sentence was written. A provider that sells personal data about people it has no relationship with is in scope, and the registry is public. Searching it for a supplier's legal entity name takes a minute and occasionally produces a surprise.
Ask what they refuse to do. The useful answers are specific: no data behind a login where the terms forbid it, no circumvention of technical access controls, robots.txt honored on these crawlers and deliberately not on those. RFC 9309 standardized the Robots Exclusion Protocol in September 2022 and states plainly that "these rules are not a form of access authorization". Respecting it is a policy choice a provider makes, and you buy that choice along with the data.
One claim to stop repeating. The line that scraping public data was settled as legal by hiQ Labs v. LinkedIn is half of a story. The Ninth Circuit rulings were about a preliminary injunction, the 2019 opinion was vacated by the Supreme Court in 2021 and sent back after Van Buren, and on the merits hiQ lost: in November 2022 the district court found it had breached LinkedIn's user agreement, and the matter ended in settlement rather than the precedent it gets cited for. The durable lesson is narrower and more useful. Computer-fraud statutes are weak against logged-out collection, so platforms sue on contract, on copyright and on circumvention instead, which is a different question from whether scraping is legal in the abstract.
Industry certification is a weak signal, not a strong one. The Ethical Web Data Collection Initiative, run with the i2Coalition, publishes a member list that currently includes Coresignal, Zyte, Proxyrack, Oxylabs, Decodo, Evomi, Rayobyte, Byteful and ProxyEmpire, and it was still issuing statements in July 2026. Membership means a company signed up to a framework of four principles. It is not an audit, and it is not SOC 2 or ISO 27001, which are. If a provider claims either, ask for the report or the certificate number and check it with the certifying body.
Check whether the data is already for sale
Before you commission a crawl, price the licensed route. Sometimes it wins outright, and sometimes the comparison tells you what you are paying for.
Google's own Places API charges $5 per 1,000 Place Details calls on the Essentials tier after 10,000 free a month, $17 per 1,000 on Pro, and $32 per 1,000 for Pro Text Search. Apify's Google Maps scraper, with 552,000 users and a last modification two days before this was written, starts at $1.50 per 1,000 scraped places. The gap is not a technology gap. One route bundles fields that would take several billable Google calls to assemble and carries no licence to redistribute; the other comes with terms, support and a contract. Knowing both numbers is what makes Google Maps data a decision instead of a default.
Ready-made datasets sit in the same category. Bright Data's marketplace starts at $0.0025 per record with a $250 minimum order and discounts of up to 80% for a monthly refresh subscription. For a standard entity type on a well-covered domain, a dataset you can sample before buying beats a custom pipeline on price and on time to first row. Custom collection earns its cost when the fields, the target list or the cadence are yours specifically, which is the case for competitor price monitoring and review monitoring, where the point is watching your own set of listings rather than a slice of a market.
When building it in-house still wins
DIY is the right answer more often than vendors admit. Build in-house when you have a small number of stable sources, engineers with genuine spare capacity, and a policy reason for keeping collection internal. Two or three static sites with no bot management are a weekend of work and an hour a month afterwards. Paying a retainer for that is paying for a problem you do not have.
The costs are simply less visible than a monthly invoice. Residential bandwidth is billed by the gigabyte: at $3.00 per GB, a 500 KB page costs $0.0015 to fetch once, so 250,000 of them is 125 GB and $375 a month before a single retry, and pulling images alongside the HTML multiplies it. Add solver fees, a browser farm if the targets need one, monitoring, and the person who gets paged when a redesign lands on a Friday.
The build side has genuinely got cheaper. Model-driven extractors return structured output without a hand-written selector for every field, which removes much of the brittleness that made maintenance a weekly chore. Access has not got cheaper. Rendering JavaScript stopped being a differentiator years ago, and getting a real response out of a well-defended site is still the expensive part, which is what you rent from an API or a managed extraction service.
Outsource when web data is an input to your business rather than your product, when targets are numerous or heavily defended, when somebody has to be accountable for a feed landing on time, and when the honest in-house cost exceeds the fee. The last clause is the one people fudge.
Questions that separate the quotes
Ask these, and listen for specifics rather than reassurance.
- What is your success rate on this URL list, and how do you define success?
- What happened the last time one of your targets changed layout? Name the site, the detection method and the hours to restore. A provider without a specific incident to describe has no monitoring or no memory.
- Which sub-processors touch our data, and how are we told when that list changes?
- What is in the SLA besides uptime, and what does a miss pay?
- Do we get the extractors if we leave?
- What will you refuse to collect, and why?
Four of those six are about failure. The ratio is deliberate.
The decision, in one line
Choosing a managed web scraping service is a total-cost-of-ownership decision made against your own maintenance load. The managed floor is around $500 a month, the self-serve floor is around $16, and the distance between them is measured in engineer-hours you either have or do not. Weigh data quality, compliance posture and maintenance response, and price them last. A one-off pull and a maintained accountable feed are different purchases sharing one name, and the data as a service shape only pays for itself on the second.