LinkedIn's own API will not return a single member profile unless that member has signed into your app and granted it permission. Bright Data will sell you 905 million LinkedIn profiles at up to $0.0025 each, minimum order $250. Both statements were true on 13 August 2026, and the distance between them is the subject of this article.
Social media scraping is the automated collection of public data from social platforms against criteria you define. A program walks pages or calls an API, gathers followers, likers, commenters, posts, comments and public contact details, and hands back a structured file: profile URLs, usernames or IDs, engagement records, post metrics. That file goes into a CRM, an outreach tool, a dashboard or a model. The mechanics are the cheap part. Everything expensive lives in which route you take to the data and what you are permitted to keep afterwards.
Every price, quota and eligibility rule below was read from the platform's or the vendor's own developer documentation or pricing page on 13 August 2026. Where a number could not be read from the party that owns it, the text says so instead of guessing.
What people actually build with it
Lead generation on B2B networks is the biggest single use and the one that attracts every legal problem in this article. A prospect list assembled from public LinkedIn profiles is job titles, companies, seniority and tenure, enriched with a work email from a verification service, dropped into a sequencer. It works. It is also the exact activity that produced a $500,000 judgment in 2022, a shuttered company in 2025, and a California obligation that took effect twelve days before this sentence was written. All three are below.
Brand monitoring and sentiment. Mentions, comments and hashtags feed social listening. This is the use case where the ground moved most: CrowdTangle, the tool a generation of analysts used for exactly this on Facebook and Instagram, was switched off by Meta on 14 August 2024, and its replacement is open only to nonprofit researchers. Commercial monitoring now runs on scraping, on paid listening suites, or on nothing.
Competitor and audience analysis. Who engages with a rival account, how often, and whether that overlaps with your own audience.
Content analytics. Which posts land, measured across accounts you do not own and therefore cannot see in native analytics.
Influencer vetting. Follower counts and engagement rates before you sign a contract, because the creator's screenshot is not evidence.
Warm-audience filtering. Activity in the last 30 days as a proxy for present interest.
Set against those, one use case deserves a warning rather than a description. Tools that mass-message, auto-follow, mass-like or inflate follower counts are not scraping. They are automation against the platform's own account rules, they get accounts banned rather than throttled, and no amount of proxy hygiene protects a burned account.
Three doors into the same data
Almost every argument about social media scraping collapses once you notice which of three doors the collection went through, because the law, the cost and the failure modes are different for each.
The official API. Sanctioned, rate-limited, metered, and usually narrower than you expect. You get exactly the fields the platform decided to expose, you get them in a stable schema, and you pay per call or per record. Nobody bans you.
Logged-out page scraping. You fetch public pages the way an anonymous visitor does. This is where the strongest legal position sits, and where the technical work is hardest: at volume it needs rotating proxies, fingerprint-consistent browsers and CAPTCHA handling, all of which cost money per page.
Logged-in scraping. You sign in and collect what the account can see. Technically it is the same four lines of code. Legally it is a different act, because signing in means you accepted the terms, and contract is the claim platforms actually win on. The mechanics of session cookies, CSRF tokens and silent session expiry, plus the case law that treats authenticated access differently, are covered in scraping login-protected sites and are not repeated here.
Most commercial social scraping runs through door two or door three while its marketing copy describes door one. Knowing which one your vendor uses is the single most useful question you can ask them.
What each network actually costs
This is where the older version of this article, and most competing ones, are simply out of date. The numbers below were read from each platform's own documentation on 13 August 2026.
| Network | Sanctioned route | What it costs | What it will not give you |
|---|---|---|---|
| X | X API v2, pay-per-usage | $0.005 per Post read | Volume: 2M Post reads per month cap |
| Data API, OAuth required | 100 queries/min free tier | Non-OAuth traffic is blocked outright | |
| YouTube | Data API v3 | Free, quota-limited | More than 100 search calls per day |
| Meta | Graph API, business accounts | Free | Anything about accounts you do not manage |
| TikTok | Research API | Free | Access, unless you are a nonprofit researcher |
| Open Permissions only | Free | Any profile that has not authorized your app |
X
The claim that X sits "behind a paid API with tiered access" was true and is not any more. X API v2 now runs on pay-per-usage credits. The documentation states it plainly: "No subscriptions" and "Pay only for what you use." Reading posts is priced at $0.005 per resource; reading your own account's data is $0.001 per resource; pay-per-usage accounts are "capped at 2 million Post reads per monthly billing cycle." Full-archive search runs at 1 request per second and 300 per 15 minutes per app. Read the X API pricing documentation before budgeting, because the shape of the bill changed, not just the amount.
Do the arithmetic before you reach for a scraper. A thousand posts through the official route costs $5. A thousand posts through a third-party actor costs somewhere between $1 and $2 and carries a block risk. That is a narrower gap than the one that existed in 2023, and for small jobs the sanctioned path is now the cheaper answer once you price your own maintenance time.
Reddit's own help documentation gives free access "100 queries per minute (QPM) per OAuth client id" and states that "Traffic not using OAuth or login credentials will be blocked." That second sentence matters more than the first: Reddit closed the anonymous door, so any Reddit collection at all now involves credentials, which puts it through door three and under Reddit's user terms. The Reddit Data API rate limits are published; the commercial per-call rate is not on a public pricing page, and the figure circulating in every roundup traces back to press statements from 2023 rather than to a Reddit document. Treat it as unverified.
Reddit is also the platform currently doing the most litigating, which is covered below.
YouTube
The most generous of the six, and the one where scraping is usually unnecessary. The YouTube Data API gives each project "100 search.list calls, 100 videos.insert calls, and 10,000 units per day combined for all other endpoints." Reads are cheap: videos.list, channels.list, commentThreads.list and playlistItems.list all cost one unit each. So 10,000 comment-thread pages a day, free, with a stable schema. Search is the scarce resource at 100 calls, and that is the gap people scrape to fill, along with deep comment threads and historical statistics the API does not retain.
Meta: Facebook and Instagram
The Graph API is built around accounts you manage. For anything else the surface is small. The Instagram API with Facebook Login lets you "find hashtagged media, and get basic metadata and metrics about other Instagram Businesses and Creators". It "Requires an Instagram Business or Creator Account linked to a Facbook Page", and the misspelling there is Meta's own, still sitting in the Instagram Platform documentation. Personal accounts are out of reach entirely.
The research door closed and reopened narrower. CrowdTangle was retired on 14 August 2024. The Meta Content Library that replaced it is broader in coverage: Facebook Pages, posts, comments, profiles, groups and events, plus Instagram accounts, posts and comments, across more than 100 searchable fields. The gate is who gets in. "Applicants must be affiliated with a qualified academic institution or a qualified research institution", and both categories require nonprofit status. A marketing team cannot apply. That single change is why so much Meta social listening moved to scraping between 2024 and 2026.
TikTok
Two official products, both narrow. The Research API is open to academic institutions in the US, EEA, UK or Switzerland and to nonprofit or independent research institutions in the EU; applicants must show "independence from commercial interests", disclose funding and provide evidence of ethical review, with a stated turnaround of four weeks. The Commercial Content API covers advertising data and is geographically bounded: "in this phase we are ONLY including data from EU countries, while a researcher/ professional who is requesting it can be located in any country." Ads stay queryable for one year after they were last shown.
For commercial creator and hashtag analytics, there is no sanctioned route. Everything in that market is scraping.
The most valuable network for B2B and the only one with effectively no public data API. Microsoft's own documentation is unambiguous: "Open Permissions are the only permissions that are available to all developers without special approval", and those permissions cover signing a member into your app and posting on their behalf. Everything else — Marketing, Sales Navigator, Talent, Learning — needs partner approval, and the Compliance program is closed to new requests. Nothing in the LinkedIn API access documentation returns data about a member who has not authorized your application.
So the entire LinkedIn data economy is unsanctioned by construction. That is worth holding in mind while reading the next section.
The law, and the four things every roundup gets wrong
hiQ did not beat LinkedIn. This is the most repeated error in the category, and the earlier version of this article repeated it too. hiQ won three preliminary-injunction rounds on the Computer Fraud and Abuse Act between 2017 and April 2022, and then lost the case. On 4 November 2022 the Northern District of California granted LinkedIn partial summary judgment on breach of contract. A stipulated judgment followed in December 2022: $500,000 against hiQ, all other monetary relief waived, plus a permanent injunction barring hiQ from scraping or accessing LinkedIn in violation of its User Agreement and from "developing, using, selling, or distributing any software or code for data collection from LinkedIn." The surviving lesson is narrow and durable: the CFAA is a weak weapon against logged-out collection, so platforms sue on contract instead.
Bright Data beat Meta because it was logged out, not because public data is free. In Meta Platforms v. Bright Data (N.D. Cal., 23-cv-00077), decided by Judge Edward Chen on 23 January 2024, Meta lost on summary judgment because its Terms bind users, and Bright Data scraped Facebook and Instagram while signed out. The court held that Bright Data "was not 'using' Facebook as contemplated by the Terms when it scraped public data while not logged-in", and struck down a survival clause that carried no termination date. Commentary filed this as a win for scraping. It is more precisely a win for logged-out scraping, and the corollary is uncomfortable: with an account you are a user, you did accept the terms at registration, and the claim that failed against Bright Data is live against you.
The newest weapon is not the CFAA at all. On 31 July 2026, Judge Paul Engelmayer of the Southern District of New York largely denied motions to dismiss in Reddit's suit against Perplexity AI and SerpApi (No. 25-cv-8736). Reddit's theory is anti-circumvention under DMCA §1201: that the defendants got at Reddit content by circumventing Google's SearchGuard protection on search results pages, rather than by scraping Reddit directly. The court let the §1201 claims proceed against both defendants while dismissing the §1201(b) trafficking claim against SerpApi and the state-law unfair competition and unjust enrichment counts. A DMCA anti-circumvention claim carries statutory damages and does not require the plaintiff to prove a $5,000 loss, which is what made the CFAA so hard to win on. Thirteen days old as of this writing and unresolved on the merits, but it is the case to watch, because it aims at the scraping-as-a-service layer rather than at the end user.
Terms of service are not the toothless thing they are described as. The received wisdom that violating a platform's terms "rarely lands you in court, it just gets you banned" held up until platforms started suing the vendors. LinkedIn filed against Proxycurl, one of the best-known LinkedIn profile APIs, in January 2025. Proxycurl shut down on 4 July 2025, its founder writing that the company was "shutting Proxycurl down to comply with the legal settlement with LinkedIn" because fighting a well-funded opponent with no prospect of recovering fees was not survivable. Account and IP bans remain the common outcome; they are the floor, not the ceiling.
The safest project collects public, non-personal data — aggregate engagement, content metrics, public business profiles — for research and analytics, while logged out, from a route the platform has not technically closed. The riskiest harvests personal data from behind a login and uses it for outreach. Most real projects sit somewhere in between, and knowing exactly where is worth more than any tool comparison.
None of this is legal advice, and all of it moves.
Personal data runs on a separate track, and it moved this month
Everything above is about access. Privacy law is about the record itself, and it does not care that a profile was public.
GDPR. A name, a photograph, a job history and a username are personal data whether or not the person made them visible. The EDPB adopted Opinion 28/2024 on 18 December 2024 setting out what a legitimate-interest basis for scraping personal data into AI models would actually have to satisfy. Two months earlier, on 28 October 2024, seventeen data protection authorities signed a concluding joint statement on data scraping built around one sentence: "Personal information, even when it is publicly accessible, is subject to privacy laws and must be adequately protected." The enforcement record is the argument: the Dutch DPA fined Clearview AI €30.5 million, with up to €5.1 million in further penalties for continued non-compliance, over a facial recognition database the regulator described as more than 30 billion photos scraped without consent. Clearview's own site today advertises "60+ Billion Images". The company kept building; the fines kept arriving.
California, as of twelve days ago. The Delete Act defines a data broker as "a business that knowingly collects and sells to third parties the personal information of a consumer with whom the business does not have a direct relationship." Read that against a prospect list built from public profiles and sold on. It fits. California residents have been able to file a single deletion request to every registered broker through the Delete Request and Opt-out Platform since 1 January 2026, and brokers were required to begin processing those requests on 1 August 2026. From then on a broker must check DROP at least once every 45 days and report the status of each request within 45 days of retrieving it. Non-compliance runs at $200 per request per day.
Read that penalty structure carefully. It is per request, per day, and a deletion queue you are not checking accrues it silently in parallel across every requester. This is the compliance obligation most likely to catch a marketing team that thinks of itself as a marketing team rather than as a data broker.
The lawful route nobody markets. Article 40 of the EU Digital Services Act gives vetted researchers a right to platform data, and the delegated act putting it into operation was adopted by the Commission on 2 July 2025, with the first researchers vettable from October 2025. Applications go through national Digital Services Coordinators and require research-organization affiliation, independence from commercial interests and a commitment to publish. For academic work on very large platforms, this is now a real alternative to scraping. For commercial work, it is not available, and pretending otherwise wastes weeks.
The Custom Audiences trap
The most common plan for a scraped social list is the one that cannot be executed. Meta's ad platform does not accept a list of profiles. A customer-list Custom Audience is built from identifiers you already hold, hashed client-side before upload. Meta's marketing documentation is precise about the mechanism: "You must hash data as SHA256; we don't support other hashing mechanisms." It is equally precise about whose data it expects, describing the advertiser as "the owner of your business's data" and therefore "responsible for creating and managing this data." A scraped profile URL is not an identifier Meta will match. A scraped email you never collected from the person is not data you own.
The size thresholds are smaller than the ones in circulation. Meta's lookalike audiences documentation states that "if you have a Custom Audience with at least 100 people, you can build lookalike audiences based on it", and for conversion-based lookalikes suggests "200 or more members who converted". The "you need at least 1,000" figure that appears in most guides, this article's earlier version included, is not in Meta's documentation.
Scraped social data earns its keep in research, enrichment and directly addressed outreach. It does not go into an ad account.
Tools, and what they cost
Prices below were read from each vendor's own pricing page on 13 August 2026.
Scraping platforms and APIs
Apify sells a marketplace of actors on top of a credit model: Free at $0 with $5 of platform credit, Starter $29/month, Scale $199/month, Business $999/month, each including credit of the same value and each billed "+ pay as you go" beyond it. Individual actors then price per result. The Instagram Scraper runs from $1.50 per 1,000 results on Business, rising to $2.70 on the free plan, and reports 39,000 monthly active users. The most-used TikTok actor runs from $1.70 per 1,000 results. Both had been modified within hours of this check, which is the useful signal about a marketplace actor: a social scraper that has not shipped in a month is already decaying.
Bright Data sells the same job three ways. Its social media scrapers cover Facebook, Instagram, LinkedIn, TikTok, X, Pinterest, Quora, YouTube, Vimeo and Bluesky at $1.50 per 1,000 records pay-as-you-go, dropping to $1.30 on a plan that includes 384,000 records a month. Its general Web Scraper API starts at $0.75 per 1,000 records, Web Unlocker at $1 per 1,000 requests, residential proxies at $2.50 per GB. Or you skip collection entirely and buy the finished dataset: 905 million LinkedIn profiles or 1.1 billion Instagram records, from $250 per 100,000 records and up to $0.0025 per record.
That last option deserves more attention than it gets. For a one-off market study, buying 100,000 records for $250 beats building a LinkedIn scraper by every measure including legal exposure, because the collection risk sits with the vendor.
General-purpose scraping APIs — Zyte, ScrapingBee, Oxylabs, ScraperAPI — handle proxy pools, headless browsers and anti-bot for you while you supply the parsing. They are the right layer when you already know exactly which pages you want.
Audience and lead tools
Phantombuster prices on execution time rather than records: Start $69/month for 20 hours and 5 phantom slots, Grow $159 for 80 hours and 15 slots, Scale $439 for 300 hours and 50 slots, with AI and email-finding credits attached to each. Time-based pricing punishes inefficient targeting in a way per-record pricing does not, which is either a discipline or a tax depending on how tight your filters are.
The Sales Navigator export tools price per lead. Evaboot runs one plan from $9/month on credits, where one exported lead is one credit and one found-and-verified email is another. Wiza offers a free tier of 20 valid emails, then $49/month for 100 emails and 100 phone numbers, $99 for 500 emails, $199 for 500 of each, with annual unlimited-email plans at $83 and $166 a month. TexAu, Captain Data and Lusha occupy the same band. Every one of these tools operates on a network with no public data API and terms that forbid the activity, which is the trade you are making.
Listening and analytics
Two corrections to the standard list. Brandwatch and Talkwalker are no longer independent competitors of the companies they are usually listed beside: Brandwatch has been a Cision company since 2021, and Hootsuite announced its acquisition of Talkwalker on 8 April 2024. A comparison table that puts Talkwalker in "enterprise" and Hootsuite in "mid-market" is comparing a company with itself. Meltwater, Sprout Social, Brand24 and Mention remain separate. None of them publishes list prices, so any figure you see for them came from a sales call, not a pricing page.
Social Blade is the cheap public option for creator statistics, and its coverage is narrower than the roundups say: its own site lists YouTube, Twitch, Instagram and Twitter. TikTok is not among them, despite appearing in most write-ups of this tool, this article's earlier version included.
No-code and desktop
Octoparse publishes real numbers: Free with 10 tasks and 50,000 exported rows a month, Standard $69/month or $58 billed annually, Professional $249/month or $209 annually with 20 concurrent cloud runs. ParseHub still answers and its help center is up. Every marketing and pricing page renders through JavaScript alone, though, and returns an empty document to a plain fetch, so its current plan prices could not be read from the vendor on 13 August 2026. It is the one tool in this article whose price is unverified, and that is stated rather than filled in from a third-party review site.
| Category | Best for | Verified entry price | Named tools |
|---|---|---|---|
| Scraping platforms | Scale, custom pipelines | $29/mo plus $1.50/1k results | Apify, Bright Data |
| Finished datasets | One-off studies | $250 per 100k records | Bright Data datasets |
| Lead export tools | B2B prospecting | $9/mo credits, $49/mo tiers | Evaboot, Wiza, Phantombuster |
| Listening suites | Sentiment, brand monitoring | Not published by any vendor | Brandwatch, Meltwater, Brand24 |
| No-code scrapers | Small ad-hoc jobs | Free tier, $58/mo annual | Octoparse |
What breaks when you scale it
A hundred profiles is a script. A hundred thousand is an operation, and four things change.
Blocks stop being errors and start being silence. A blocked social platform rarely returns 403. It returns a login wall, an empty feed, or a truncated list that looks exactly like a small account. Your job finishes, your row count is plausible, and your dataset is wrong. Assert on shape, not on status codes: if a profile that had 4,000 followers yesterday returns 12 today, fail the run rather than write the row.
Cost per record dominates. At $1.50 per 1,000 records, a million records is $1,500 and the decision is trivial. Build it yourself and the residential proxy bill alone runs $2.50 per GB, before proxy rotation logic, CAPTCHA solving and the engineer maintaining it against a page structure that changes without notice. One wasted retry per profile across a million-profile crawl is not a rounding error, it is a second crawl.
Schema drift outruns your parser. Social platforms ship UI changes continuously and without changelogs. A selector-based scraper written in June is a coin flip in August. This is the argument for a marketplace actor or a managed API over your own code: you are buying somebody else's obligation to fix it by Tuesday.
Deletion becomes a running obligation, not a one-time task. Once you hold personal data at volume, every subject who deletes their profile, and every Californian who files a DROP request, creates work in your pipeline. Handling that means keeping provenance per record: where it came from, when, and under what basis. Retrofitting provenance onto a dataset you already shipped is the most expensive mistake in this article, and a managed data extraction service or a delivered data-as-a-service feed exists partly because that obligation is easier to carry once than fifty times.
Where this stops working
Scraping is the wrong instrument in four situations, and recognizing them early saves months.
Private data. Anything requiring a friend connection, a group membership or a paid account. The technique works; the exposure changes completely once you sign in.
Anything you plan to put into an ad account. Covered above. The ad platforms check.
Regulated personal data at volume, without a lawful basis and a deletion path. Not a technical problem and not solvable with better proxies.
Real-time monitoring at second-level latency. Scrapers poll. If your use case genuinely needs alerts within seconds, you want a streaming product or a platform webhook, and you should pay for one. Review and mention monitoring on a scheduled cadence is a different product from a firehose, and confusing the two produces an expensive scraper that is always slightly late.
Choosing, in four questions
Does a sanctioned route exist for this network? For YouTube it does and it is generous. For Meta it does if you manage the account. For LinkedIn it does not, and no vendor's marketing changes that.
Is the data personal? If yes, the tool choice is the small decision and the lawful basis is the large one.
Is this one study or a standing feed? For a one-off, a finished dataset at $250 per 100,000 records or a free tier will beat anything you build. For a standing feed, budget for maintenance from day one, because the scraper will break and the only question is whether it breaks loudly.
Who carries the block risk? Your engineer, a marketplace actor's author, or a vendor with a contractual uptime commitment. That answer determines your real cost far more than the sticker price does.
Test on a trial before you buy, and test against a target you already know the answer for. A scraper that returns 200 followers for an account with 2,000 is easier to catch on a known account than on 10,000 unknown ones.
What it comes down to
The technical difficulty of social media scraping has been falling for a decade while the legal and compliance cost has been rising. In 2026 the collection itself is a commodity priced between $0.75 and $2.70 per thousand records, available from vendors who carry the anti-block work for you. What is not a commodity is knowing which door you went through, what basis you hold the data on, and what happens when somebody asks you to delete it.
So the ordering that survives contact with reality is backwards from the one most guides use. Decide what you are allowed to collect, then how you will delete it, then which tool collects it. A list you cannot lawfully keep is worth less than no list, and it costs the same to build.