Resources · Blog

Guides & Basics

Core concepts and how web scraping works.

Guides & Basics
Web Scraping Best Practices: Core Principles

Web scraping best practices rechecked in August 2026: robots.txt under RFC 9309, a Python stdlib parser that got three rules of four wrong until May 2026, Retry-After, Decimal money, and measured storage and parsing costs for a ten-thousand-page run.

30 May · 19 min
Guides & Basics
Best Web Scraping Books: A Curated Reading List

A web scraping reading list rechecked in August 2026: core Python titles with dates, page counts and prices, the regex canon, storage books, and the parts that have quietly rotted.

24 May · 15 min
Guides & Basics
Best Programming Language for Web Scraping: Full Comparison

Eleven languages compared for web scraping, rechecked in August 2026: measured parse times, current library versions and release dates, dead projects flagged, and what the choice actually costs.

22 May · 27 min
Guides & Basics
Web Scraping Project Ideas: Choosing the Right Technology

Sixteen web scraping projects from beginner to advanced, each with a stack, a failure mode and a price per thousand pages. Includes a measured parser benchmark, proxy and scraping-API rates read from vendor pricing pages on 10 August 2026, and the three ideas on this list that stopped working since it was written.

30 April · 26 min
Guides & Basics
Web Crawling vs Web Scraping vs Parsing: Key Differences

Web crawling finds pages, scraping fetches and extracts them, parsing turns markup into structure. Where the line falls, why robots.txt covers only one of the three, measured fetch-versus-parse timings from August 2026, and three parsers returning three different answers for the same tag.

14 April · 23 min
Guides & Basics
Web Crawling vs Web Scraping in E-commerce

Crawling finds product URLs, scraping lifts price and stock. What four large retailers allow in robots.txt, which JSON endpoints stores publish on purpose, what a pass over 200,000 pages costs at list prices, and where the pipeline breaks. Sources read 10 August 2026.

29 March · 15 min
Guides & Basics
DOM Parser: How HTML Becomes a Tree of Nodes

How a DOM parser turns markup into a node tree, and why two parsers build two different trees from the same bytes: the tbody rule that breaks copied selectors, measured parse time and memory for lxml, BeautifulSoup, cheerio, jsdom and lexbor, and library versions read from each vendor in August 2026.

1 March · 17 min
Guides & Basics
HTTP Protocol Basics Every Web Scraper Should Know

HTTP for scrapers, rechecked against the current specifications on 13 August 2026: RFC 9110-9114 instead of the obsolete 7230 series, exact section numbers, measured default headers from four clients, cookie limits, caching, and version fingerprints.

19 December · 34 min
Guides & Basics
HTTP vs HTTPS: Key Differences for Web Scraping

HTTP vs HTTPS rechecked on 13 August 2026: the RFC 9110-9114 series that retired RFC 2818, the 200-day certificate cap in force since 15 March 2026, post-quantum ClientHellos on over 60% of Cloudflare HTTPS traffic, and why TLS fingerprinting blocks a client whose headers are perfect.

17 December · 20 min