The task
The agency works with a handful of premium residential complexes and has to stay on top of every new listing and price change in exactly those developments. Watching several portals manually every day is unrealistic, and the best listings go to whoever sees them first.
What we did and what data we collected
We set up regular real estate listings collection from several portals with filtering down to the target complexes: price, area, floor, layout, status, seller, date. We configured change monitoring: a new listing, a price drop, a delisting — everything lands in the agency's feed with a history for every lot.
Challenges and how we solved them
First, attribution to a specific complex. Addresses in listings are written loosely, sometimes only the district is mentioned. We matched listings to complexes by a combination of signals (address, coordinates, complex name, landmarks) rather than a single address line.
Second, duplicates. The same lot is posted on several portals and by different agents. We deduplicated listings so the feed shows real properties, not their copies.
Third, change tracking and anti-bot restrictions. Some sites serve data dynamically and throttle automated requests; we collected carefully, at a sensible frequency and with proxies.
The result
The agency got a single feed for its premium developments with new listings and price changes over time — the team reacts to interesting properties faster, and monitoring no longer depends on manually checking the portals.
Services in this case study: Real estate scraping · Website and Data Change Monitoring · Data for the Real Estate Market
Useful reading: Scraping dynamic sites · Web scraping proxies · Data normalization