Dubizzle Property Scraper
Gemini API-powered scraper extracting structured property data from Dubizzle — images, pricing, amenities, legal IDs, agent contacts.
The Challenge
Property listings are written for humans — prices as prose, amenities as a paragraph, legal identifiers buried in a footer. Selector-based scraping breaks the moment a listing deviates from the template, and most of them do.
Our Solution
Built a two-stage extractor: BeautifulSoup pulls the listing region, then Gemini reads it as language and returns structured fields. Because the second stage interprets rather than pattern-matches, listings that break the DOM template still resolve into clean records covering pricing, amenities, legal identifiers, agent contacts and image sets.
Key Highlights
Gemini interprets listing text, so extraction survives inconsistent markup
Legal and permit identifiers captured as first-class fields
Agent contact details resolved alongside the property record
Image URLs collected per listing for downstream use
Tech Stack
The next great brand starts with a conversation.
Tell us where you are and where you want to be. We'll tell you exactly how we'd get you there — no pitch decks, no fluff, just a clear plan.
hello@webversearena.com
Phone
+91 8220115779
Location
Chennai, India
Response time
< 48 hours
