German-Austrian B2B Directory Scrape

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I need a senior-level web scraping specialist to collect company information at scale from several large public B2B directories that list firms operating in Germany and Austria. The data will fuel an upcoming lead-generation initiative, so accuracy, breadth, and the ability to refresh the crawl later are essential.
The sites are public but heavily rate-limited; expect to work with rotating proxies, headless browsers (e.g. Selenium or Playwright), and a robust crawler framework such as Scrapy. I will share the exact URLs, field map, and volume targets once we agree on terms, but plan for well over a million records.

Key expectations
• Build a repeatable, well-documented script or small tool (Python preferred, but Node.js is fine) that I can rerun without code changes.
• Capture standard company fields—name, full address, phone, email when available, website, and business category—and return the dataset as CSV or XLSX.
• Implement thorough deduplication and basic data cleaning so the final file is ready for direct import into a CRM.
• Respect robots.txt where legally required while still achieving full coverage through smart throttling and proxy rotation.

Deliverables
1. Working scraper source code with clear setup instructions.
2. One complete data export covering the agreed directories.
3. Short technical report summarising crawl duration, record count, and any skipped pages.

Acceptance criteria
• At least 98 % of listed companies captured per directory sample checks.
• Duplicate rate below 1 %.
• Script runs end-to-end on my machine with the provided README.

If you have handled similar high-volume extractions in the DACH region—or anywhere with strict anti-bot measures—let’s talk.
python web scraping software architecture data mining scrapy data extraction selenium data management
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.