Web Scraper - Practitioner Directory

via Freelancer ·

Budget / SalaryA$30–250
TypeFreelance project
LocationRemote
Posted3 hours ago
Website Scraping Project

Overview
Experienced web automation/data-extraction developer to build a reliable, and reusable tool that extracts data from a public practitioner register:

https://bams.vba.vic.gov.au/bams/s/practitioner-search
The public register operates on a Salesforce Experience Cloud site.

Scope
Required data fields:
• Record type: individual or company
• Practitioner name
• Business or company name
• Registration number
• Registration category and class
• Registration status
• Registration commencement, Anniversary and expiry date
• Conditions and limitations (if displayed)
• Business Address
• Contact Details
• Source detail-page URL
• Date and time extracted

Technical requirements
Playwright with Python is preferred. Another approach may be proposed if the developer explains why it is quicker to build and more reliable.

The solution must include:
• Configuration file for practitioner classes, statuses, search partitions, pacing & output paths
• Deterministic and repeatable search coverage
• Correct handling of pagination and result limits
• Documented method for subdividing searches to reach the site’s maximum result count
• Conservative sequential request pacing
• Configurable delay between page interactions
• Exponential backoff for temporary failures
• Limited and configurable retries
• Pause or termination if the site returns 403, 429, CAPTCHA, access-denied or similar
• Checkpointing so an interrupted run can resume without restarting
• Local caching where appropriate to avoid unnecessary repeated requests
• Deduplication using registration number
• Preservation of practitioners holding multiple registration classes
• Structured error handling
• Detailed run and validation logs
• UTF-8 CSV output compatible with Microsoft Excel

Search coverage and validation
The solution must not rely on random names or searches.
The developer must develop a systematic coverage strategy based on the filters and search capabilities exposed by the website. This may include name, registration category, class, status, alphabetic prefix, location or another deterministic partitioning method.

Every run must record:
• Search parameters used
• Search start and completion time
• Number of results reported by each search
• Number of result links discovered
• Number of detail pages successfully processed
• Number of records saved
• Duplicate records encountered
• Failed, skipped and retried pages
• Search partitions that reached a result cap
• Search partitions that did not complete

The final reconciliation report must identify:
• Total searches performed
• Total result links discovered
• Total detail pages processed
• Total unique registration records
• Duplicate count
• Failed pages
• Incomplete searches
• Records by practitioner category, class, status and record type
• Any known coverage limitations

Deliverables
The completed project must include:
• Full, uncompiled source code
• Dependency and version files
• Configuration file with explanatory comments
• README with installation and execution instructions
• Search coverage methodology
• UTF-8 CSV final dataset
• Human readable validation and reconciliation report
• Separate CSV containing failed or incomplete records
• Sample configuration for a small test run
• Correction of reproducible defects identified within 14 days of acceptance

The software must run locally without an ongoing subscription, proprietary cloud service or on a developer controlled account.
python data processing excel web scraping data mining data extraction automation data management
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.