Government Scraper Fix & Testing
Budget / Salary$10–50
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a Python-based, Dockerised service that collects public records from an online government portal. The crawler currently suffers from queue-handling bugs and occasional container issues, so the first milestone is to stabilise the existing codebase and make the scraping run end-to-end again.
Once stability is restored I want solid quality gates in place:
• Write and wire up full e2e tests that spin the stack in Docker, hit the target site and verify data is written to our queues/storage.
• Manually exercise the run, capture a short screen recording that shows the scrape completing and records being persisted.
• Wrap all work in a clean Git workflow—open a feature branch, commit changes and raise a pull request so the diff is easy to review.
The portal uses a custom CAPTCHA (not ReCAPTCHA or hCaptcha). I don’t yet have a bypass strategy, so I’m open to your ideas—whether that’s third-party solving services, ML-based pattern recognition, or another creative approach. For this phase it’s enough to stub out a solution or spike a proof-of-concept that shows the CAPTCHA can eventually be conquered without manual input.
Tech you’ll touch: Python 3.x, Docker Compose, pytest (or similar) for e2e, and GitHub Actions is available if you’d like to wire in CI.
Deliverables:
1. Fixed, runnable scraper container.
2. Automated e2e test suite with clear pass/fail output.
3. Screen-recorded proof of a successful scrape.
4. Pull request with well-documented commits and brief README update explaining the CAPTCHA plan.
If that sounds straightforward for you, let’s get this scraper humming again and lay the groundwork for a hands-free CAPTCHA bypass in the next iteration.
Once stability is restored I want solid quality gates in place:
• Write and wire up full e2e tests that spin the stack in Docker, hit the target site and verify data is written to our queues/storage.
• Manually exercise the run, capture a short screen recording that shows the scrape completing and records being persisted.
• Wrap all work in a clean Git workflow—open a feature branch, commit changes and raise a pull request so the diff is easy to review.
The portal uses a custom CAPTCHA (not ReCAPTCHA or hCaptcha). I don’t yet have a bypass strategy, so I’m open to your ideas—whether that’s third-party solving services, ML-based pattern recognition, or another creative approach. For this phase it’s enough to stub out a solution or spike a proof-of-concept that shows the CAPTCHA can eventually be conquered without manual input.
Tech you’ll touch: Python 3.x, Docker Compose, pytest (or similar) for e2e, and GitHub Actions is available if you’d like to wire in CI.
Deliverables:
1. Fixed, runnable scraper container.
2. Automated e2e test suite with clear pass/fail output.
3. Screen-recorded proof of a successful scrape.
4. Pull request with well-documented commits and brief README update explaining the CAPTCHA plan.
If that sounds straightforward for you, let’s get this scraper humming again and lay the groundwork for a hands-free CAPTCHA bypass in the next iteration.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.