Government Scraper Fix & Testing

via Freelancer ·

Budget / Salary$10–50
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a Python-based, Dockerised service that collects public records from an online government portal. The crawler currently suffers from queue-handling bugs and occasional container issues, so the first milestone is to stabilise the existing codebase and make the scraping run end-to-end again.

Once stability is restored I want solid quality gates in place:
• Write and wire up full e2e tests that spin the stack in Docker, hit the target site and verify data is written to our queues/storage.
• Manually exercise the run, capture a short screen recording that shows the scrape completing and records being persisted.
• Wrap all work in a clean Git workflow—open a feature branch, commit changes and raise a pull request so the diff is easy to review.

The portal uses a custom CAPTCHA (not ReCAPTCHA or hCaptcha). I don’t yet have a bypass strategy, so I’m open to your ideas—whether that’s third-party solving services, ML-based pattern recognition, or another creative approach. For this phase it’s enough to stub out a solution or spike a proof-of-concept that shows the CAPTCHA can eventually be conquered without manual input.

Tech you’ll touch: Python 3.x, Docker Compose, pytest (or similar) for e2e, and GitHub Actions is available if you’d like to wire in CI.

Deliverables:
1. Fixed, runnable scraper container.
2. Automated e2e test suite with clear pass/fail output.
3. Screen-recorded proof of a successful scrape.
4. Pull request with well-documented commits and brief README update explaining the CAPTCHA plan.

If that sounds straightforward for you, let’s get this scraper humming again and lay the groundwork for a hands-free CAPTCHA bypass in the next iteration.
php python web scraping software architecture git docker docker compose ci/cd
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.