Python Crawler for Building Case Archive

via Freelancer ·

Budget / Salary$250–750
TypeFreelance project
LocationRemote
Posted1 hour ago
Title: Python crawler for Norwegian municipal case archive (Saksinnsyn) — Docker, Postgres, resume, provenance

Brief:

I need a senior Python developer to build a polite, resumable crawler for Oslo municipality's public building-case archive (innsyn.pbe.oslo.kommune.no/saksinnsyn). Plain HTTP (requests/httpx + lxml), no browser automation. I have a working single-file prototype (v0.29) that covers discovery, case pages, journalpost lists, attachment download, manifest and block detection; you may reuse or rewrite it, your choice, the price is fixed either way.

Scope

Discover all byggesaker 2019–present via the site's search (I'll give you the working query patterns and the case-number rules, which differ before/after 2025 and must be config per source and year, not hardcoded).
Per case: save the raw case page HTML plus parsed fields (saksnummer as shown + normalized, title, sakstype, status, dates, address, gnr/bnr, unit).
Per case: full journalpost list, raw HTML + parsed rows (number, date, title, inn/ut/intern, sender/recipient).
Per journalpost: every attachment (file id, filename, file). Items that are login-gated or withheld are recorded with status and reason, not fetched.
Linked cases (kobling): record the link and fetch the linked case.
Every download logged: URL, time, HTTP status, size, sha256, run id.
Output: raw files on disk (verbatim, never rewritten) plus a Postgres table case → journalpost → document → file → status (fetched / gated / withheld / missing / blocked).
Idempotent and resumable: re-running never duplicates; stop/start any time.
Multi-municipality by design: every record carries kommunenr and source; Oslo fetching is one module behind a common data model. Nothing downstream may know about Oslo's pages.
Delivered as Docker Compose (crawler + Postgres) running on my Windows PC, CLI-driven. No web UI.

Site behaviour you must handle (this is the hard part)
The server intermittently blocks file downloads (showfile.asp) with a captcha page or a redirect to ID-porten login. This is server-side and not IP-based; it also hits a normal browser. Required behaviour: detect both responses, mark the file blocked with reason/time/response hash, keep fetching metadata pages (which stay open), probe at most once per hour, resume automatically when the gate reopens. Fixed User-Agent string that I supply, single session, rate limit configurable (default 1 request / 3 s), daily file budget configurable. No captcha solving, no ID-porten automation, no proxy/IP/session/User-Agent rotation. Bids that propose any of these will not be considered. I'll test this with a mock server returning the block page.

Acceptance

End-to-end run on 200 cases I specify produces the status table and matching files.
Full 2024 run matches sha256 of ~3,200 files I already hold.
Block simulation: crawler stops, logs, resumes, no duplicates.

Terms
Fixed price, milestones 30% start / 40% at 200-case run / 30% at 2024 verified. Code in my GitHub repo from day one, IP assigned to me. Timeline 3 weeks.

You should have: 5+ years Python, prior crawling of government or public-sector archives, Postgres, Docker. Say in your bid what you'd do when the crawler gets the captcha page — I read that line first.
php python web scraping software architecture postgresql docker data extraction http
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.