Wayback Data Extraction Workflow
Budget / Salary100–800 SGD
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a personal list of several thousand archived URLs on the Wayback Machine and need a repeatable workflow that can process them in bulk. For every snapshot the workflow should:
• Pull out the full text content, all embedded images, and any available metadata (titles, descriptions, dates, alt tags, etc.). Give separate selectors so I'm able to extract only the required data.
• Save a clean, self-contained copy of the page so it can be opened offline as a standalone webpage. A option to go discover for subpages and extract data from them too.
• Generate a parallel JSON file that structures the same information in a machine-readable way.
Speed and resilience matter; the run must cope with large batches without timing out or getting blocked by the Wayback servers. If an individual URL fails, the process should log the error and continue, with a summary report at the end.
Please build this in an easily reproducible environment—Python with requests/BeautifulSoup or Scrapy, Node with Puppeteer, or another language you can justify—and document any dependencies so I can spin it up on my own machine. A simple command-line interface that accepts my URL list, an optional date range, and an output directory is ideal. It has to work across the whole timeline.
Deliverables
1. Source code and any helper scripts.
2. Instructions to install and run (readme or screencast).
3. A short test run on 200 sample URLs demonstrating the JSON output and offline pages.
4. Error-handling and logging strategy.
I’ll consider the job complete once I can replicate your results on my side with no manual tweaks.
I'll share the 200 urls that need to be tested.
• Pull out the full text content, all embedded images, and any available metadata (titles, descriptions, dates, alt tags, etc.). Give separate selectors so I'm able to extract only the required data.
• Save a clean, self-contained copy of the page so it can be opened offline as a standalone webpage. A option to go discover for subpages and extract data from them too.
• Generate a parallel JSON file that structures the same information in a machine-readable way.
Speed and resilience matter; the run must cope with large batches without timing out or getting blocked by the Wayback servers. If an individual URL fails, the process should log the error and continue, with a summary report at the end.
Please build this in an easily reproducible environment—Python with requests/BeautifulSoup or Scrapy, Node with Puppeteer, or another language you can justify—and document any dependencies so I can spin it up on my own machine. A simple command-line interface that accepts my URL list, an optional date range, and an output directory is ideal. It has to work across the whole timeline.
Deliverables
1. Source code and any helper scripts.
2. Instructions to install and run (readme or screencast).
3. A short test run on 200 sample URLs demonstrating the JSON output and offline pages.
4. Error-handling and logging strategy.
I’ll consider the job complete once I can replicate your results on my side with no manual tweaks.
I'll share the 200 urls that need to be tested.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.