E-commerce Text and Image Scraper
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I need a rock-solid scraper that can pull both product text and all related images from a JavaScript-heavy e-commerce site. The pages rely on lazy loading and endless scrolling, so I’ve been leaning on Playwright to render each view while using Scrapy’s crawl engine to fan out across categories and pagination. The goal is straightforward: capture every SKU’s name, price, description, variants, and the full-resolution image set, then export the lot to a clean JSON (or CSV) plus a mirrored folder of images.
Here’s how I picture the engagement: you build a Playwright-powered spider that logs in if required, waits for dynamic content, and hands the DOM to Scrapy for high-speed harvesting. The solution should respect robots.txt delays where possible and retry gracefully on captchas or throttling. I’ll run this on a headless Ubuntu box, so a requirements.txt or Pipfile and a brief README are essential.
Deliverables
• Fully commented Playwright + Scrapy script(s)
• Sample output file with 50–100 products demonstrating text and image links
• Local image archive matching the sample output
• README with setup, run command, and troubleshooting tips
Acceptance criteria: the script must finish a full category crawl without manual intervention, save all images at original quality, and structure data exactly as outlined above. If you have prior experience bypassing dynamic loading on retail platforms, that’s a big plus—please mention it when you bid.
Here’s how I picture the engagement: you build a Playwright-powered spider that logs in if required, waits for dynamic content, and hands the DOM to Scrapy for high-speed harvesting. The solution should respect robots.txt delays where possible and retry gracefully on captchas or throttling. I’ll run this on a headless Ubuntu box, so a requirements.txt or Pipfile and a brief README are essential.
Deliverables
• Fully commented Playwright + Scrapy script(s)
• Sample output file with 50–100 products demonstrating text and image links
• Local image archive matching the sample output
• README with setup, run command, and troubleshooting tips
Acceptance criteria: the script must finish a full category crawl without manual intervention, save all images at original quality, and structure data exactly as outlined above. If you have prior experience bypassing dynamic loading on retail platforms, that’s a big plus—please mention it when you bid.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.