Text Data Cleaning Pipeline -- 2

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I have two active text streams—pages collected through web scraping and records extracted directly from our database—and neither is ready for analysis. Encoding problems, leftover HTML tags, inconsistent casing, duplicates, and other typical messiness need to be resolved before the data reaches our analytics team.

I’m looking for a repeatable, well-documented pipeline that:
• pulls the raw text from both sources on demand,
• applies thorough cleaning steps (standardise encoding, strip markup, normalise whitespace and punctuation, remove or flag duplicates/corrupted rows), and
• exports a single, tidy UTF-8 file (CSV or JSON) with a consistent schema.

Python with pandas, BeautifulSoup/Scrapy and a standard SQL connector is my preferred stack, but I’m open to any open-source approach provided the final script can be run end-to-end on my machine without manual tweaks. When you reply, outline the quality checks you’ll build in and the time you’ll need. I’ll sign off once your script reliably produces a clean dataset ready for downstream work.
javascript python sql web scraping data mining scrapy beautifulsoup pandas
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.