Data Collection & Cleaning
Budget / SalaryA$5,000–10,000
TypeFreelance project
LocationRemote
Posted1 hour ago
We are looking for 2-3 freelancers to help with data collection and cleaning across a series of projects.
Scope
The job includes a variety of data sources, for example:
- sourcing data from public websites and building a database with time-stamps
- calls to public APIs
- administrative data that you will be given access to if necessary
You will be responsible for collecting and cleaning the data, and delivering a final product of a usable dataset ready for statistical work. You may need to do manual searches and/or write code to do API calls or similar work. Your workflow should comfortably move between these structures without loss of information.
Deliverables
• A script (or set of reproducible commands) that collects the raw data from each source.
• A cleaned, well-documented dataset in Stata format, ready for analysis.
• A short log or README explaining every transformation, validation check, and any assumptions made.
Acceptance Criteria
1. All source connections (scraper, database query, API call) must run end-to-end without manual tweaks.
2. Field names, data types, and row counts match the specifications we agree on up front.
3. Missing or anomalous values are flagged and handled according to the documented rules.
4. Final Stata file opens without errors and reproduces the summary statistics we validate together.
Scope
The job includes a variety of data sources, for example:
- sourcing data from public websites and building a database with time-stamps
- calls to public APIs
- administrative data that you will be given access to if necessary
You will be responsible for collecting and cleaning the data, and delivering a final product of a usable dataset ready for statistical work. You may need to do manual searches and/or write code to do API calls or similar work. Your workflow should comfortably move between these structures without loss of information.
Deliverables
• A script (or set of reproducible commands) that collects the raw data from each source.
• A cleaned, well-documented dataset in Stata format, ready for analysis.
• A short log or README explaining every transformation, validation check, and any assumptions made.
Acceptance Criteria
1. All source connections (scraper, database query, API call) must run end-to-end without manual tweaks.
2. Field names, data types, and row counts match the specifications we agree on up front.
3. Missing or anomalous values are flagged and handled according to the documented rules.
4. Final Stata file opens without errors and reproduces the summary statistics we validate together.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.