Daily Glassdoor Review Scraper
Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted2 hours ago
The project involves web-crawling Glassdoor review text and other relevant information for a predefined list of U.S. publicly listed companies. The main objective is to extend an existing Glassdoor dataset, which currently covers data through 2021, to include observations from 2021 through the end of 2025.
The work will include identifying and matching the correct Glassdoor company pages for each firm, collecting available employee reviews and associated metadata, and organizing the collected information into a structured dataset. The new data should follow the structure of the existing dataset as closely as possible.
I will provide a sample file from the existing dataset. The current data include fields such as company identifiers, review ID and date, review title, employment status, job title, overall rating, recommendation, CEO approval, business outlook, pros and cons, helpfulness counts, and category-specific ratings such as work-life balance, culture, diversity and inclusion, career opportunities, compensation and benefits, and senior management.
The primary period of interest is 2021–2025. Duplicate observations should be identified and removed, particularly where the newly collected data overlap with observations already included in the existing dataset.
The final deliverables should include:
1. The complete original Python code used for web crawling, data collection, parsing, cleaning, and processing. The code should be sufficiently documented and reproducible so that I can rerun or modify the collection process independently.
2. The raw or minimally processed output generated by the crawler, whenever feasible.
3. The cleaned and structured review-level dataset, preferably in CSV format, using variable names and formatting consistent with the sample dataset provided.
4. A record of firms for which data could not be collected or reliably matched to a Glassdoor company page, together with any relevant notes or errors encountered during collection.
5. Any auxiliary files or mapping tables used to match the provided company list to Glassdoor firms or company IDs.
Both the source code and all collected data/output files must be shared with me upon completion of the project. I will provide the target company list and an example file from the existing dataset as references for the desired output structure.
The work will include identifying and matching the correct Glassdoor company pages for each firm, collecting available employee reviews and associated metadata, and organizing the collected information into a structured dataset. The new data should follow the structure of the existing dataset as closely as possible.
I will provide a sample file from the existing dataset. The current data include fields such as company identifiers, review ID and date, review title, employment status, job title, overall rating, recommendation, CEO approval, business outlook, pros and cons, helpfulness counts, and category-specific ratings such as work-life balance, culture, diversity and inclusion, career opportunities, compensation and benefits, and senior management.
The primary period of interest is 2021–2025. Duplicate observations should be identified and removed, particularly where the newly collected data overlap with observations already included in the existing dataset.
The final deliverables should include:
1. The complete original Python code used for web crawling, data collection, parsing, cleaning, and processing. The code should be sufficiently documented and reproducible so that I can rerun or modify the collection process independently.
2. The raw or minimally processed output generated by the crawler, whenever feasible.
3. The cleaned and structured review-level dataset, preferably in CSV format, using variable names and formatting consistent with the sample dataset provided.
4. A record of firms for which data could not be collected or reliably matched to a Glassdoor company page, together with any relevant notes or errors encountered during collection.
5. Any auxiliary files or mapping tables used to match the provided company list to Glassdoor firms or company IDs.
Both the source code and all collected data/output files must be shared with me upon completion of the project. I will provide the target company list and an example file from the existing dataset as references for the desired output structure.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.