Facebook Text Data Collection
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I need a reliable contributor who can gather English-language text from Facebook and package it into a clean, well-structured dataset that can be used for downstream AI/NLP experimentation. The assignment is exclusively text-based; no images, audio, or mixed-media assets are required.
Scope of work
You will capture publicly available posts and their associated comments coming from Facebook pages or groups that allow legal, compliant scraping. Content must be written in English and sourced directly from the platform—no third-party aggregators. I am interested in a broad topical spread rather than a single niche so the final corpus reflects natural, everyday language as it appears in social media.
Compliance & privacy
All data has to be collected in accordance with Facebook’s terms of service and local regulations. Strip or hash any personally identifiable information (names, profile links, IDs) before delivery. I will request a brief explanation of your compliance workflow so I can document provenance on my side.
Preferred approach
Whether you use the official Graph API, Selenium, or a custom Python script, I only care that the method is reproducible and the output is consistent. Please note the exact toolchain and versions you rely on so I can replicate the pull later if needed.
Deliverables
• CSV or JSON file containing raw post text, comment text, timestamps, and an anonymised author identifier
• A short README detailing collection method, cleaning steps, and field definitions
• Log file or notebook demonstrating the scraping script executing successfully on a small sample run
Acceptance criteria
The dataset must reach at least 50,000 unique English sentences, pass an automated language-detection check (≥ 95 % English confidence), and contain no unredacted personal data.
If this sounds straightforward and you have past experience scraping Facebook text at scale, let’s discuss timeline and milestones so we can get moving quickly.
Scope of work
You will capture publicly available posts and their associated comments coming from Facebook pages or groups that allow legal, compliant scraping. Content must be written in English and sourced directly from the platform—no third-party aggregators. I am interested in a broad topical spread rather than a single niche so the final corpus reflects natural, everyday language as it appears in social media.
Compliance & privacy
All data has to be collected in accordance with Facebook’s terms of service and local regulations. Strip or hash any personally identifiable information (names, profile links, IDs) before delivery. I will request a brief explanation of your compliance workflow so I can document provenance on my side.
Preferred approach
Whether you use the official Graph API, Selenium, or a custom Python script, I only care that the method is reproducible and the output is consistent. Please note the exact toolchain and versions you rely on so I can replicate the pull later if needed.
Deliverables
• CSV or JSON file containing raw post text, comment text, timestamps, and an anonymised author identifier
• A short README detailing collection method, cleaning steps, and field definitions
• Log file or notebook demonstrating the scraping script executing successfully on a small sample run
Acceptance criteria
The dataset must reach at least 50,000 unique English sentences, pass an automated language-detection check (≥ 95 % English confidence), and contain no unredacted personal data.
If this sounds straightforward and you have past experience scraping Facebook text at scale, let’s discuss timeline and milestones so we can get moving quickly.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.