Bi-Weekly Facebook & Website Scraper
Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted1 hour ago
I need a fully automated pipeline that fetches fresh data from specific Facebook assets and a public website twice a week and drops the results into a structured destination I can immediately work with.
Scope
• Facebook: once the job starts I will supply a short list of page, group and/or profile URLs. The solution should pull all visible posts, their comments, basic engagement metrics and any media links for each run. Please stay within Facebook’s Terms of Service and use the official Graph API wherever possible (access tokens will be provided).
• Website: a single-domain crawl that captures several predefined sections. I will point you to the exact pages; think product blocks, review snippets and article text.
Automation cadence
Run every Monday and Thursday (UTC). A simple cron, GitHub Action, Cloud Function or comparable scheduler is fine as long as it is hands-off for me.
Output
CSV is preferred, with each run stored in its own timestamped folder in Google Drive. If a small SQLite or Postgres database is easier for deduping, feel free to propose it.
Deliverables
- Clean, well-commented code (Python is ideal, open to Node/Go)
- Setup instructions and a one-click deployment script or Dockerfile
- A short README outlining field mappings and runtime steps
- Basic logging so I can verify each execution
Acceptance criteria
1. Pipeline runs unattended on the agreed schedule for one full week during hand-off.
2. Each target record appears exactly once per run, with no missing fields.
3. Any rate-limit or auth error is reported in the log and does not crash the job.
Please let me know which libraries or scraping tools you plan to use and any questions you have about access.
Scope
• Facebook: once the job starts I will supply a short list of page, group and/or profile URLs. The solution should pull all visible posts, their comments, basic engagement metrics and any media links for each run. Please stay within Facebook’s Terms of Service and use the official Graph API wherever possible (access tokens will be provided).
• Website: a single-domain crawl that captures several predefined sections. I will point you to the exact pages; think product blocks, review snippets and article text.
Automation cadence
Run every Monday and Thursday (UTC). A simple cron, GitHub Action, Cloud Function or comparable scheduler is fine as long as it is hands-off for me.
Output
CSV is preferred, with each run stored in its own timestamped folder in Google Drive. If a small SQLite or Postgres database is easier for deduping, feel free to propose it.
Deliverables
- Clean, well-commented code (Python is ideal, open to Node/Go)
- Setup instructions and a one-click deployment script or Dockerfile
- A short README outlining field mappings and runtime steps
- Basic logging so I can verify each execution
Acceptance criteria
1. Pipeline runs unattended on the agreed schedule for one full week during hand-off.
2. Each target record appears exactly once per run, with no missing fields.
3. Any rate-limit or auth error is reported in the log and does not crash the job.
Please let me know which libraries or scraping tools you plan to use and any questions you have about access.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.