Dark Web Scraping & Intelligence Dashboard System (Docker Stack)
Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted1 hour ago
Project Description:
We are looking for a skilled Engineer / Developer to build a fully localized, containerized Scraping & Intelligence Platform deployed via Docker Compose.
The system automates .onion site discovery, recursive multi-engine scraping, content extraction, and automated categorization using a free, local tech stack.
(IMPORTANT): Automated bids will be immediately ignored and reported, Briefly outline your experience with Python asyncio, Tor routing/proxies, Playwright/FastAPI. and your suggestions.
1/ Core Stack Requirements:
- Backend & API: Python 3.12 (FastAPI), SQLAlchemy / PostgreSQL, Celery + Redis for task queues.
- Scraping & Tor Routing: httpx, Playwright (Headless Chromium), custom SOCKS5 Tor circuit rotation logic, and proxy verification.
- Local AI/ML Engine: Local LLM integration (via Ollama or others) for content summarization, entity extraction (crypto wallets, PGP keys, emails), and automated 20-category classification.
- Frontend Dashboard: Lightweight, single-screen dashboard (Streamlit or React) with full-text search, visual uptime tracking, and bulk data export capabilities.
2/ Key System Capabilities:
- Adaptive Scraping: Auto-switch between lightweight HTTP parsing and JS rendering depending on page complexity.
- Scam/Clone Fingerprinting: Structural/content fingerprinting algorithm to isolate duplicate scam sites from original platforms.
- Automated Monitoring: Periodic uptime status checks, content diff versioning, and automated alerts for status changes.
- Open-Source Infrastructure: Entire stack must be 100% free/open-source and deployable on a single local machine or Linux VPS via Docker Compose.
3/ Additional Info:
Comprehensive step-by-step documentation, architecture specs, and tool lists will be provided immediately upon hiring.
We are looking for a skilled Engineer / Developer to build a fully localized, containerized Scraping & Intelligence Platform deployed via Docker Compose.
The system automates .onion site discovery, recursive multi-engine scraping, content extraction, and automated categorization using a free, local tech stack.
(IMPORTANT): Automated bids will be immediately ignored and reported, Briefly outline your experience with Python asyncio, Tor routing/proxies, Playwright/FastAPI. and your suggestions.
1/ Core Stack Requirements:
- Backend & API: Python 3.12 (FastAPI), SQLAlchemy / PostgreSQL, Celery + Redis for task queues.
- Scraping & Tor Routing: httpx, Playwright (Headless Chromium), custom SOCKS5 Tor circuit rotation logic, and proxy verification.
- Local AI/ML Engine: Local LLM integration (via Ollama or others) for content summarization, entity extraction (crypto wallets, PGP keys, emails), and automated 20-category classification.
- Frontend Dashboard: Lightweight, single-screen dashboard (Streamlit or React) with full-text search, visual uptime tracking, and bulk data export capabilities.
2/ Key System Capabilities:
- Adaptive Scraping: Auto-switch between lightweight HTTP parsing and JS rendering depending on page complexity.
- Scam/Clone Fingerprinting: Structural/content fingerprinting algorithm to isolate duplicate scam sites from original platforms.
- Automated Monitoring: Periodic uptime status checks, content diff versioning, and automated alerts for status changes.
- Open-Source Infrastructure: Entire stack must be 100% free/open-source and deployable on a single local machine or Linux VPS via Docker Compose.
3/ Additional Info:
Comprehensive step-by-step documentation, architecture specs, and tool lists will be provided immediately upon hiring.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.