Web News Curation PDF Exporter

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I am building a Python-based tool that emulates the core experience of Dow Jones Factiva, but focused exclusively on articles that surface through Google. The application must be a web-based UI with clear multi-page navigation: one page for searching, another for reviewing results, and a final page for assembling and downloading a hyperlinked PDF.

Core workflow I need implemented
1. Users enter a query and optional Boolean operators (AND / OR).
• Advanced flags such as “hlp=” (headline + lead para) and “hl=” (headline only) must be honoured, returning results accordingly.
2. Results display in a tidy table with checkbox selection. Live filters (date, source, keyword, etc.) should refine the list in real time.
3. Selected articles are compiled into a single PDF in which every headline links back to the original URL. Pagination, basic styling, and clickable hyperlinks are non-negotiable.

Tech stack suggestions
Python 3.x is required. I am comfortable if you reach for common libraries such as Flask or Django for the UI, BeautifulSoup or wrapper for scraping/aggregation, and ReportLab, WeasyPrint or a similar library for PDF generation—feel free to propose alternatives if they achieve the same result cleanly.

*Limitations
No use of http extraction/any api whatsover even if google news etc

Acceptance criteria
• Search page accepts plain and Boolean queries and returns accurate Google article metadata (title, snippet, source, date, link).
• Navigation is intuitive between search, results, and export pages with no page refresh glitches.
• PDF download triggers within a few seconds for 25 articles or fewer, retaining all hyperlinks and respecting article order chosen by the user.
• Codebase is clean, commented, and ready to run locally via a single requirements.txt and README.

Hand-off deliverables
• Complete source code in a Git-ready folder structure
• requirements.txt and README with setup instructions
• Deployed demo on any free tier hosting (Heroku, Render, etc.) so I can test instantly

If you have prior experience with news scraping, search interfaces, or PDF generation, that will make collaboration smoother, but quality code and clear communication matter most. Let me know how you would approach the Boolean filtering and PDF link preservation, and an estimated timeline to reach a working MVP.
php javascript python web scraping django web development flask beautifulsoup
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.