Patent Data Extraction and Analytics

via Freelancer ·

Budget / Salary£250–750
TypeFreelance project
LocationRemote
Posted2 hours ago
Build a geocoded USPTO patent dataset (submarine cable technologies) — data collection + cleaning + geocoding
Overview
I'm building a structured dataset of US patents in a defined technology area (submarine fibre-optic cable systems and related optical/materials technologies) for a quantitative analysis of where this innovation happens and has happened geographically. The output feeds an academic project modelled on recent work in the economics of innovation that maps patenting to geography — see [attached paper / link], e.g. the commuting-zone patent maps in Figure 13. That's the kind of geocoded output I'm after.
What you'll build
Pull US patent records for a set of search terms and CPC classes I provide, from PatentsView, Google Patents (BigQuery), or USPTO bulk data — whichever you can justify.
For each patent: patent number, title, abstract, filing and grant dates, assignee(s), inventor(s), and inventor/assignee locations.
Clean and disambiguate assignee names (e.g. collapsing "Corning Inc" / "Corning Incorporated" / "CORNING INC." into one entity).
Geocode inventor and assignee locations to city-level lat/long, preserving the raw address components (city, state, country as separate fields) so locations can later be mapped to US commuting zones.
Deliver as CSV + Excel, plus the Python or R script so I can re-run and extend it.
Workflow
Start with a sample of a few hundred patents, fully processed end-to-end, so we confirm the pipeline is correct before scaling to the full corpus. I'd rather get the sample exactly right than receive a large, messy dataset.
To apply, tell me briefly:
Which patent data source you've used, and on what specific project.
How you'd geocode inventor addresses to city level, and how you'd handle missing or malformed location data.
How you'd approach assignee name disambiguation.
A link to a past data project counts for far more than a written description. Exposure to the geography of innovation or innovation economics is a strong plus but not essential — clean, reproducible data work is what matters most.
Timeline
The dataset build needs to be completed between now and 25 September 2026.
python data processing excel data scraping data extraction data analysis patent landscape bigquery data management pandas
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.