AI-Driven Marine Data Aggregation & Analysis

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
Senior Data / AI Engineer - Marine Data, APIs, Web Scraping & Machine Learning

We are building a new AI-powered travel technology platform that works with large amounts of marine, oceanographic, weather, environmental, geographic and destination data.

We are looking for a strong Data / AI / Machine Learning Engineer to help build the data and intelligence infrastructure behind the platform.

This is not primarily a frontend or UI development role.

We need someone who is very comfortable working with large datasets, APIs, web scraping, automated data pipelines, geospatial data, historical and forecast data, oceanographic datasets, and scoring/recommendation models.

Experience working with marine weather, oceanographic, coastal or environmental data is a major advantage.

What you will work on

You will help us:

* Collect and structure data from multiple public and commercial data sources
* Research and identify reliable marine, weather, environmental and geographic data sources
* Build reliable web scraping and automated data collection pipelines
* Integrate multiple third-party APIs
* Work with oceanographic, marine weather, environmental and location-based datasets
* Collect historical, seasonal, real-time and forecast data
* Clean, normalize, deduplicate and validate large datasets
* Design scalable database structures for destinations and geographic locations
* Match information from different sources to the correct physical locations
* Build automated systems to keep data continuously updated
* Develop scoring, ranking and recommendation logic
* Transform complex raw data into structured information that can be used by a consumer application
* Build data-quality checks, anomaly detection and confidence systems
* Document data sources, transformations, assumptions and limitations
* Work closely with the founder and development team to translate product requirements into reliable data infrastructure

Technical skills we are looking for

Strong experience with:

* Python
* SQL
* PostgreSQL
* APIs / REST APIs
* Web scraping
* ETL / ELT pipelines
* Data cleaning and normalization
* Pandas / NumPy
* Large structured datasets
* Automation
* Scheduled data pipelines
* Geospatial data and coordinate matching
* Git / GitHub

Experience with some of the following is a major advantage:

* Machine learning
* Recommendation engines
* Ranking and scoring algorithms
* Time-series data
* Marine weather data
* Oceanographic datasets
* Environmental APIs
* Meteorological data
* Historical and forecast datasets
* GIS / PostGIS
* Geospatial analysis
* Open-Meteo or similar environmental data services
* Supabase
* AWS / GCP or other cloud infrastructure
* Data validation and anomaly detection
* LLM / AI integrations
* Vector databases or embeddings

The person we want

We are not looking for someone who simply follows instructions or writes individual scripts.

We want someone who can look at a difficult data problem and say:

“Here is how I would structure this, here are the best sources, here is what data we can trust, here is what is missing, and here is how I would automate the entire system.”

You should be:

* Highly analytical
* Extremely organized with data
* Comfortable researching unfamiliar datasets and APIs
* Good at identifying unreliable, incomplete or conflicting information
* Able to design systems, not just scripts
* Proactive in suggesting better technical approaches
* Very careful about data accuracy
* Comfortable working independently
* Able to explain technical decisions clearly to a non-data-specialist founder
* Interested in potentially working with us longer term if the initial project goes well

Data quality is extremely important

We do not want someone who simply collects thousands of records and considers the project finished.

Accuracy, traceability and maintainability matter more than raw volume.

Important data should ideally have:

* A known source
* A clear definition
* Source attribution
* A timestamp or update frequency where relevant
* Validation rules
* Confidence indicators where appropriate
* A process for detecting stale or incorrect information

The system needs to be maintainable and scalable as the platform expands internationally.

Initial project

The first stage will involve auditing an existing database and data architecture.

You will:

1. Review the existing database, schema and datasets
2. Identify missing, unreliable, duplicated or incorrectly structured data
3. Recommend improvements to the data architecture
4. Research and identify the best available data sources and APIs
5. Determine which sources should be used for which types of information
6. Build or improve automated data collection pipelines
7. Normalize, validate and document the collected information
8. Help develop the first version of our scoring and recommendation data engine
9. Create a scalable foundation for future data sources and models

If the collaboration is successful, this can become a significant ongoing role as the platform grows.

When applying

Please answer the following questions directly:

1. Tell us about a project where you collected and combined data from multiple APIs, datasets or websites. What was technically difficult about it?

2. Have you worked with marine, oceanographic, weather, environmental, geographic or other time-series data? Please explain exactly what you worked with.

3. How would you approach matching information from several different sources when the same physical location may have different names, coordinates or IDs?

4. How do you determine whether information collected from an API, open dataset or scraped source is reliable enough to use in a production system?

5. Have you built a scoring, ranking or recommendation system before? Please describe the logic and your role.

6. How would you architect a system that collects information from multiple sources and automatically updates different datasets at different frequencies?

7. How would you track data provenance so we always know where an individual piece of information originated?

8. Please include links to relevant GitHub repositories, technical projects or examples of previous work if available.

Start your proposal with the words:

DATA FIRST

This lets us know you have read the full project description.

Generic AI-generated proposals without specific answers to the questions above will not be considered.

We care much more about technical thinking, data quality, architecture and problem-solving ability than a long list of technologies.
python sql web scraping postgresql data science data architecture automation
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.