AI-Driven Marine Data Aggregation & Analysis
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
Senior Data / AI Engineer - Marine Data, APIs, Web Scraping & Machine Learning
We are building a new AI-powered travel technology platform that works with large amounts of marine, oceanographic, weather, environmental, geographic and destination data.
We are looking for a strong Data / AI / Machine Learning Engineer to help build the data and intelligence infrastructure behind the platform.
This is not primarily a frontend or UI development role.
We need someone who is very comfortable working with large datasets, APIs, web scraping, automated data pipelines, geospatial data, historical and forecast data, oceanographic datasets, and scoring/recommendation models.
Experience working with marine weather, oceanographic, coastal or environmental data is a major advantage.
What you will work on
You will help us:
* Collect and structure data from multiple public and commercial data sources
* Research and identify reliable marine, weather, environmental and geographic data sources
* Build reliable web scraping and automated data collection pipelines
* Integrate multiple third-party APIs
* Work with oceanographic, marine weather, environmental and location-based datasets
* Collect historical, seasonal, real-time and forecast data
* Clean, normalize, deduplicate and validate large datasets
* Design scalable database structures for destinations and geographic locations
* Match information from different sources to the correct physical locations
* Build automated systems to keep data continuously updated
* Develop scoring, ranking and recommendation logic
* Transform complex raw data into structured information that can be used by a consumer application
* Build data-quality checks, anomaly detection and confidence systems
* Document data sources, transformations, assumptions and limitations
* Work closely with the founder and development team to translate product requirements into reliable data infrastructure
Technical skills we are looking for
Strong experience with:
* Python
* SQL
* PostgreSQL
* APIs / REST APIs
* Web scraping
* ETL / ELT pipelines
* Data cleaning and normalization
* Pandas / NumPy
* Large structured datasets
* Automation
* Scheduled data pipelines
* Geospatial data and coordinate matching
* Git / GitHub
Experience with some of the following is a major advantage:
* Machine learning
* Recommendation engines
* Ranking and scoring algorithms
* Time-series data
* Marine weather data
* Oceanographic datasets
* Environmental APIs
* Meteorological data
* Historical and forecast datasets
* GIS / PostGIS
* Geospatial analysis
* Open-Meteo or similar environmental data services
* Supabase
* AWS / GCP or other cloud infrastructure
* Data validation and anomaly detection
* LLM / AI integrations
* Vector databases or embeddings
The person we want
We are not looking for someone who simply follows instructions or writes individual scripts.
We want someone who can look at a difficult data problem and say:
“Here is how I would structure this, here are the best sources, here is what data we can trust, here is what is missing, and here is how I would automate the entire system.”
You should be:
* Highly analytical
* Extremely organized with data
* Comfortable researching unfamiliar datasets and APIs
* Good at identifying unreliable, incomplete or conflicting information
* Able to design systems, not just scripts
* Proactive in suggesting better technical approaches
* Very careful about data accuracy
* Comfortable working independently
* Able to explain technical decisions clearly to a non-data-specialist founder
* Interested in potentially working with us longer term if the initial project goes well
Data quality is extremely important
We do not want someone who simply collects thousands of records and considers the project finished.
Accuracy, traceability and maintainability matter more than raw volume.
Important data should ideally have:
* A known source
* A clear definition
* Source attribution
* A timestamp or update frequency where relevant
* Validation rules
* Confidence indicators where appropriate
* A process for detecting stale or incorrect information
The system needs to be maintainable and scalable as the platform expands internationally.
Initial project
The first stage will involve auditing an existing database and data architecture.
You will:
1. Review the existing database, schema and datasets
2. Identify missing, unreliable, duplicated or incorrectly structured data
3. Recommend improvements to the data architecture
4. Research and identify the best available data sources and APIs
5. Determine which sources should be used for which types of information
6. Build or improve automated data collection pipelines
7. Normalize, validate and document the collected information
8. Help develop the first version of our scoring and recommendation data engine
9. Create a scalable foundation for future data sources and models
If the collaboration is successful, this can become a significant ongoing role as the platform grows.
When applying
Please answer the following questions directly:
1. Tell us about a project where you collected and combined data from multiple APIs, datasets or websites. What was technically difficult about it?
2. Have you worked with marine, oceanographic, weather, environmental, geographic or other time-series data? Please explain exactly what you worked with.
3. How would you approach matching information from several different sources when the same physical location may have different names, coordinates or IDs?
4. How do you determine whether information collected from an API, open dataset or scraped source is reliable enough to use in a production system?
5. Have you built a scoring, ranking or recommendation system before? Please describe the logic and your role.
6. How would you architect a system that collects information from multiple sources and automatically updates different datasets at different frequencies?
7. How would you track data provenance so we always know where an individual piece of information originated?
8. Please include links to relevant GitHub repositories, technical projects or examples of previous work if available.
Start your proposal with the words:
DATA FIRST
This lets us know you have read the full project description.
Generic AI-generated proposals without specific answers to the questions above will not be considered.
We care much more about technical thinking, data quality, architecture and problem-solving ability than a long list of technologies.
We are building a new AI-powered travel technology platform that works with large amounts of marine, oceanographic, weather, environmental, geographic and destination data.
We are looking for a strong Data / AI / Machine Learning Engineer to help build the data and intelligence infrastructure behind the platform.
This is not primarily a frontend or UI development role.
We need someone who is very comfortable working with large datasets, APIs, web scraping, automated data pipelines, geospatial data, historical and forecast data, oceanographic datasets, and scoring/recommendation models.
Experience working with marine weather, oceanographic, coastal or environmental data is a major advantage.
What you will work on
You will help us:
* Collect and structure data from multiple public and commercial data sources
* Research and identify reliable marine, weather, environmental and geographic data sources
* Build reliable web scraping and automated data collection pipelines
* Integrate multiple third-party APIs
* Work with oceanographic, marine weather, environmental and location-based datasets
* Collect historical, seasonal, real-time and forecast data
* Clean, normalize, deduplicate and validate large datasets
* Design scalable database structures for destinations and geographic locations
* Match information from different sources to the correct physical locations
* Build automated systems to keep data continuously updated
* Develop scoring, ranking and recommendation logic
* Transform complex raw data into structured information that can be used by a consumer application
* Build data-quality checks, anomaly detection and confidence systems
* Document data sources, transformations, assumptions and limitations
* Work closely with the founder and development team to translate product requirements into reliable data infrastructure
Technical skills we are looking for
Strong experience with:
* Python
* SQL
* PostgreSQL
* APIs / REST APIs
* Web scraping
* ETL / ELT pipelines
* Data cleaning and normalization
* Pandas / NumPy
* Large structured datasets
* Automation
* Scheduled data pipelines
* Geospatial data and coordinate matching
* Git / GitHub
Experience with some of the following is a major advantage:
* Machine learning
* Recommendation engines
* Ranking and scoring algorithms
* Time-series data
* Marine weather data
* Oceanographic datasets
* Environmental APIs
* Meteorological data
* Historical and forecast datasets
* GIS / PostGIS
* Geospatial analysis
* Open-Meteo or similar environmental data services
* Supabase
* AWS / GCP or other cloud infrastructure
* Data validation and anomaly detection
* LLM / AI integrations
* Vector databases or embeddings
The person we want
We are not looking for someone who simply follows instructions or writes individual scripts.
We want someone who can look at a difficult data problem and say:
“Here is how I would structure this, here are the best sources, here is what data we can trust, here is what is missing, and here is how I would automate the entire system.”
You should be:
* Highly analytical
* Extremely organized with data
* Comfortable researching unfamiliar datasets and APIs
* Good at identifying unreliable, incomplete or conflicting information
* Able to design systems, not just scripts
* Proactive in suggesting better technical approaches
* Very careful about data accuracy
* Comfortable working independently
* Able to explain technical decisions clearly to a non-data-specialist founder
* Interested in potentially working with us longer term if the initial project goes well
Data quality is extremely important
We do not want someone who simply collects thousands of records and considers the project finished.
Accuracy, traceability and maintainability matter more than raw volume.
Important data should ideally have:
* A known source
* A clear definition
* Source attribution
* A timestamp or update frequency where relevant
* Validation rules
* Confidence indicators where appropriate
* A process for detecting stale or incorrect information
The system needs to be maintainable and scalable as the platform expands internationally.
Initial project
The first stage will involve auditing an existing database and data architecture.
You will:
1. Review the existing database, schema and datasets
2. Identify missing, unreliable, duplicated or incorrectly structured data
3. Recommend improvements to the data architecture
4. Research and identify the best available data sources and APIs
5. Determine which sources should be used for which types of information
6. Build or improve automated data collection pipelines
7. Normalize, validate and document the collected information
8. Help develop the first version of our scoring and recommendation data engine
9. Create a scalable foundation for future data sources and models
If the collaboration is successful, this can become a significant ongoing role as the platform grows.
When applying
Please answer the following questions directly:
1. Tell us about a project where you collected and combined data from multiple APIs, datasets or websites. What was technically difficult about it?
2. Have you worked with marine, oceanographic, weather, environmental, geographic or other time-series data? Please explain exactly what you worked with.
3. How would you approach matching information from several different sources when the same physical location may have different names, coordinates or IDs?
4. How do you determine whether information collected from an API, open dataset or scraped source is reliable enough to use in a production system?
5. Have you built a scoring, ranking or recommendation system before? Please describe the logic and your role.
6. How would you architect a system that collects information from multiple sources and automatically updates different datasets at different frequencies?
7. How would you track data provenance so we always know where an individual piece of information originated?
8. Please include links to relevant GitHub repositories, technical projects or examples of previous work if available.
Start your proposal with the words:
DATA FIRST
This lets us know you have read the full project description.
Generic AI-generated proposals without specific answers to the questions above will not be considered.
We care much more about technical thinking, data quality, architecture and problem-solving ability than a long list of technologies.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.