Senior Data Engineer (Databricks Migration)
TypeFull-time job
LocationPortugal
Posted1 hour ago
Participate in the migration of a large-scale analytical platform from BigQuery to Databricks
Design and implement scalable Lakehouse architectures using Databricks and Delta Lake
Analyze existing ETL / ELT workloads and define migration approaches
Develop and optimize data pipelines processing large volumes of retail and analytical data
Implement incremental processing strategies and scalable transformation frameworks
Build and maintain Spark-based data processing solutions using PySpark
Design and maintain medallion architecture layers including Bronze, Silver, and Gold
Implement data governance and security best practices using Unity Catalog
Collaborate with Data Science, Analytics, Product, and Customer Engineering teams
Participate in architecture discussions and technical solution design
Develop reusable data platform components and engineering standards
Conduct code reviews and contribute to platform reliability and maintainability
Troubleshoot and optimize complex SQL and Spark workloads
Support production deployments and platform modernization activities
5+ years of professional experience as a Data Engineer
Strong programming skills in Python and advanced SQL
Hands-on commercial experience with Databricks
Strong knowledge of Apache Spark, primarily PySpark
Experience designing and building modern cloud-based data platforms
Experience developing ETL / ELT pipelines and large-scale data processing solutions
Hands-on experience with Delta Lake
Experience with Spark Declarative Pipelines
Experience with cluster monitoring, metrics analysis, and performance optimization
Strong understanding of distributed data processing architectures
Solid understanding of data warehousing concepts and dimensional modeling
Experience with Airflow or similar orchestration tools
Experience optimizing complex analytical SQL workloads
Experience implementing CI / CD practices for data engineering platforms
Strong troubleshooting and performance optimization skills
Ability to work collaboratively in cross-functional international teams
Upper-Intermediate or higher English level
WILL BE A PLUS
Experience working with GCP cloud services
Experience with AWS or Azure cloud platforms
Experience in retail analytics or pricing optimization domains
Experience supporting machine learning or AI-related data workloads
Experience with platform modernization and cloud migration initiatives
PERSONAL PROFILE
Strong analytical and problem-solving mindset
Proactive and ownership-driven approach
Ability to work independently and collaboratively
Good communication and stakeholder collaboration skills
Passion for scalable data engineering and modern data platforms
Interest in continuous learning and technology innovation
Are you a Senior Data Engineer passionate about building scalable, high-performance data platforms and working with modern Lakehouse technologies? Join Sigma Software’s Data Engineering Center of Excellence and contribute to the modernization of an enterprise-scale analytics ecosystem for the retail domain.
We are looking for a Senior specialist with strong Databricks, PySpark, and cloud data engineering expertise to participate in the migration of a large-scale analytical platform from BigQuery to Databricks. You will collaborate with international teams, contribute to architectural decisions, and help shape reliable and scalable data solutions.
We at Sigma Software create opportunities for continuous learning, technology growth, and meaningful engineering impact while working on complex international projects.
CUSTOMER
Our Customer is a leading retail technology company specializing in AI-driven pricing optimization solutions for enterprise retailers. The company helps businesses improve profitability and competitiveness through advanced analytics, automation, and intelligent pricing strategies. Their platform combines business intelligence with sophisticated algorithms to support data-informed pricing decisions at scale for global retail organizations.
PROJECT
The project focuses on the strategic migration of a large-scale analytical platform from a legacy BigQuery ecosystem to a modern Databricks Lakehouse architecture. The platform processes high-volume retail datasets, machine learning workloads, analytics pipelines, and customer-specific business logic.
As part of the modernization initiative, the engineering team is implementing scalable Spark-based processing, Delta Lake architecture, medallion data layers, and modern governance practices. The role offers an opportunity to work with distributed data processing systems, optimize large-scale workloads, and contribute to the evolution of an enterprise-grade data platform.
Key Technologies: Databricks, Apache Spark, PySpark, Delta Lake, Python, SQL, Airflow, GCP, CI/CD, Unity Catalog
Originally posted on Himalayas
Design and implement scalable Lakehouse architectures using Databricks and Delta Lake
Analyze existing ETL / ELT workloads and define migration approaches
Develop and optimize data pipelines processing large volumes of retail and analytical data
Implement incremental processing strategies and scalable transformation frameworks
Build and maintain Spark-based data processing solutions using PySpark
Design and maintain medallion architecture layers including Bronze, Silver, and Gold
Implement data governance and security best practices using Unity Catalog
Collaborate with Data Science, Analytics, Product, and Customer Engineering teams
Participate in architecture discussions and technical solution design
Develop reusable data platform components and engineering standards
Conduct code reviews and contribute to platform reliability and maintainability
Troubleshoot and optimize complex SQL and Spark workloads
Support production deployments and platform modernization activities
5+ years of professional experience as a Data Engineer
Strong programming skills in Python and advanced SQL
Hands-on commercial experience with Databricks
Strong knowledge of Apache Spark, primarily PySpark
Experience designing and building modern cloud-based data platforms
Experience developing ETL / ELT pipelines and large-scale data processing solutions
Hands-on experience with Delta Lake
Experience with Spark Declarative Pipelines
Experience with cluster monitoring, metrics analysis, and performance optimization
Strong understanding of distributed data processing architectures
Solid understanding of data warehousing concepts and dimensional modeling
Experience with Airflow or similar orchestration tools
Experience optimizing complex analytical SQL workloads
Experience implementing CI / CD practices for data engineering platforms
Strong troubleshooting and performance optimization skills
Ability to work collaboratively in cross-functional international teams
Upper-Intermediate or higher English level
WILL BE A PLUS
Experience working with GCP cloud services
Experience with AWS or Azure cloud platforms
Experience in retail analytics or pricing optimization domains
Experience supporting machine learning or AI-related data workloads
Experience with platform modernization and cloud migration initiatives
PERSONAL PROFILE
Strong analytical and problem-solving mindset
Proactive and ownership-driven approach
Ability to work independently and collaboratively
Good communication and stakeholder collaboration skills
Passion for scalable data engineering and modern data platforms
Interest in continuous learning and technology innovation
Are you a Senior Data Engineer passionate about building scalable, high-performance data platforms and working with modern Lakehouse technologies? Join Sigma Software’s Data Engineering Center of Excellence and contribute to the modernization of an enterprise-scale analytics ecosystem for the retail domain.
We are looking for a Senior specialist with strong Databricks, PySpark, and cloud data engineering expertise to participate in the migration of a large-scale analytical platform from BigQuery to Databricks. You will collaborate with international teams, contribute to architectural decisions, and help shape reliable and scalable data solutions.
We at Sigma Software create opportunities for continuous learning, technology growth, and meaningful engineering impact while working on complex international projects.
CUSTOMER
Our Customer is a leading retail technology company specializing in AI-driven pricing optimization solutions for enterprise retailers. The company helps businesses improve profitability and competitiveness through advanced analytics, automation, and intelligent pricing strategies. Their platform combines business intelligence with sophisticated algorithms to support data-informed pricing decisions at scale for global retail organizations.
PROJECT
The project focuses on the strategic migration of a large-scale analytical platform from a legacy BigQuery ecosystem to a modern Databricks Lakehouse architecture. The platform processes high-volume retail datasets, machine learning workloads, analytics pipelines, and customer-specific business logic.
As part of the modernization initiative, the engineering team is implementing scalable Spark-based processing, Delta Lake architecture, medallion data layers, and modern governance practices. The role offers an opportunity to work with distributed data processing systems, optimize large-scale workloads, and contribute to the evolution of an enterprise-grade data platform.
Key Technologies: Databricks, Apache Spark, PySpark, Delta Lake, Python, SQL, Airflow, GCP, CI/CD, Unity Catalog
Originally posted on Himalayas
Apply on Himalayas →
Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.