Data Engineering and Cloud Transformation
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
AISHWARYA KHEKARE
Lead Data Engineer • 4.5 Years Experience • 5x Snowflake Certified
Results-driven Lead Data Engineer with 4.5+ years designing end-to-end cloud data platforms on Snowflake, dbt, and Apache
Airflow. Proven record of automating client workflows, building self-serve Streamlit apps, scaling dbt pipelines to 600+ clients, and
improving query performance by up to 95%. Expert in RBAC, data governance, CI/CD, and Snowpark-driven transformations across
AWS and Azure.
CORE TECHNICAL SKILLS
Cloud /Dataplatforms Snowflake, dbt (Core, Cloud), SnowSQL, Snowpark, Databricks (Lakehouse), AWS, Azure
ETL / Orchestration
Programming
BI & Applications
Apache Airflow, Fivetran, Prefect, Snowpipe, Streams & Tasks
Python, SQL, PySpark, Stored Procedures, UDFs, Snowpark Python, Pandas, NumPy
Tableau, Streamlit (Snowflake Native Apps), Power BI
DevOps & CI/CD
Governance & Security
Databases
Git, GitLab CI/CD, GitHub Actions, Agile, Query Optimization, Performance Tuning
RBAC, PII Masking, Row-Level Security, Data Quality Frameworks, Data Lineage, GDPR, CCPA
MySQL, Azure SQL Server, Snowflake
WORK EXPERIENCE
Snowflake Architecture & RBAC Setup
• Architected Snowflake from scratch for new clients - configured warehouses, databases, schemas, RBAC roles, and resource
monitors across dev/staging/prod environments with least-privilege access.
• Built Snowflake UDFs and stored procedures (Python & SQL) for complex business logic; migrated 300+ legacy SPs to
modular, testable dbt models with equivalent logic and better maintainability.
• Implemented Snowpipe + Streams + Tasks for event-driven, near-real-time ingestion pipelines; replaced batch jobs with
dynamic tables for continuous incremental processing.
• Designed and deployed Snowpark Python pipelines for in-warehouse data transformation, eliminating external compute
overhead.
• Configured clustering keys, search optimization, and materialized views - reducing average query execution time by 40–60%
across high-traffic tables.
• Enforced PII data masking policies, column-level security, and row-level access filters to meet GDPR/CCPA compliance
across all client environments.
dbt — Setup, Optimization & Macros
• Set up full dbt environments for 10+ clients from zero - profiles, sources, seeds, exposures, schema tests, and documentation;
enabled dbt Cloud CI jobs with PR-level test gates.
• Optimized dbt model execution by restructuring DAGs, switching materializations (view → incremental → table), and
eliminating multiple query patterns to achieving up to 95% reduction in run times.
• Built reusable dbt macros for SCD Type 2, dynamic schema dispatch, audit column injection, soft deletes, and cross-database
ref resolution — reducing model duplication by ~60%.
• Scaled a single-client dbt project to support 600+ clients using parameterized models, env-level vars, and dynamic source
definitions without per-client forks.
• Implemented dbt tests (schema, custom SQL, dbt-expectations) as a Data Quality (DQ) framework layer — automated checks
for completeness, uniqueness, referential integrity, and threshold breaches with Slack/email alerting.
Apache Airflow — End-to-End Setup & DAG Engineering
• Set up Apache Airflow infrastructure end-to-end (environment, connections, variables, pools, custom operators) and owned
DAG development for all client pipelines.
• Designed multi-stage DAGs: FiveTran REST API sync (MS SQL → Snowflake) → Snowflake procedure execution → dbt run
via Prefect → AAS model refresh via TMSL/REST API — fully automated with retry and failure alerting.
• Built Python operators with dynamic task mapping for client-parameterized workflows; implemented SLA monitoring, dead
letter queuing, and Slack alerts on task failure.
Streamlit App & Full Client Automation
• Built a Streamlit-on-Snowflake native application for finance clients — enabling self-serve reconciliation checks, pipeline status
views, and on-demand data refresh triggers; eliminated recurring manual support requests.
• Automated end-to-end client finance workflows (ingestion → transformation → reconciliation → reporting → alerting) using
Airflow + dbt + Snowflake stored procedures — cutting weekly manual effort by 10+ hours per client.
• Created a metadata-driven ETL framework handling dynamic source configs, transformation rules, logging, and anomaly alerts
— enabling zero-code onboarding of new data sources.
Cloud, CI/CD & Governance
• Integrated AWS (S3, Lambda, Fivetran, DMS) and Azure (ADF, Synapse, Blob) into unified Snowflake pipelines with
governance controls at each layer.
• Implemented GitLab CI/CD pipelines for dbt deployments — automated linting, testing, and promotion across dev → staging →
prod environments.
• Built a custom Python logging & alerting framework (email + Slack) with structured error codes, reducing incident resolution
time from 2 hours to 20 minutes.
• Conducted Snowflake cost optimization audits — auto-suspend tuning, warehouse right-sizing, and query pruning strategies,
reducing compute costs by 25%+.
Data Engineer | Persistent Systems — Pune, India | Jan 2022 – June 2024
• Established Snowflake infrastructure (RBAC, 3 warehouses, dev/staging/prod) and performed MySQL → Snowflake migration
via Fivetran — migrating 200+ GB with zero data loss.
• Developed dbt models for raw-to-consumption layer transformations; implemented SCD Type 2 across 50+ dimension tables;
added dbt tests for automated data validation.
• Built Snowflake stored procedures and Tasks for 2M+ daily record pipelines; implemented Snowpipe for streaming ingestion
from S3 and Azure Blob.
• Used Databricks (Lakehouse) for PySpark-based large-scale data processing and exploratory transformation jobs alongside
Snowflake pipelines.
• Orchestrated ETL workflows with Apache Airflow and Azure Data Factory; automated CI/CD deployments using GitLab
pipelines.
• Integrated Snowflake with Tableau — built real-time dashboards and optimized extract refresh schedules; migrated 25+
dashboards from Redshift to Snowflake with zero downtime.
• Configured Snowflake auto-scaling and auto-suspend policies, achieving 25% reduction in monthly compute costs.
• Applied PII data masking on 20+ sensitive tables; implemented data lineage tracking and query performance optimization via
clustering and pruning.
CERTIFICATIONS
• Snowflake SnowPro Advanced: Data Engineer
• Snowflake SnowPro Advanced: Architect
• Snowflake SnowPro Advanced: Data Analyst
• Snowpro Specialty: Gen AI
• Snowflake SnowPro Core
EDUCATION
• Microsoft Azure Data Fundamentals (DP-900)
• Microsoft Azure Fundamentals (AZ-900)
• Microsoft Azure AI Fundamentals (AI-900)
• Databricks Lakehouse Fundamentals
• dbt Fundamentals
B.Tech — Electronics & Telecommunications | KIT's College of Engineering, Kolhapur
2018 – 2022 • CGPA: 9.3 / 10
Lead Data Engineer • 4.5 Years Experience • 5x Snowflake Certified
Results-driven Lead Data Engineer with 4.5+ years designing end-to-end cloud data platforms on Snowflake, dbt, and Apache
Airflow. Proven record of automating client workflows, building self-serve Streamlit apps, scaling dbt pipelines to 600+ clients, and
improving query performance by up to 95%. Expert in RBAC, data governance, CI/CD, and Snowpark-driven transformations across
AWS and Azure.
CORE TECHNICAL SKILLS
Cloud /Dataplatforms Snowflake, dbt (Core, Cloud), SnowSQL, Snowpark, Databricks (Lakehouse), AWS, Azure
ETL / Orchestration
Programming
BI & Applications
Apache Airflow, Fivetran, Prefect, Snowpipe, Streams & Tasks
Python, SQL, PySpark, Stored Procedures, UDFs, Snowpark Python, Pandas, NumPy
Tableau, Streamlit (Snowflake Native Apps), Power BI
DevOps & CI/CD
Governance & Security
Databases
Git, GitLab CI/CD, GitHub Actions, Agile, Query Optimization, Performance Tuning
RBAC, PII Masking, Row-Level Security, Data Quality Frameworks, Data Lineage, GDPR, CCPA
MySQL, Azure SQL Server, Snowflake
WORK EXPERIENCE
Snowflake Architecture & RBAC Setup
• Architected Snowflake from scratch for new clients - configured warehouses, databases, schemas, RBAC roles, and resource
monitors across dev/staging/prod environments with least-privilege access.
• Built Snowflake UDFs and stored procedures (Python & SQL) for complex business logic; migrated 300+ legacy SPs to
modular, testable dbt models with equivalent logic and better maintainability.
• Implemented Snowpipe + Streams + Tasks for event-driven, near-real-time ingestion pipelines; replaced batch jobs with
dynamic tables for continuous incremental processing.
• Designed and deployed Snowpark Python pipelines for in-warehouse data transformation, eliminating external compute
overhead.
• Configured clustering keys, search optimization, and materialized views - reducing average query execution time by 40–60%
across high-traffic tables.
• Enforced PII data masking policies, column-level security, and row-level access filters to meet GDPR/CCPA compliance
across all client environments.
dbt — Setup, Optimization & Macros
• Set up full dbt environments for 10+ clients from zero - profiles, sources, seeds, exposures, schema tests, and documentation;
enabled dbt Cloud CI jobs with PR-level test gates.
• Optimized dbt model execution by restructuring DAGs, switching materializations (view → incremental → table), and
eliminating multiple query patterns to achieving up to 95% reduction in run times.
• Built reusable dbt macros for SCD Type 2, dynamic schema dispatch, audit column injection, soft deletes, and cross-database
ref resolution — reducing model duplication by ~60%.
• Scaled a single-client dbt project to support 600+ clients using parameterized models, env-level vars, and dynamic source
definitions without per-client forks.
• Implemented dbt tests (schema, custom SQL, dbt-expectations) as a Data Quality (DQ) framework layer — automated checks
for completeness, uniqueness, referential integrity, and threshold breaches with Slack/email alerting.
Apache Airflow — End-to-End Setup & DAG Engineering
• Set up Apache Airflow infrastructure end-to-end (environment, connections, variables, pools, custom operators) and owned
DAG development for all client pipelines.
• Designed multi-stage DAGs: FiveTran REST API sync (MS SQL → Snowflake) → Snowflake procedure execution → dbt run
via Prefect → AAS model refresh via TMSL/REST API — fully automated with retry and failure alerting.
• Built Python operators with dynamic task mapping for client-parameterized workflows; implemented SLA monitoring, dead
letter queuing, and Slack alerts on task failure.
Streamlit App & Full Client Automation
• Built a Streamlit-on-Snowflake native application for finance clients — enabling self-serve reconciliation checks, pipeline status
views, and on-demand data refresh triggers; eliminated recurring manual support requests.
• Automated end-to-end client finance workflows (ingestion → transformation → reconciliation → reporting → alerting) using
Airflow + dbt + Snowflake stored procedures — cutting weekly manual effort by 10+ hours per client.
• Created a metadata-driven ETL framework handling dynamic source configs, transformation rules, logging, and anomaly alerts
— enabling zero-code onboarding of new data sources.
Cloud, CI/CD & Governance
• Integrated AWS (S3, Lambda, Fivetran, DMS) and Azure (ADF, Synapse, Blob) into unified Snowflake pipelines with
governance controls at each layer.
• Implemented GitLab CI/CD pipelines for dbt deployments — automated linting, testing, and promotion across dev → staging →
prod environments.
• Built a custom Python logging & alerting framework (email + Slack) with structured error codes, reducing incident resolution
time from 2 hours to 20 minutes.
• Conducted Snowflake cost optimization audits — auto-suspend tuning, warehouse right-sizing, and query pruning strategies,
reducing compute costs by 25%+.
Data Engineer | Persistent Systems — Pune, India | Jan 2022 – June 2024
• Established Snowflake infrastructure (RBAC, 3 warehouses, dev/staging/prod) and performed MySQL → Snowflake migration
via Fivetran — migrating 200+ GB with zero data loss.
• Developed dbt models for raw-to-consumption layer transformations; implemented SCD Type 2 across 50+ dimension tables;
added dbt tests for automated data validation.
• Built Snowflake stored procedures and Tasks for 2M+ daily record pipelines; implemented Snowpipe for streaming ingestion
from S3 and Azure Blob.
• Used Databricks (Lakehouse) for PySpark-based large-scale data processing and exploratory transformation jobs alongside
Snowflake pipelines.
• Orchestrated ETL workflows with Apache Airflow and Azure Data Factory; automated CI/CD deployments using GitLab
pipelines.
• Integrated Snowflake with Tableau — built real-time dashboards and optimized extract refresh schedules; migrated 25+
dashboards from Redshift to Snowflake with zero downtime.
• Configured Snowflake auto-scaling and auto-suspend policies, achieving 25% reduction in monthly compute costs.
• Applied PII data masking on 20+ sensitive tables; implemented data lineage tracking and query performance optimization via
clustering and pruning.
CERTIFICATIONS
• Snowflake SnowPro Advanced: Data Engineer
• Snowflake SnowPro Advanced: Architect
• Snowflake SnowPro Advanced: Data Analyst
• Snowpro Specialty: Gen AI
• Snowflake SnowPro Core
EDUCATION
• Microsoft Azure Data Fundamentals (DP-900)
• Microsoft Azure Fundamentals (AZ-900)
• Microsoft Azure AI Fundamentals (AI-900)
• Databricks Lakehouse Fundamentals
• dbt Fundamentals
B.Tech — Electronics & Telecommunications | KIT's College of Engineering, Kolhapur
2018 – 2022 • CGPA: 9.3 / 10
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.