Databricks Streaming Pipeline Setup
Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted1 hour ago
I’m ready to move my real-time data processing into Databricks and need a specialist who can build a robust streaming pipeline end-to-end. The source will be Apache Kafka; events arrive continuously and must be processed, transformed, and stored inside the Databricks environment with low latency and solid fault tolerance.
You’ll design the job structure, configure Structured Streaming, handle schema evolution, and implement checkpoints so that the pipeline restarts cleanly after failures. I also want concise notebook-based documentation that shows how the pipeline is triggered, how it scales, and where metrics can be monitored in the workspace.
Deliverables
• A working Databricks notebook (or repo) that ingests from our existing Kafka topics, performs necessary transformations, and writes the results to our designated target (Delta format inside our lakehouse).
• Automated cluster or job configuration files so the pipeline can be deployed to test and production workspaces without manual tweaks.
• README or in-notebook commentary that explains setup steps, configuration parameters, and how to extend the pipeline with additional topics in the future.
I’ll test the hand-off by running the notebook against a staged Kafka topic; acceptance is complete when new events appear in Delta within our agreed SLA and the pipeline passes a restart test. If you’ve tuned Structured Streaming against Kafka before and can demonstrate previous success with Databricks, let’s get started.
You’ll design the job structure, configure Structured Streaming, handle schema evolution, and implement checkpoints so that the pipeline restarts cleanly after failures. I also want concise notebook-based documentation that shows how the pipeline is triggered, how it scales, and where metrics can be monitored in the workspace.
Deliverables
• A working Databricks notebook (or repo) that ingests from our existing Kafka topics, performs necessary transformations, and writes the results to our designated target (Delta format inside our lakehouse).
• Automated cluster or job configuration files so the pipeline can be deployed to test and production workspaces without manual tweaks.
• README or in-notebook commentary that explains setup steps, configuration parameters, and how to extend the pipeline with additional topics in the future.
I’ll test the hand-off by running the notebook against a staged Kafka topic; acceptance is complete when new events appear in Delta within our agreed SLA and the pipeline passes a restart test. If you’ve tuned Structured Streaming against Kafka before and can demonstrate previous success with Databricks, let’s get started.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.