Databricks Data Specialist - R01569707

BrillioBengaluru, KarnatakaOn-siteFull-timeJunior, 1–2 yearsListed 1 month ago

Apply now

About this role

Data Specialist

Primary Skills

Databricks Engineer

Role Overview

We are seeking a highly skilled Databricks Engineer to design, develop, and optimize scalable data engineering and analytics solutions on the Databricks Lakehouse Platform . The ideal candidate will possess strong expertise in Databricks, PySpark, and SQL , with hands-on experience building batch and real-time data pipelines, implementing Lakehouse architectures, and ensuring data governance and performance optimization.

Key Responsibilities

- Design, develop, and maintain end-to-end data pipelines using Databricks and PySpark .

- Build and implement Lakehouse architectures utilizing Bronze, Silver, and Gold data layers.

- Develop and manage Delta Lake solutions with ACID transactions, schema enforcement, and data reliability features.

- Create, monitor, and optimize Delta Live Tables (DLT) pipelines.

- Implement scalable and efficient data ingestion processes using Auto Loader .

- Develop and manage real-time data processing solutions using Structured Streaming .

- Orchestrate, schedule, and monitor data workflows using Databricks Workflows .

- Design and implement Lakehouse data models to support reporting, analytics, and business intelligence requirements.

- Establish and enforce data governance, security, and access controls using Unity Catalog .

- Optimize Spark jobs, SQL queries, and overall platform performance to ensure efficiency and scalability.

- Collaborate with cross-functional teams, including data analysts, architects, and business stakeholders, to deliver high-quality data solutions.

Required Skills (Must Have)

- Databricks Platform
- Delta Lake

- Delta Live Tables (DLT)

- Unity Catalog

- Databricks Workflows

- PySpark and Apache Spark

- Structured Streaming

- Auto Loader

- SQL

- Lakehouse Data Modeling

- Strong understanding of data engineering best practices and scalable data architectures

Preferred Skills (Good to Have)

Azure Ecosystem

- Azure Data Factory (ADF)

- Azure Synapse Analytics

- Microsoft Purview

- Microsoft Fabric

AWS Ecosystem

- AWS Glue

- AWS Lambda

- AWS Step Functions

Data Engineering & Integration

- Apache Airflow

- DBT

- Fivetran

- Informatica

Streaming & Analytics

- Apache Kafka

- Power BI

Data Governance

- Collibra

- Alation

GCP

- BigQuery

Qualifications

- Bachelor's or Master's degree in Computer Science, Data Engineering, Information Technology, or a related discipline.

- Proven experience in designing and implementing cloud-based data engineering solutions and scalable data pipelines.

- Strong analytical, troubleshooting, and problem-solving capabilities.

- Experience working in agile and collaborative environments.

- Excellent communication and stakeholder management skills.

Preferred Candidate Profile

- Hands-on experience with modern Lakehouse architectures and enterprise-scale data platforms.

- Strong understanding of data governance, security, and compliance frameworks.

- Experience delivering both batch and real-time data processing solutions.

- Ability to work independently while collaborating effectively across global teams.

Key Technologies

Databricks | PySpark | Apache Spark | Delta Lake | Delta Live Tables (DLT) | Unity Catalog | Structured Streaming | Auto Loader | SQL | Lakehouse Architecture | Azure | AWS | Airflow | Kafka | Power BI

Specialization

- Databricks Engineering: Lead Data Engineer

Job requirements

Databricks Engineer

Role Overview

We are seeking a highly skilled Databricks Engineer to design, develop, and optimize scalable data engineering and analytics solutions on the Databricks Lakehouse Platform . The ideal candidate will possess strong expertise in Databricks, PySpark, and SQL , with hands-on experience building batch and real-time data pipelines, implementing Lakehouse architectures, and ensuring data governance and performance optimization.

Key Responsibilities

- Design, develop, and maintain end-to-end data pipelines using Databricks and PySpark .

- Build and implement Lakehouse architectures utilizing Bronze, Silver, and Gold data layers.

- Develop and manage Delta Lake solutions with ACID transactions, schema enforcement, and data reliability features.

- Create, monitor, and optimize Delta Live Tables (DLT) pipelines.

- Implement scalable and efficient data ingestion processes using Auto Loader .

- Develop and manage real-time data processing solutions using Structured Streaming .

- Orchestrate, schedule, and monitor data workflows using Databricks Workflows .

- Design and implement Lakehouse data models to support reporting, analytics, and business intelligence requirements.

- Establish and enforce data governance, security, and access controls using Unity Catalog .

- Optimize Spark jobs, SQL queries, and overall platform performance to ensure efficiency and scalability.

- Collaborate with cross-functional teams, including data analysts, architects, and business stakeholders, to deliver high-quality data solutions.

Required Skills (Must Have)

- Databricks Platform
- Delta Lake

- Delta Live Tables (DLT)

- Unity Catalog

- Databricks Workflows

- PySpark and Apache Spark

- Structured Streaming

- Auto Loader

- SQL

- Lakehouse Data Modeling

- Strong understanding of data engineering best practices and scalable data architectures

Preferred Skills (Good to Have)

Azure Ecosystem

- Azure Data Factory (ADF)

- Azure Synapse Analytics

- Microsoft Purview

- Microsoft Fabric

AWS Ecosystem

- AWS Glue

- AWS Lambda

- AWS Step Functions

Data Engineering & Integration

- Apache Airflow

- DBT

- Fivetran

- Informatica

Streaming & Analytics

- Apache Kafka

- Power BI

Data Governance

- Collibra

- Alation

GCP

- BigQuery

Qualifications

- Bachelor's or Master's degree in Computer Science, Data Engineering, Information Technology, or a related discipline.

- Proven experience in designing and implementing cloud-based data engineering solutions and scalable data pipelines.

- Strong analytical, troubleshooting, and problem-solving capabilities.

- Experience working in agile and collaborative environments.

- Excellent communication and stakeholder management skills.

Preferred Candidate Profile

- Hands-on experience with modern Lakehouse architectures and enterprise-scale data platforms.

- Strong understanding of data governance, security, and compliance frameworks.

- Experience delivering both batch and real-time data processing solutions.

- Ability to work independently while collaborating effectively across global teams.

Key Technologies

Databricks | PySpark | Apache Spark | Delta Lake | Delta Live Tables (DLT) | Unity Catalog | Structured Streaming | Auto Loader | SQL | Lakehouse Architecture | Azure | AWS | Airflow | Kafka | Power BI