About this role
Key Responsibilities
- Design develop and maintain scalable data pipelines using Azure Databricks and Azure Data Factory
- Develop ETLELT solutions for ingesting transforming and loading data from multiple sources
- Build and optimize PySpark applications and Spark SQL transformations for largescale data processing
- Create and manage Databricks notebooks workflows jobs and clusters
- Implement Delta Lakebased data solutions ensuring data quality reliability and performance
- Develop data ingestion frameworks using Azure Data Lake Storage ADLS Gen2
- Monitor troubleshoot and optimize data pipelines and Databricks workloads
- Implement CICD pipelines for Databricks artifacts using Azure DevOps or GitHub
- Collaborate with data architects business analysts and stakeholders to deliver enterprise data solutions
- Ensure security governance and compliance using Unity Catalog Azure Key Vault and related Azure services
Mandatory Skills
- Strong programming experience in Python
- Expertise in PySpark and Spark SQL
- Handson experience with Azure Databricks
- Experience with Azure Data Factory ADF
- Knowledge of Azure Data Lake Storage ADLS Gen2
- Strong SQL and data modeling skills
- Experience in ETLELT development and optimization
- Exposure to Delta Lake Databricks Workflows and Notebook development
- Version control using GitGitHub
- Understanding of performance tuning and troubleshooting Spark jobs
- GoodtoHave Skills
- Azure Synapse Analytics
- Unity Catalog
- Azure DevOps CICD
- Azure Key Vault
- Snowflake
- Data Governance and Data Quality frameworks
- Monitoring and Logging solutions in Azure
- CDC and Realtime Data Processing
- Preferred Certifications
- Microsoft Azure Data Engineer Associate DP203
- Databricks Data Engineer Associate