About this role
1.Lead the design and implementation of scalable Databricks Lakehouse data solutions supporting analytics and business applications.
2. Design and optimize high-volume batch and near-real-time data pipelines using PySpark, SQL, Python, Delta Lake, and Databricks workflows.
3. Design and develop end to end frameworks using patterns such as CDC, incremental processing, deduplication, SCD, and schema evolution.
4. Design data models and processing frameworks that provide reliable, performant, and reusable datasets for analytics and downstream applications.
5. Contribute to data governance, security, access control, and lineage using Databricks and Unity Catalog.
6. Drive engineering best practices including Git, code reviews, automated testing, CI/CD, documentation, and production support.
7. Collaborate with data architects, analysts, BI teams, data scientists, and business stakeholders to translate requirements into scalable data solutions.
8.Experience with Airflow or other modern orchestration tools.
9. Mentor intermediate engineers and provide technical guidance on data-engineering practices and platform capabilities.
Programming & Data Engineering
• Strong Python and advanced SQL skills.
• Strong hands-on experience with Apache Spark/PySpark.
• Strong hands-on experience with Databricks and Delta Lake.
Strong hands on experience in creating data models and data governance framework
• Strong understanding of batch data processing and streaming concepts.
• Strong knowledge of CDC, incremental processing, deduplication, SCD, and schema evolution.
• Experience designing scalable and maintainable data pipelines and data models.
• Good analytical and debugging skill
Databricks
• Experience designing and implementing Databricks-based data solutions.
• Experience with Databricks Jobs/workflows and production pipeline management.
• Experience with Databricks Lakeflow Pipelines, Lakeflow Connect, Auto Loader, or Structured Streaming.
• Exposure to Terraform or infrastructure-as-code practices.
• Experience with data observability, pipeline monitoring, and cloud cost/performance optimization.
• Experience modernizing or migrating legacy ETL/data warehouse workloads to Databricks.
• Exposure to Data Mesh, Data Products, or domain-oriented data architecture.
• Exposure to Docker/Kubernetes or other cloud-native technologies.
AI / GenAI
• Experience designing data platforms or pipelines supporting ML/AI/GenAI workloads.
• Experience using AI-powered development tools or GenAI coding assistants.
• Exposure to Databricks AI/ML capabilities or equivalent cloud AI platforms
