About this role
Key Responsibilities:
Data Engineering & Pipeline Development
- Design, develop, and maintain scalable data pipelines using Databricks, PySpark, Python, and SQL .
- Build ETL/ELT pipelines to ingest data from internal and external source systems.
- Load and manage data in Databricks Delta Tables .
- Handle data from multiple and evolving source systems.
- Implement data validation, quality checks, error handling, and monitoring.
- Troubleshoot and optimize data pipelines for performance, reliability, and scalability.
Data Transformation & Modeling
- Analyze source data and business requirements to design appropriate data transformations.
- Develop standardized and curated datasets for reporting, dashboards, and analytics.
- Design and maintain logical and physical data models .
- Apply dimensional modeling concepts such as fact and dimension tables, star schema, and normalization/denormalization .
- Ensure data models are scalable and aligned with business requirements.
Solution Architecture
- Design end-to-end data architectures and data movement strategies.
- Evaluate source systems and integration patterns and recommend appropriate solutions.
- Design scalable, secure, maintainable, and reliable data solutions.
- Create technical documentation covering architecture, data flows, mappings, and integration processes.
- Contribute to reusable data engineering frameworks and best practices.
Business & Stakeholder Management
- Work directly with business users, data consumers, and technical teams to understand and clarify requirements.
- Translate business requirements into technical specifications and data models.
- Understand business processes and determine appropriate data transformation and architecture approaches.
- Communicate technical solutions effectively to both technical and non-technical stakeholders.
Data Governance & Quality
- Implement data quality and validation frameworks.
- Follow data governance, naming conventions, taxonomy, metadata, and lineage standards.
- Ensure data accuracy, consistency, traceability, and auditability.
- Support enterprise data management and governance initiatives.
Key Skills – Must Have:
- Databricks – Strong hands-on experience
- PySpark – Strong
- Python
- SQL – Strong
- Delta Lake / Delta Tables
- ETL / ELT
- Data Engineering
- Data Modeling
- Data Architecture / Solution Design
- Data Pipeline Development
- Data Transformation and Integration
- Performance Tuning and Optimization
- Data Quality and Validation
- Strong analytical and problem-solving skills
- Strong stakeholder communication skills
Good-to-Have Skills
- AWS / Azure
- Azure Data Factory / AWS Glue
- Power BI / Tableau
- Kafka / Streaming
- CI/CD and DevOps
- Git / Azure DevOps
- Databricks Unity Catalog
- Data Governance
- Metadata and Data Lineage
- Data Taxonomy
- Enterprise Data Standards
Experience:
- 8+ years of overall IT experience
- Strong recent hands-on experience in Databricks and PySpark
- Experience working on enterprise-scale data engineering projects
- Experience designing end-to-end data solutions
- Experience interacting directly with business and technical stakeholders
Work Details:
Work Location: Hyderabad
Work Mode: Hybrid – 3 days per week from office
Working Hours: 2:00PM to 11:00 PM.
Role Type: Contract