About this role
Role: Data Engineer – Microsoft Fabric
Location: India (Bangalore, Hyderabad / Mumbai / Gurugram)
Position: Senior Associate
Required Experience: 3–6 years
Role Summary
We’re looking for a skilled Data
Engineer with Microsoft Fabric experience to join our growing data and AI team.
In this role, you will design and build modern data platforms leveraging
Microsoft Fabric, enabling scalable analytics, AI-driven insights, and
enterprise-grade data solutions for global clients. This is an excellent
opportunity to work on next-generation data architectures, contribute to
AI-driven transformation programs, and grow into advanced data engineering and
platform leadership roles.
Key Responsibilities
• Design, build, and
maintain scalable data pipelines using Microsoft Fabric Data Factory
(Pipelines), Dataflows Gen2, Fabric Notebooks, and Lakehouse.
• Develop and manage
OneLake/Fabric Lakehouse architectures for structured and semi-structured
healthcare and enterprise data.
• Build and optimize batch
and API-based ingestion pipelines from multiple enterprise data sources,
databases, files, and application systems.
• Develop robust ETL/ELT
workflows using SQL, Python, and PySpark for data transformation and
enrichment.
• Implement data quality,
validation, reconciliation, completeness, consistency, and anomaly-detection
checks across source and target datasets.
• Develop data
standardization, deduplication, entity-resolution, record-linkage, and matching
workflows across heterogeneous data sources.
• Design analytics-ready
datasets and data models for data scientists, analysts, ML pipelines, and
downstream applications.
• Collaborate closely with
data scientists and ML engineers to prepare reliable feature-ready and
model-consumable datasets.
• Integrate Microsoft Fabric
with Azure data services such as Azure Data Lake Storage, Azure SQL, Synapse
components, and Power BI where required.
• Implement incremental
ingestion, change-data handling, error handling, retry mechanisms, and pipeline
recovery strategies where applicable.
• Implement observability
and monitoring using Azure Log Analytics, alerts, action groups, and
pipeline-level monitoring.
• Support data lineage,
metadata management, governance, security, and discoverability using Fabric
Catalog and/or Microsoft Purview.
• Optimize SQL queries,
Spark transformations, storage strategies, and pipeline execution for
large-scale datasets.
• Support CI/CD, Git-based
version control, deployment automation, and DataOps practices for Fabric data
solutions.
Required Qualifications
• Bachelor’s/Master’s degree
in Computer Science, Engineering, Data Engineering, or a related field, or
equivalent practical experience.
• 3–6 years of hands-on data
engineering experience, including practical Microsoft Fabric experience.
• Strong hands-on experience
with Microsoft Fabric Lakehouse, OneLake, Data Factory/Pipelines, Dataflows
Gen2, and Fabric Notebooks.
• Strong SQL skills,
including complex queries, data transformation, optimization, and data
modeling.
• Strong Python and/or
PySpark experience for data transformation and pipeline development.
• Experience designing and
implementing ETL/ELT pipelines and integrating data from APIs, databases,
files, and enterprise systems.
• Experience implementing
data quality, validation, reconciliation, and error-handling workflows.
• Experience with data
matching, entity resolution, deduplication, record linkage, or similar data
integration workflows.
• Working knowledge of Azure
Data Lake Storage, Azure SQL, and/or Synapse Analytics.
• Understanding of data
warehousing concepts, dimensional modeling, and analytics-ready data design.
• Familiarity with Git,
CI/CD, deployment automation, and DataOps practices.
• Understanding of data
governance, security, lineage, metadata, and performance optimization.
Preferred Qualifications
• Experience working with
healthcare/provider data.
• Experience integrating
Power BI with enterprise data platforms.
• Exposure to Microsoft
Purview or Fabric Catalog.
• Experience with
incremental/CDC ingestion and event-driven or near-real-time data pipelines.
• Exposure to Eventstream,
Azure Event Hubs, or other streaming ingestion frameworks.