Data Engineer

Fulcrum DigitalDublin, LeinsterOn-siteFull-timeSenior, 5–8 yearsListed 4 weeks ago

Apply now

About this role

Role Overview

We are looking for a highly skilled Data Quality Engineer with strong Data Engineering expertise to ensure the accuracy, reliability, and scalability of enterprise data platforms. The ideal candidate will possess hands-on experience with Databricks, PySpark, Hadoop, Hive, and Cloud technologies, along with advanced SQL skills to validate data across large-scale data pipelines and Lakehouse architectures.

Required Experience

- 3+ years of experience in Data Engineering, Data Quality Engineering, or Data Testing.
- Hands-on experience with Databricks and PySpark.
- Strong experience in Hadoop ecosystem components such as Hive, HDFS, Spark, and related Apache technologies.
- Advanced SQL expertise for large-scale data validation and analysis.
- Experience working with Data Warehouses, Data Lakes, and Lakehouse architectures.
- Understanding of Star Schema, Snowflake Schema, and dimensional modeling.
- Experience with cloud platforms (Azure, AWS, or GCP).

Requirements

Key Responsibilities

Data Quality & Validation

- Perform end-to-end validation of data pipelines across ingestion, transformation, and consumption layers.
- Execute source-to-target reconciliation and data quality checks.
- Identify, investigate, and resolve data anomalies and inconsistencies.
- Define and implement data quality frameworks, metrics, and controls.

Data Engineering & Processing

- Develop and validate data pipelines using PySpark and Databricks.
- Work with large-scale datasets in Hadoop, Hive, and Lakehouse environments.
- Support ETL/ELT workflows and ensure data integrity throughout the data lifecycle.
- Optimize data processing jobs for performance and scalability.

SQL & Analytics

- Write advanced SQL queries for data profiling, reconciliation, and root cause analysis.
- Perform complex joins, window functions, CTEs, aggregations, and query optimization.
- Validate business rules and transformation logic against source systems.

Lakehouse & Cloud Platforms

- Validate and monitor data across Databricks Lakehouse architecture.
- Work with cloud platforms such as Azure, AWS, or GCP.
- Collaborate with Data Engineers, Architects, and Analysts to ensure reliable data delivery.

Defect & Incident Management

- Analyze production issues and conduct root cause analysis.
- Track and manage data defects through resolution.
- Implement proactive monitoring and automated quality checks.

Required Experience

- 6+ years of experience in Data Engineering, Data Quality Engineering, or Data Testing.
- Hands-on experience with Databricks and PySpark.
- Strong experience in Hadoop ecosystem components such as Hive, HDFS, Spark, and related Apache technologies.
- Advanced SQL expertise for large-scale data validation and analysis.
- Experience working with Data Warehouses, Data Lakes, and Lakehouse architectures.
- Understanding of Star Schema, Snowflake Schema, and dimensional modeling.
- Experience with cloud platforms (Azure, AWS, or GCP).

Preferred Skills

- Automated data testing frameworks.
- Data observability and monitoring tools.
- CI/CD implementation for data pipelines.
- Experience with Delta Lake, Unity Catalog, or similar technologies.
- Knowledge of Airflow, Kafka, or other Apache ecosystem tools.