Data Engineer

CapgeminiCairo, CairoOn-siteFull-timeMid level, 2–5 yearsListed 2 hours ago

Apply now

About this role

Job Description

We are seeking a highly skilled and motivated Senior Data Engineer to join our Data & Analytics practice. The successful candidate will build and operate enterprise-scale data solutions using PySpark, Apache Spark and modern data engineering technologies, with a strong grounding in data warehousing, lakehouse architecture and large-scale data integration.

The target environment uses Cloudera Data Platform (CDP). Prior Cloudera experience is preferred but is not a mandatory requirement. Candidates with strong PySpark/Spark expertise and relevant experience on comparable enterprise data platforms will be considered, and Cloudera CDP upskilling will be provided to the selected candidate.

This role requires hands-on engineering capability, technical ownership and effective collaboration with architects, analysts, source-system teams and client stakeholders. Experience in banking and financial services is highly desirable.

Critical hiring priority: Strong, production-grade PySpark/Spark engineering skills and the ability to rapidly learn and work effectively within the Cloudera CDP environment.

##

Key Responsibilities

Data Engineering & Development

Design, develop, test and maintain scalable batch and streaming data pipelines using PySpark.

Build and optimize ETL/ELT processes for large-volume enterprise data workloads.

Develop reusable ingestion, transformation and validation frameworks.

Implement data quality controls, reconciliation checks, monitoring and operational logging.

Support dimensional, warehouse, data lake and lakehouse data modeling activities.

Apply Spark performance optimization techniques including partitioning, join optimization, caching and handling data skew.

Cloudera Platform Engineering

Develop and run Spark workloads within the Cloudera Data Platform (CDP) environment.

Work with components including Cloudera Data Engineering (CDE), Cloudera Data Warehouse (CDW), Hive, Impala, Ozone, Ranger, Atlas and NiFi.

Configure, execute, monitor and troubleshoot Spark workloads.

Performance-tune PySpark jobs, Spark configurations, SQL queries and Hive/Impala workloads.

Diagnose platform, ingestion, transformation, storage and workload execution issues.

Apply security, access-control, metadata and lineage practices using Ranger and Atlas.

Candidates without prior Cloudera experience will be expected to complete the required Cloudera enablement/upskilling and apply their existing Spark and data engineering knowledge to the platform.

Architecture, Governance & Delivery

Contribute to lakehouse implementations using modern table formats such as Apache Iceberg.

Apply engineering standards covering modular design, code quality, testing, version control and CI/CD.

Participate in requirements analysis, solution design, technical estimation and design reviews.

Support system integration testing, UAT, production deployment and operational readiness.

Produce clear technical documentation and deliver structured knowledge-transfer sessions.

Collaborate effectively with distributed, multicultural teams and technical and business stakeholders.

Candidate Profile

Education & Experience

Bachelor's degree in Computer Science, Information Systems, Engineering or a related discipline.

At least 4 years of professional experience in data engineering.

At least 3 years of hands-on PySpark / Apache Spark development experience.

Experience delivering large-scale enterprise data pipelines and ETL/ELT solutions.

Cloudera CDP experience is an advantage, but not mandatory. Strong candidates from comparable Spark-based data platforms are encouraged to apply.

##

Mandatory Technical Skills

✓ PySpark
✓ Apache Spark
✓ Python
✓ Advanced SQL and query optimization
✓ ETL/ELT development
✓ Data warehousing concepts
✓ Linux
✓ Git
✓ Data modeling fundamentals
✓ Experience working with large-scale enterprise data platforms

Preferred / Upskillable Skills

Cloudera CDP

Cloudera CDE and CDW

Hive and Impala

Ranger and Atlas

Apache Iceberg

Apache NiFi and Kafka

Apache Airflow

Denodo and Informatica

Oracle and Teradata

Azure data services

DevOps and CI/CD pipelines

Cloudera CDP knowledge is preferred, not mandatory. The selected candidate will be provided with Cloudera enablement/upskilling where required.

Preferred Domain Background

Banking and financial services

Digital banking platforms

Customer analytics

Regulatory reporting

Enterprise data warehousing

Data lakehouse implementation and migration programs

Behavioural Competencies

Strong analytical thinking and structured problem solving.

Ownership mindset with the ability to work independently and deliver reliably.

Clear written and verbal communication with technical and non-technical stakeholders.

Ability to mentor junior engineers and contribute to team capability building.

Ability and willingness to quickly learn new data platforms and technologies.

Comfort working in distributed and multicultural delivery teams.