About this role
Join a team that designs and develops scalable and secure distributed architectures and solutions, focusing on data ingestion and processing utilizing appropriate cloud native technologies and services.
As a Lead Data Engineer, within our Corporate Technology Team, you will design, implement, and maintain data pipelines that efficiently collect, process, and store large volumes of data from various sources, ensuring data timeliness, quality, and completeness, and ensure that data solutions comply with relevant data residency and privacy regulations, and implement best practices for securing data at rest and in transit in compliance with financial regulations and firm wide policies.
Job Qualifications:
- Design and develop scalable and secure distributed architectures and solutions, focusing on data ingestion and processing - utilizing appropriate cloud native technologies and services.
- Data pipeline development: Design, implement, and maintain data pipelines that efficiently collect, process, and store large volumes of data from various sources, ensuring data timeliness, quality, and completeness.
- Security and compliance: Ensure that data solutions comply with relevant data residency and privacy regulations, and implement best practices for securing data at rest and in transit in compliance with financial regulations and firm wide policies.
- Engages technical teams and business stakeholders to discuss and propose technical approaches to meet current and future needs
- Defines the technical target state of their product and drives achievement of the strategy
- Evaluates recommendations and provides feedback on new technologies
- Executes creative software solutions, design, and development
Required Qualifications, capabilities, and skills:
- Programming: Comfortable with Java/Python including sound testing and code review practices.
- SQL expertise: Joins, aggregations, subqueries, window functions
- Data pipelines: Design, build, and optimize production ETL/ELT pipelines (batch + streaming) using a popular framework (Spark, Flink, Dataflow, etc).
- Streaming: Hands-on with Kafka (topics, keys, partitions, consumer groups) at-least-once semantics, and schema registry basics.
- Warehousing/Lakehouse: Data modelling, partitioning, clustering. Hands-on with one of Snowflake, Databricks, etc, and cloud storage or HDFS.
- Cloud: Production experience with at least one major cloud provider (GCP/AWS) using native data services . FinOps-aware with cost-effective design.
- Reliability: Data quality checks, backfills, incorporating SLIs with observability and reporting and lakehouse platforms and table formats (Delta/Iceberg/Avro/Parquet) and time-travel.
Preferred qualifications, capabilities, and skills:
- Experience with Kafka, Flink, or other streaming technologies.
- Familiarity with AI/ML technologies including LLMs, prompt engineering, vector search, and responsible AI practices and experience using AI-assisted software development tools such as GitHub Copilot, Claude, or similar technologies.
- Financial services industry experience and understanding of large-scale enterprise data environments.
- Experience mentoring engineers and leading technical delivery initiatives.