Quantexa Data Engineer

Unison GroupSingaporeOn-siteFull-timeStaff, 8–12 yearsListed 1 day ago

Apply now

About this role

Role Overview

We are seeking a talented and experienced Data Engineer with strong expertise in Quantexa, Hadoop, Scala, Apache Spark, Elasticsearch, OpenShift Container Platform (OCP), and DevOps practices .

The successful candidate will be responsible for designing, developing, optimizing, and maintaining scalable big data solutions using Apache Spark, Scala, Hadoop, and Elasticsearch . The role involves collaborating with cross-functional teams to build efficient data processing pipelines and search applications.

Knowledge and experience in the Compliance / AML domain will be an added advantage.

Key Responsibilities

- Design, develop, and implement Spark/Scala applications and data processing pipelines for large volumes of structured and unstructured data.
- Implement data transformation, aggregation, enrichment, and computation processes to support analytics and machine learning initiatives.
- Collaborate with cross-functional teams to understand data requirements and translate them into effective data engineering solutions.
- Integrate Elasticsearch with Spark for efficient data indexing, querying, and retrieval.
- Implement transformations and aggregations using Spark RDDs, DataFrames, Datasets, and Spark SQL .
- Develop scalable, reliable, and fault-tolerant Spark applications following industry best practices and coding standards.
- Optimize and tune Spark jobs and Elasticsearch queries to improve performance, scalability, and resource utilization.
- Monitor job performance, identify bottlenecks, troubleshoot issues, and implement appropriate optimizations.
- Troubleshoot and resolve issues related to data processing, data quality, Spark performance, and Elasticsearch integration .
- Ensure data quality, consistency, accuracy, and integrity throughout the data processing lifecycle.
- Design and deploy data engineering solutions on OpenShift Container Platform (OCP) using containerization and orchestration technologies.
- Optimize data engineering workflows for containerized environments and efficient resource utilization.
- Collaborate with DevOps teams to streamline deployments and implement CI/CD pipelines .
- Implement data governance, data lineage, and metadata management practices to ensure data accuracy, traceability, and compliance.
- Implement monitoring and logging mechanisms to ensure the health, availability, and performance of data infrastructure.
- Monitor and optimize end-to-end data pipeline performance and implement required enhancements.
- Document data engineering processes, workflows, architecture, and infrastructure configurations for knowledge sharing and future reference.

Requirements

- Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related field .
- Quantexa Certified Data Engineer / Data Architect with hands-on experience and strong proficiency in the Quantexa platform.
- Proven experience as a Data Engineer working with Hadoop, Spark, and large-scale data processing technologies.
- Strong proficiency in Scala and familiarity with functional programming concepts.
- In-depth understanding of Apache Spark architecture, RDDs, DataFrames, Datasets, and Spark SQL .
- Strong expertise in Hadoop ecosystem technologies , including HDFS, Hive, Pig, and related tools.
- Hands-on experience with Elasticsearch , including data indexing, search applications, data modeling, indexing strategies, and query optimization.
- Experience with OpenShift Container Platform (OCP) and Kubernetes-based container orchestration.
- Strong programming skills in Scala, Python, Java, and/or Spark .
- Good understanding of DevOps practices, CI/CD pipelines, and infrastructure automation .
- Experience with tools such as Docker, Jenkins, Ansible, and Bitbucket .
- Experience with distributed computing, parallel processing, and large-scale datasets .
- Strong experience in performance tuning and optimization of Spark applications and Elasticsearch queries.
- Experience with Git and collaborative software development workflows.
- Strong analytical and problem-solving skills with the ability to troubleshoot complex technical issues.
- Excellent communication and collaboration skills with the ability to work effectively with cross-functional teams.
- Experience with Grafana, Prometheus, and Splunk will be an added advantage.
- Exposure to cloud platforms such as AWS, Azure, or GCP and their data services will be a plus.
- Knowledge or experience in the Compliance / AML domain will be an added advantage.