Senior Data Engineer

SRM TechnologiesRemoteFull-timeStaff, 8–12 yearsListed 3 days ago

Apply now

About this role

This is a remote position.

Summary:

Role: Senior Data Engineer

Experience: 8+ Years

Mandatory/Core: Python, PySpark, Snowflake, dbt, Apache Iceberg, AWS, SQL

Preferred: AWS Glue, S3, EMR, Lambda, Airflow, Snowpipe/Snowpark, CI/CD, Terraform, Data Modeling

Role Type: Senior hands-on Data Engineer

Focus: Cloud Data Engineering, Lakehouse, Data Transformation, Performance Optimization and Production Engineering

Detailed information:

Senior Data Engineer:

Experience: 8+ years of overall Data Engineering experience , with strong hands-on experience building enterprise-scale cloud data platforms and pipelines.

Primary Skills:

- Python

- PySpark / Apache Spark

- Snowflake

- dbt (Data Build Tool)

- Apache Iceberg

- AWS Data Services

- Advanced SQL

- Data Engineering / ETL / ELT

- Data Lake / Lakehouse architecture

Secondary / Preferred Skills:

- AWS services such as:

- S3

- AWS Glue

- EMR

- Lambda

- Step Functions

- CloudWatch

- IAM

- Apache Airflow or other workflow orchestration tools

- Snowflake performance optimization and cost optimization

- Snowpipe / Snowpark

- Spark performance tuning

- Data modeling and dimensional modeling

- Parquet and other columnar data formats

- Data quality frameworks and automated validation

- CI/CD for data pipelines

- Git / GitHub / GitLab

- Infrastructure as Code such as Terraform or AWS CDK

- Docker / containerization

- Data governance, lineage, security, and access control

- Agile/Scrum delivery experience

Job Description:

We are looking for a Senior Data Engineer with strong hands-on expertise in Python, PySpark, Snowflake, dbt, Apache Iceberg, and AWS to design, develop, and maintain scalable enterprise data solutions.

The candidate should have strong experience working with high-volume data processing, cloud-based data platforms, modern lakehouse architectures, ETL/ELT pipelines, data modeling, performance optimization, and production-grade engineering practices.

The ideal candidate should be capable of independently owning complex data-engineering components, contributing to technical design and architecture decisions, troubleshooting production issues, and providing technical guidance to other engineers.

Key Responsibilities:

1. Data Pipeline Engineering

- Design, develop, test, and maintain scalable ETL/ELT data pipelines .

- Develop production-quality data-processing solutions using Python and PySpark .

- Build reusable frameworks and components for ingestion, transformation, validation, and publishing of data.

- Process large structured, semi-structured, and distributed datasets.

- Implement incremental and batch-processing patterns where appropriate.

2. Snowflake Development

- Design and develop scalable data solutions using Snowflake .

- Develop complex SQL transformations, data models, views, and reusable data structures.

- Optimize Snowflake workloads for performance, scalability, and cost.

- Implement appropriate data-loading and transformation patterns between AWS data platforms and Snowflake.

- Troubleshoot performance and data-quality issues across Snowflake workloads.

3. dbt Development

- Build and maintain transformation pipelines using dbt .

- Develop modular, reusable, maintainable dbt models.

- Implement dbt tests and documentation.

- Follow appropriate development practices for source, staging, intermediate, and business-layer transformations.

- Support automated deployment and CI/CD practices for dbt projects.

4. Apache Iceberg / Lakehouse

- Design and implement data-lake and lakehouse solutions using Apache Iceberg .

- Build scalable table structures for large analytical datasets.

- Work with partitioning, schema evolution, incremental processing, and table-maintenance strategies.

- Integrate Iceberg-based datasets with Spark and AWS-based data-processing services.

- Ensure efficient storage and query patterns for high-volume datasets.

5. AWS Data Engineering

- Design and implement cloud-native data solutions on AWS .

- Build data-processing workloads leveraging services such as S3, Glue, EMR and Lambda where appropriate.

- Implement secure access patterns using AWS IAM.

- Monitor data workloads and troubleshoot operational issues.

- Participate in designing scalable, reliable, secure, and cost-efficient cloud data architectures.

6. Performance & Scalability

- Diagnose and optimize Spark/PySpark jobs , SQL queries, Snowflake workloads, and data pipelines.

- Identify bottlenecks involving compute, storage, partitioning, data skew, transformations, and queries.

- Design solutions capable of supporting increasing data volumes without unnecessary infrastructure cost.

7. Data Quality & Governance

- Implement automated data-quality checks across ingestion and transformation layers.

- Establish proper logging, monitoring, exception handling, and reconciliation mechanisms.

- Follow organizational standards for data security, governance, lineage, and access controls.

- Ensure production pipelines are reliable, auditable, and maintainable.

8. Engineering Best Practices

- Write clean, modular, reusable, testable, and maintainable code.

- Perform code reviews and enforce engineering standards.

- Implement unit, integration, and data-validation testing.

- Use Git-based version control and CI/CD practices.

- Create and maintain appropriate technical documentation.

9. Senior-Level Responsibilities

- Independently drive technically complex data-engineering requirements from design through production deployment.

- Participate in solution design and architecture discussions.

- Evaluate alternative implementation approaches and recommend appropriate solutions.

- Troubleshoot complex production and performance issues.

- Mentor junior and mid-level data engineers.

- Collaborate with Architects, Product Owners, Business Analysts, Data Scientists, QA, DevOps, and application teams.

- Translate business/data requirements into scalable technical solutions.

- Identify technical risks and proactively recommend improvements.

Core Skills Expected

A strong candidate should demonstrate deep hands-on capability , not merely theoretical exposure, in the following areas:

Area

Expected Capability

Python

Advanced, production-quality data engineering development

PySpark

Large-scale distributed processing, optimization and troubleshooting

Snowflake

Development, modeling, optimization and performance tuning

dbt

Models, tests, macros, documentation and deployment practices

Apache Iceberg

Lakehouse/table design, partitioning, schema evolution and optimization

AWS

Hands-on cloud data platform development

SQL

Advanced SQL, query optimization and analytical processing

Data Engineering

ETL/ELT, batch/incremental pipelines, data quality and orchestration

Data Architecture

Data Lake, Data Warehouse and Lakehouse concepts

Engineering Practices

Git, testing, code reviews, CI/CD and production support

Preferred Qualifications

- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.

- Strong experience delivering enterprise-scale cloud data platforms .

- Experience migrating legacy data workloads to modern AWS/Snowflake architectures.

- Experience working with very large datasets and distributed processing.

- Knowledge of data security and governance practices.

- Experience working in Agile delivery environments.

- AWS and/or Snowflake certification is an added advantage.