Sr. Data Engineer

ARC-One SolutionsUnited StatesRemoteFull-timeSenior, 5–8 yearsListed 1 week ago

Apply now

About this role

Overview

Manages and evolves the enterprise data lake and data warehouse while ensuring the reliable, secure, and efficient flow of high-quality data. Implements data processes, managing data architecture, designing ETL processes, and analyzing data for business insights.

The base salary range for this position is $99,937-$157,044.

Actual pay will be determined based upon a candidate’s job-related knowledge, skills, education, experience, geographic location, and may include other job-related factors such as certification(s), professional licensure, or internal equity considerations.

Responsibilities

- Design, implement and maintain scalable data pipnes on WS using S3, DMS, Glue, lambda, step function/MWAA & Redshift.

- Develop robust batch and near-real-time ETL/ELT workflow to ingest, cleanse, transform and load data from databases, legacy applications and event streams using Python & Pyspark.

- Design incremental/CDC mechanism, including restart ability, idempotency, duplicate handling and recovery.

- Implement automated controls for completeness, accuracy, reconciliation, schema changes & lineage.

- Optimize Glue/Spark, Athena, Redshift & S3 workload through partitioning, columnar formats, query tuning and appropriate storage/compute design.

- Design near real time/event-driven pipelines using Kinesis/Kafka where required, covering ordering, retry, idempotency and failure recovery.

- Implement AWS data security, least privilege access, data classification and governance controls.

- Monitor pipelines such as CloudWatch, troubleshoot failure and resolving production data incidents.

- Enforce Git/version control, code review, automated testing and CI/CD practices.

- Work with product owners, architect, reporting and business stakeholders to translate requirements into scalable data solutions.

- Document pipelines and operational procedures.

Qualifications

Qualifications Required

- Bachelor's degree in a computer-related field from an accredited college or university and five (5) or more years of experience in data engineering, building scalable and distributed ETL data pipelines in enterprise environments.

- Experience building and operating scalable AWS-based data platforms and pipelines using services including Lambda, Glue, Athena, S3, Redshift, DMS, MWAA (Airflow), and Step Functions, supporting batch, CDC, and near real-time data processing.

- Advanced proficiency in Python, SQL, and PySpark with hands-on experience developing reusable ETL/ELT frameworks, data warehouses, data marts, and integrations across databases, APIs, event streams, and analytics environments.

- Experience implementing data quality, governance, and optimization best practices, including automated validation frameworks, Lake Formation and Glue Data Catalog, performance tuning, and cost optimization across AWS data services.

- Strong communication skills with the ability to translate complex data concepts for business stakeholders; experience in healthcare, life sciences, and other highly regulated environments with HIPAA, GDPR, FDA, or similar compliance requirements preferred.

- Experience with metadata management, data lineage, data observability, master data management, or enterprise data catalog solutions.

- Knowledge with data modeling & analytical data model, schema design, schema evolution, and data structure optimized for reporting and analytics.

- Knowledge of data lake and data warehouse architecture include data partitioning and columnar storage format such as Parquet.

- Relevant AWS certification, such as AWS Certified Data Engineer – Associate, or an equivalent cloud or data engineering certification.

WORKING CONDITIONS

- Flexible work hours in fun collaborative environment

- Working remote requires a reliable internet connection

- Must have the ability to travel, as needed for company meetings