About this role
Responsibilities:
- Design, develop, and maintain data pipelines that ingest process, and deliver data from various sources, ensuring data quality and reliability.
- Data Modeling: Create and maintain data models to support reporting, analytics, and business intelligence needs, optimizing data structures for performance and efficiency.
- Implement ETL processes to transform raw data into meaningful insights, handling data transformation, aggregation, and enrichment.
- Monitor and address data quality issues, implement data validation processes, and establish data governance practices.
- Manage and optimize data storage, processing, and distribution systems, ensuring scalability and performance.
- Collaborate with data scientists, analysts, and cross-functional teams to understand data requirements and deliver solutions that meet business needs.
- Document data engineering processes, pipelines, and systems to maintain clear and accessible knowledge for team members.
Requirements:
- Bachelor’s or Master’s degree in Computer Science, Data Science, or a related field.
- Minimum of 3 years of experience in data engineering or a related field.
- Proven experience with Redshift and Snowflake.
- Experience with DBT.
- Experience with Apache Airflow for workflow orchestration.
- Proficiency in Python for data pipeline development and scripting.
- Experience with AWS cloud services, including S3, EC2, and EMR.
- Familiarity with Kafka for real-time data streaming.
- Familiarity with data visualization and reporting tools (e.g. Tableau, Power BI or Looker).
- Project management skills using Jira or similar tools.
- Ability to collaborate effectively with cross-functional teams and understand business data needs.
- Strong problem-solving skills, a proactive approach to troubleshooting data issues, and critical thinking abilities.
- Adaptable and open to learning new technologies and methodologies.