About this role
ABOUT THE ROLE
We are looking for an experienced Data Enginee r for our partners. an outsourcing company that works with a client in the energy industry. You will design, build and maintain the data pipelines and cloud-based architectures that support business intelligence, analytics and machine learning.
You will work with Python, SQL, Apache Spark, Kafka and Airflow, turning data from a variety of sources into reliable, well-structured data lakes and warehouses.
COLLABORATION
- B2B collaboration;
- Hybrid remote in Bucharest.
DUTIES AND RESPONSIBILITIES
- Design, develop, and maintain robust data pipelines and architectures to support business intelligence, analytics, and machine learning models.
- Collaborate with data scientists, analysts, and other engineers to ensure seamless integration of data across platforms.
- Collaborate with business stakeholders to understand their data needs, translating business requirements into scalable and efficient data solutions.
- Communicate complex technical concepts clearly to non-technical stakeholders, ensuring alignment between engineering and business teams.
- Implement efficient data processing solutions, ensuring high performance and scalability.
- Optimize data storage, retrieval, and ETL processes to ensure data availability and integrity.
- Lead the design and development of cloud-based data solutions.
- Maintain version control and collaborate using GitLab to manage code and deployment pipelines.
- Develop and optimize SQL queries and integrate data from a variety of sources into central repositories (e.g., data lakes, data warehouses).
- Automate and streamline data pipeline processes with shell scripting and other relevant tools.
- Provide mentorship to junior data engineers, guiding them through technical challenges and helping then grow in their roles.
- Lead code reviews and ensure best practices are followed across the engineering team.
REQUIREMENTS
Technical requirements
- Bachelor degree in STEM subjects;
- 4+ years in Data Engineer role or similar role with proven track record of providing large-scale data solutions;
- Strong experience in SQL, including complex queries;
- Strong programming experience with Python for data engineering tasks, including building custom data processing solutions and automation;
- In-depth experience with Apache Spark and Apache Kafka for distributed data processing and batch processing;
- Hands-on experience with Apache Airflow for orchestration of data flows;
- Solid experience in AWS (or another Cloud provider);
- Proficiency in using GitLab for version control and CI/CD processes;
- Hands-on experience Docker and Kubernetes for containerization;
- Experience with DBT for data transformation;
- Experience in Shell scripting (e.g. Bash).
Nice to have
- Experience with at least one other programming language Java, Scala, Apache Hive;
- Understand data governance, security best practices, and compliance regulations;
- Familiarity with agile methodologies and working in cross-functional teams;
- Familiarity with Machine Learning concepts.
Key Attributes
- Strong problem-solving skills and ability to work with large, complex datasets;
- Excellent communication and collaboration skills to work effectively across teams;
- Ability to mentor junior engineers and provide technical leadership.