AI Data Engineer - Senior

Cummins Inc.Pune, MaharashtraOn-siteFull-timeSenior, 5–8 yearsListed 1 hour ago

Apply now

About this role

Job Summary:

Leads projects for design, development and maintenance of a data and analytics platform. Effectively and efficiently process, store and make data available to analysts and other consumers. Works with key business stakeholders, IT experts and subject-matter experts to plan, design and deliver optimal analytics and data science solutions. Works on one or many product teams at a time.

Key Responsibilities:

Designs and automates deployment of our distributed system for ingesting and transforming data from various types of sources (relational, event-based, unstructured). Designs and implements framework to continuously monitor and troubleshoot data quality and data integrity issues. Implements data governance processes and methods for managing metadata, access, retention to data for internal and external users. Designs and provide guidance on building reliable, efficient, scalable and quality data pipelines with monitoring and alert mechanisms that combine a variety of sources using ETL/ELT tools or scripting languages. Designs and implements physical data models to define the database structure. Optimizing database performance through efficient indexing and table relationships. Participates in optimizing, testing, and troubleshooting of data pipelines. Designs, develops and operates large scale data storage and processing solutions using different distributed and cloud based platforms for storing data (e.g. Data Lakes, Hadoop, Hbase, Cassandra, MongoDB, Accumulo, DynamoDB, others). Uses innovative and modern tools, techniques and architectures to partially or completely automate the most-common, repeatable and tedious data preparation and integration tasks in order to minimize manual and error-prone processes and improve productivity. Assists with renovating the data management infrastructure to drive automation in data integration and management. Ensures the timeliness and success of critical analytics initiatives by using agile development technologies such as DevOps, Scrum, Kanban Coaches and develops less experienced team members.

Experience:

5 to 8 years of experience in data engineering, with strong expertise in building and optimizing scalable data pipelines, ETL/ELT processes, and data integration solutions. Skilled in designing robust architectures that support advanced analytics, reporting, and data-driven applications.

Technical Skills:

Required:

- Knowledge of the latest technologies and trends in data science is highly preferred.
- Hands on experiences in the following are preferred:
- Exposure to Big Data open source
- Clustered compute cloud-based implementation experience
- Familiarity analyzing complex business systems, industry requirements, and/or data regulations
- Understanding of AI/ML concepts and tools
- Experience in ETL/ELT Data Engineering Technologies
- Background in processing and managing large data sets
- Design and development for a Big Data platform using open source and third-party tools
- Proficiency in Python, SQL, and Spark (PySpark preferred).
- Hands-on experience integrating with platforms like Palantir, Snowflake, Neo4j, etc.
- Solid knowledge of machine learning workflows, model deployment, and advanced analytics (regression, clustering, time-series analysis).
- Understanding of data governance, data cataloging tools (e.g., Azure Purview, Alation), and metadata management.
- SQL query language
- Clustered compute cloud-based implementation experience
- Experience developing applications requiring large file movement for a Cloud-based environment and other data extraction tools and methods from a variety of sources
- Take full ownership of the developed data pipelines, providing ongoing support for enhancements and performance optimization

Nice to have:

- Experience with graph data modeling and graph databases (e.g., Neo4j, TigerGraph) and familiarity with Palantir Ontology is a strong plus
- Understanding of data governance, data cataloging tools (e.g., Azure Purview, Alation), and metadata management.

Additional Key Responsibilities:

- Stay current with AI trends and suggest improvements to existing systems and workflows
- Excellent verbal and written communication skills
- Demonstrated self-starter with a proactive, problem-solving mindset

Candidate need to work from Cummins Pune IOC-B office for 3 days a week. There will be an overlap of few hours with US EST time zone.