About this role
Data Engineer - Commodities
About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millennium’s mission is to deliver results for our investors.
Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.
Meet the Team
The Commodities Technology team builds and operates a data platform that aggregates and curates critical commodities data, including weather, supply and demand, storage, transportation, and other fundamental and alternative datasets. The team’s curated content layer supports Portfolio Managers and researchers in understanding markets and constructing trades. As part of Millennium’s Information Technology organization, the team develops flexible, scalable technology and proprietary systems that support advanced analytical and trading capabilities.
What You'll Do
• Design, implement, and maintain end-to-end ETL workflows in Python and SQL to ingest and transform commodities data from multiple vendor and internal sources.
• Build and maintain standardized data models, schemas, and metadata that make commodities datasets easy to understand, discover, and reuse.
• Schedule, monitor, and manage data pipelines using Airflow or similar workflow orchestration tools to ensure reliable, timely data delivery.
• Implement validation, reconciliation, and anomaly-detection controls to maintain data completeness, accuracy, and consistency.
• Apply AI to automate schema inference across structured and semi-structured data sources, manage schema drift, and accelerate scalable ingestion development.
• Use AI-driven data-quality, observability, and documentation capabilities to detect anomalies, monitor data health, and produce clear lineage and technical documentation.
• Maintain high-quality, production-ready code and repeatable deployments through Git, GitHub Actions, PyTest, and automated testing practices.
• Partner with commodities Portfolio Managers, researchers, data analysts, and data strategists to refine datasets, definitions, and documentation based on evolving business needs.
What You Bring
• 5+ years of experience in data engineering, analytics engineering, or a similar role building and maintaining ETL pipelines.
• Strong Python and SQL skills, including experience working with large datasets and complex data transformations.
• Hands-on experience with Airflow or comparable workflow orchestration tools.
• Experience with Git, CI/CD pipelines such as GitHub Actions, and automated testing frameworks such as PyTest.
• Strong attention to detail, data quality, and documentation, with the ability to identify edge cases and protect data integrity.
• Ability to work independently, communicate effectively with technical and non-technical stakeholders, and manage multiple concurrent initiatives.
• Knowledge of commodities markets and data, including weather, supply and demand, storage, freight, or flows, is preferred.
• Experience with data warehousing, columnar storage formats, analytic databases, data catalog or governance tools, or financial services, trading, or research-driven environments is preferred.