About this role
Data Engineer, AI Training and Evaluation
About the Role
We are looking for a Data Engineer to build the trusted data foundations behind Blue Yonder's AI agents and model-improvement workflows.
As part of Blue Yonder's Autonomy Labs, you will work with software engineers, model-training teams, researchers, and product owners. You will build pipelines and data products that make agent behavior understandable: what happened, which tools were used, whether the workflow succeeded, where failures occurred, and what evidence should inform the next model or product iteration.
This is not primarily a reporting or business-intelligence role. The data you build and maintain will support evaluation, debugging, experimentation, model training, regression testing, and continuous improvement in production.
Your Mission
Create a reliable and reproducible data foundation connecting live agent behavior with evaluation and model development. You will ensure that traces, outcomes, feedback, and training data are trustworthy, well governed, and usable by both engineering and model-training teams.
The systems you build will help turn noisy production activity into evidence: curated datasets, measurable failure modes, reproducible scenarios, and high-quality signals for improving models and agents.
What You'll Do
- Design and operate scalable pipelines for collecting agent traces, model responses, tool calls, workflow state, evaluation results, user feedback, and business outcomes.
- Define durable data models and schemas that make behavior comparable across models, prompts, tools, datasets, environments, and application versions.
- Build reliable batch and streaming workflows for validating, normalizing, enriching, joining, deduplicating, and transforming complex AI system data.
- Create curated and versioned datasets for model training, evaluation, regression testing, scenario replay, failure analysis, and experimentation.
- Implement automated data-quality checks covering completeness, consistency, provenance, schema conformance, distribution changes, and unexpected behavior.
- Build data interfaces supporting annotation, human review, synthetic data generation, evaluation, and model-development workflows.
- Establish lineage, access controls, retention policies, and privacy protections for potentially sensitive production and training data.
- Help maintain clear separation between training, validation, and evaluation datasets while preserving reproducibility and preventing leakage.
- Partner with engineers building agent environments and evaluation systems to provide reproducible scenarios, state snapshots, fixtures, and expected outcomes.
- Make data easy to discover and use through well-designed data products, documentation, metadata, and appropriate self-service capabilities.
- Improve observability across the model-improvement lifecycle, including dataset health, pipeline reliability, evaluation coverage, and feedback-loop performance.
- Work directly with model-training teams to understand data requirements, investigate behavior patterns, and convert experimental workflows into reliable production pipelines.
What We're Looking For
- Strong experience building and operating production-grade data platforms, pipelines, or data-intensive applications.
- Strong Python and SQL skills, with expertise in data modelling, orchestration, distributed processing, data contracts, testing, and observability.
- Experience working with complex event, telemetry, interaction, or workflow data—not only aggregated analytical tables.
- Experience with cloud data platforms and modern batch and streaming architectures.
- Strong understanding of data quality, provenance, lineage, governance, privacy, access control, and reproducibility.
- Experience designing data products that serve multiple technical consumers with different access patterns and quality requirements.
- Ability to collaborate closely with software engineers, researchers, and model-training teams and translate evolving experimental needs into maintainable systems.
- Strong engineering judgment and a commitment to operational reliability, clear interfaces, and well-tested data transformations.
Preferred Qualifications
- Experience working with ML training data, evaluation datasets, annotation systems, synthetic data, embeddings, experiment tracking, or feature platforms.
- Familiarity with LLM applications, AI agents, tool-use traces, model evaluation, or post-training workflows.
- Experience designing datasets and pipelines for replay, simulation, offline evaluation, or behavioral regression testing.
- Knowledge of techniques for detecting data leakage, distribution drift, label-quality issues, and dataset contamination.
- Experience with supply chain, retail, or other enterprise domains involving sensitive data and complex operational workflows.
What Makes This Role Different
The data platform is part of the learning system. Your work will not stop at delivering tables or dashboards: it will determine whether teams can understand agent behavior, reproduce results, trust evaluations, construct useful datasets, and improve models after they are deployed.
Our Values
If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.