Staff AI Engineer

Blue YonderParis, Île-de-FranceOn-siteFull-timeStaff, 8–12 yearsListed 52 minutes ago

Apply now

About this role

About the Role

We are looking for an experienced Staff AI Engineer to help build the engineering foundations behind Blue Yonder's next generation of AI-powered products.

As part of Blue Yonder's Autonomy Labs, you will work alongside model-training teams, researchers, data engineers, and product engineers to develop agents that operate within real supply chain and retail workflows. You will build the services, APIs, evaluation systems, and development environments that allow us to train, test, deploy, and continuously improve these systems.

This is a hands-on software engineering and technical leadership role close to the model-development lifecycle. You do not need to be a model-training researcher, but you should understand how production behavior, evaluation results, traces, and data can be turned into better models and more reliable products.

Your Mission

Help turn increasingly capable models into dependable software systems. You will own important architectural decisions, build critical parts of the platform, and create the engineering loops through which agents are evaluated and improved after they begin operating in real workflows.

The challenge is not simply to make a model sound knowledgeable about supply chain. It is to build agents that can use tools, manage state, follow constraints, recover from failures, and complete valuable work reliably.

What You'll Do

- Architect and build robust software systems supporting the development, evaluation, deployment, and ongoing improvement of AI agents.

- Develop services and pipelines for collecting and processing agent traces, tool calls, workflow outcomes, user feedback, and other production signals.

- Build systems that evaluate and verify model behavior at scale using deterministic checks, model-based evaluation, statistical analysis, and human review where appropriate.

- Create reliable feedback loops between production systems and model-training teams, turning observed failures and successful behaviors into evaluation cases, regression tests, and candidate training data.

- Build training and evaluation environments that represent realistic business workflows through stable APIs, tools, simulations, resettable scenarios, and reproducible state.

- Implement the integrations and supporting services required for agents to interact with enterprise systems, domain data, and multi-step workflows.

- Design for reproducibility across models, prompts, tools, datasets, environments, and application versions so that changes can be measured with confidence.

- Partner with model-training teams to integrate new models, investigate behavior failures, define engineering requirements, and determine whether problems are best addressed through software, tools, data, evaluation, or model changes.

- Improve reliability and operability through observability, automated testing, failure recovery, safe rollout mechanisms, and clear launch criteria.

- Set a high engineering bar through system design, code review, documentation, mentoring, and pragmatic technical leadership across teams.

What We're Looking For

- Around eight or more years of software engineering experience, or equivalent depth, with a track record of designing and delivering complex production systems.

- Strong Python and backend engineering expertise, including API design, distributed systems, data processing, automated testing, and maintainable service architecture.

- Hands-on experience building LLM-powered applications, AI agents, machine learning platforms, or similarly complex systems.

- Practical understanding of AI evaluation, including test-case design, trace analysis, failure classification, regression testing, and measurement of nondeterministic systems.

- Strong understanding of software engineering fundamentals, including code quality, security, observability, CI/CD, and operational reliability.

- Experience with cloud platforms such as Azure, AWS, or GCP, along with containerization and modern deployment practices.

- The ability to work effectively with researchers and model-training engineers, translating experimental needs into well-designed software and production evidence into actionable feedback.

- Strong technical judgment, clear communication, and the ability to provide direction in ambiguous problem spaces without becoming a bottleneck for delivery.

Preferred Qualifications

- Familiarity with model-training or post-training approaches such as supervised fine-tuning, preference optimization, reinforcement learning, reward or verifier design, and checkpoint evaluation.

- Experience building model-training environments, simulators, evaluation harnesses, or agent development platforms.

- Understanding of training and evaluation data practices, including curation, synthetic data, provenance, versioning, quality control, and prevention of data leakage.

- Experience with supply chain, retail, or other enterprise environments involving complex workflows, permissions, business constraints, and human approval paths.

What Makes This Role Different

This role sits between model development and production engineering. You will not be asked to conduct research in isolation or simply wrap a model in an application. You will build the systems that make model behavior measurable, reproducible, improvable, and useful in real operational environments.

Our Values

If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.