About this role
Valeo is a tech global company, designing breakthrough solutions to reinvent the mobility. We are an automotive supplier partner to automakers and new mobility actors worldwide. Our vision? Invent a greener and more secured mobility, thanks to solutions focusing on intuitive driving and reducing CO2 emissions. We are leader on our businesses, and recognized as one of the largest global innovative companies.
VALEO develops cutting-edge Artificial Intelligence solutions particularly in Computer Vision, and we are looking to strengthen our infrastructure team to support our ambitious projects." In addition ; AI is more and more in the focus on Simulation IIT Means.
- Main Mission: Reporting to the DSIS CoE DigIT Director you will be responsible for the design, implementation, management, and optimization of infrastructures (on-premise and cloud) dedicated to AI trainings and Simulation HPC Systems. You will play an essential role in the performance, scalability, and reliability of our HPC / AI and Cloud environments.
Key Responsibilities:
On-Premise Infrastructure Management:
Administer and maintain our GPU compute clusters (NVIDIA).
Manage storage and networks associated with AI/ML activities.
Ensure the availability and performance of on-premise resources.
Cloud Infrastructure Management (GCP):
Design, deploy, and manage cloud architectures on Google Cloud Platform (GCP) for AI/ML projects.
Master and optimize the use of relevant GCP services (Compute Engine, Kubernetes Engine (GKE), Cloud Storage, Networking, etc.).
Administer and operate the Vertex AI platform for model training, deployment, and management.
Simulation HPC Environment:
Master and optimize the use HPC Related infrastructure such as Rescale / CCRT / AtNorth. Slurm is the orchestrator for CCRT and AtNorth.
- Orchestration and Automation:
Develop, maintain, and improve orchestration tools (particularly in Python / Bash / Powershell / Slurm ) for managing and automating jobs across different infrastructures (on-premise and cloud).
Implement and promote Infrastructure as Code (IaC) practices (e.g., Terraform, Ansible) on the Cloud.
Support and Collaboration:
Work closely with Data Science and Machine Learning teams to understand their infrastructure needs and provide them with suitable environments.
Provide technical support on infrastructure aspects related to the entire AI stack (from development to production).
Diagnose and resolve incidents related to AI/ML and HPC infrastructures.
Optimization and Technology Monitoring:
Monitor the performance, costs, and security of the infrastructures.
Propose and implement optimizations.
Actively monitor technological advancements in hardware and software infrastructure solutions (on-premise and cloud), MLOps tools, and best practices in the field.
Profile:
Education: Engineering degree or Master's degree (Bac+5) in Computer Science or equivalent.
Experience: Confirmed experience (minimum 3-5 years suggested, to be adapted) in IT infrastructure management, including a significant part dedicated to AI/ML or High-Performance Computing (HPC) environments.
Essential Technical Skills:
Excellent expertize of Linux environments.
Solid experience in managing compute clusters, ideally with NVIDIA GPUs.
Solid experience in managing HPC Simulation Environment (Slurm)
In-depth knowledge of the Google Cloud Platform (GCP) ecosystem, including infrastructure services and the Vertex AI platform.
Advanced proficiency in Python development, particularly for automation, scripting, and orchestration.
Good knowledge of containerization technologies (Docker, Kubernetes).
Strong understanding of the end-to-end AI stack and MLOps principles.
Experience with Infrastructure as Code (IaC) tools.
Appreciated Skills (Assets):
Specific experience in infrastructure for Computer Vision.
Knowledge of other cloud platforms (AWS).
GCP Certification (Cloud Architect, Machine Learning Engineer...).
Knowledge of monitoring tools (Prometheus, Grafana...).
Personal Qualities:
Rigorous and methodical.
Autonomous and proactive.
Excellent analytical and problem-solving skills.
Technical curiosity and taste for innovation.
Good communication skills and team spirit.
Languages: Fluent French. Good level of technical English required (reading documentation, participating in technical exchanges).
Job:
IT Infrastructure Engineer
Organization:
R&D DSIS CoE
Schedule:
Full time
Employee Status:
Regular
Job Type:
Undefined term
Job Posting Date:
2026-05-21
Join Us !
Being part of our team, you will join:
- one of the largest global innovative companies, with more than 20,000 engineers working in Research & Development
- a multi-cultural environment that values diversity and international collaboration
- more than 100,000 colleagues in 31 countries... which make a lot of opportunity for career growth
- a business highly committed to limiting the environmental impact if its activities and ranked by Corporate Knights as the number one company in the automotive sector in terms of sustainable development
More information on Valeo: https://www.valeo.com