About this role
Remote Position: Hybrid
Region: Americas
Country: Canada
State/Province: Ontario
City: Toronto
Project Objectives & Role Summary
The AI DevOps Specialist is primarily responsible for the technical deployment, secure enablement, administration, and continuous optimization of Celestica’s global HPS Artificial Intelligence (AI) sandbox, modeling, and software engineering toolchains.
This specialist will bridge the gap between AI development pipelines, secure networking infrastructure, and cloud platform services. They will take a hands-on lead in configuring foundational AI cloud services (primarily Google Cloud Platform - Vertex AI), managing user-facing web interfaces (LibreChat or Gemini Enterprise Agent Platform), implementing agentic workflow architectures (Flowise AI or Gemini Enterprise Agent Platform), ensuring secure, uninterrupted developer access to AI coding assistants (such as Claude Code, Gemini Code Assist, and Claude Sonnet/Opus models), and establishing robust cloud budget governance.
Core Responsibilities & Scope of Work
1. AI Sandbox Platform Administration & Deployment
- Multi-Phase Sandbox Rollout: Own the deployment and lifecycle management of the HPS AI Sandbox environment across cloud and hybrid infrastructures.
- Front-End User Interfaces: Deploy, configure, and maintain LibreChat (or similar web-based interfaces) to provide HPS engineers with safe, compliant, and localized access to LLMs.
- Agentic Frameworks: Establish and configure Flowise AI or Gemini Enterprise Agent Platform for the design and orchestration of agentic AI workflows and LLM-backed applications.
- Model Registry & Endpoints: Administer model deployment, model endpoint configurations, and vector databases within the sandbox environment.
2. GCP & Vertex AI Model Management
- Project Governance: Oversee the HPS-designated GCP AI projects (e.g., gcp-ai-hps), including security permissions, service accounts, and IAM roles.
- Model Enablement & Quota Tuning: Coordinate the provisioning and scale-out of advanced foundational models (including Anthropic Claude 3.5/3.6/3.7 suite [Opus, Sonnet, Haiku] and Google Gemini) in Vertex AI. Actively manage, troubleshoot, and resolve API quota restrictions with cloud providers.
- Local Code Integration: Ensure developer command-line interfaces and local development terminals (such as Claude Code, VS Code, and GitKraken) seamlessly communicate with cloud model endpoints.
3. Secure DevOps & Network Integration (Zscaler & Firewall)
- Lab Network Troubleshooting: Diagnose and resolve intermittent connection issues between HPS Design Labs (such as the Innovation Lab) and external AI resources or code repositories (e.g., troubleshooting DNS, routing timeouts, and SSL inspection blocks on GitHub).
- Traffic Rules & Proxy Controls: Partner with Network Security to define, test, and troubleshoot Zscaler ZTNA app connectors, firewall rules, and proxy exceptions necessary to enable secure outbound AI API traffic while protecting proprietary codebase egress.
- CI/CD Integration: Work alongside DevOps administrators to embed automated vulnerability checks, binary scanning, and AI-assisted testing steps in Azure DevOps, Jenkins, and GitHub pipelines.
4. Cloud Budget Governance & Financial Controls
- Cost Allocation & Monitoring: Design and implement rigid budget monitoring, consumption alerts, and cost-attribution controls in GCP to track HPS developer usage.
- Usage Auditing: Develop weekly/monthly utilization dashboards to track API token consumption, model call costs, and sandbox compute runtimes.
- ROI Optimization: Provide recommendations on token limits, model pruning, caching strategies, and model choice (e.g., optimizing workloads to use Haiku or Flash models where appropriate to conserve budget).
5. AI Security, IP Protection & Compliance
- Data Sovereignty Compliance: Enforce enterprise policies ensuring that no proprietary hardware schematics, PCB layouts, firmware source code, or IP are ingested into public training models.
- Evaluation & Testing: Support the isolation of the HPS Innovation Lab to evaluate new open-source models, libraries, and AI security evaluation tools prior to general HPS rollout.
Education & Experience
- Bachelor’s degree in Computer Science, Software Engineering, DevOps, Cloud Engineering, or equivalent technical experience.
- 4+ years of hands-on experience in DevOps, Cloud Engineering, or System Administration, with at least 2 years focused specifically on AI/MLOps platform delivery.
Knowledge/Skills/Competencies
Required Technical Skills
- Cloud Platform Expertise (GCP): Advanced experience with Google Cloud Platform (GCP) and specifically Vertex AI / Model Garden, IAM, billing alerts, and VPC setups.
- AI Toolchain & LLM Tooling: Proven experience deploying and maintaining containerized LibreChat architectures, Flowise AI (or LangChain/LlamaIndex equivalents), and local/cloud LLM APIs.
- Networking & Security Engineering: Solid understanding of enterprise networking protocols (DNS, TCP/IP routing, NAT, SSL/TLS handshake) and secure access controls (Zscaler ZTNA, enterprise Firewalls).
- Containerization & Orchestration: Strong proficiency with Docker, docker-compose, and Kubernetes to deploy scalable sandbox services.
- Code Assist Integration: Familiarity with the configuration of developer-focused AI integrations like Claude Code, Gemini Code Assist, or MSFT Copilot CLI inside Linux/Mac environments.
- Automation & Scripting: Strong scripting abilities in Python (specifically utilizing AI/ML libraries, request handling, and GCP SDKs) and Bash.
Preferred Certifications
- Google Cloud Professional DevOps Engineer or Professional Cloud Architect
- Google Cloud Professional Machine Learning Engineer
- Certified Kubernetes Administrator (CKA)
- HashiCorp Certified: Terraform Associate
Working Style & Competencies
- Diagnostic Mindset: Exceptionally strong debugging skills for networking, package distribution, and cloud service interconnections.
- Proactive Collaboration: Able to work cross-functionally with HPS Hardware Design Teams, Enterprise IT, Corporate Security, and external consulting suppliers (e.g., Elastify).
- Detail-Oriented Documentation: High commitment to writing complete, clear Standard Operating Procedures (SOPs), system topologies, and budget management guidelines.
Physical Demands
Salary
The stated range includes Base Salary and target Short-Term Incentive (STI) compensation only. A comprehensive benefits package is offered in addition to this range.
The range described in this posting is an estimate by the Company, and may change based on several factors, including but not limited to a change in the duties covered by the job posting, or the credentials, experience or geographic jurisdiction of the successful candidate.
109,000 - CAD 173,000
Notes
This job description is not intended to be an exhaustive list of all duties and responsibilities of the position. Employees are held accountable for all duties of the job. Job duties and the % of time identified for any function are subject to change at any time.
Celestica is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, pregnancy, genetic information, disability, status as a protected veteran, or any other protected category under applicable federal, state, and local laws.
At Celestica we are committed to fostering an inclusive, accessible environment, where all employees and customers feel valued, respected and supported. Special arrangements can be made for candidates who need it throughout the hiring process. Please indicate your needs and we will work with you to meet them.