About this role
Date Posted: 09/29/2026
Req ID: 50347
Faculty/Division: VP-People, Finance & Digital Services
Department: Enterprise Infrastructure Solutions
Campus: St. George (Downtown Toronto)
Position Number: 00058955
Existing Vacancy: Yes
Description:
About us:
The Enterprise Infrastructure Solutions (EIS) group, part of the Information Technology Services (ITS) division, is responsible for campus core network, campus wireless, wide area network connectivity and internet connectivity for the University, including connectivity to research and education networks.
EIS is also responsible for services related to departmental network management, network, server and storage infrastructure, Windows and Linux servermanagement services, database and application integration and support, enterprise backup service, 24/7 operation of central administrative data centres and telecommunications services.
If you’re motivated and passionate about learning technologies and dedicated to improving experiences for today’s student, consider a career with us.
Your opportunity:
Reporting to the Manager, AI Engineering & Operations within the Enterprise Infrastructure Solutions group, the AI Infrastructure & Solutions Architect plays a critical role in defining the future of research and administrative computing at the University. In this role, you will lead the architectural design and deployment of secure, scalableAI platforms that serve the entire campus community, spanning the University’s own data centres and private cloud as well as the major public cloud platforms (AWS, Microsoft Azure, and Google Cloud). You will bridge the gap between high-performancehardware, cloud services, and practical user applications, ensuring that our AI platforms, from on-premises GPU clusters and sovereign data sandboxes to cloud-hosted AI services, are reliable, supportable, cost-effective, and aligned with institutional governance and ethical AI frameworks.
You will collaborate closely with AI Developers & Integration Specialists, while retaining ownership of platform architecture, workload placement, reliability, security posture, cost efficiency, and lifecycle management across the University’s on-premises infrastructure and multiple public cloud environments.
This role offers a rare opportunity to design and operate AI platforms at institutional scale across a multi-cloud estate that combines sovereign, on-premises capabilities with AWS, Microsoft Azure, and Google Cloud, supporting both cutting-edge research and mission-critical administrative use cases.
Your responsibilities will include:
- Architecting and operating container orchestration platforms (Kubernetes/K8s) on-premises and in the cloud (e.g., EKS, AKS, GKE), including GPU operators and AI-aware schedulers for efficient accelerator utilization.
- Defining and maintaining the University’s multi-cloud AI reference architecture and workload placement framework, determining where AI workloads run across the University’s on-premises infrastructure, AWS, Microsoft Azure, and Google Cloud based on data classification, sovereignty, cost, performance, and service availability.
- Designing and integrating managed cloud AI services (e.g., Amazon Bedrock and SageMaker, Azure AI Foundry and Azure OpenAI, Google Vertex AI) with on-premises platforms through common abstractions such as model gateways, federated identity, and cloud landing zones provisioned through Infrastructure as Code.
- Designing and enforcing AI platform security and governance controls across on-premises and cloud environments, including data sovereignty and residency (e.g., Canadian cloud regions), identity federation and access isolation, auditability, and compliance with privacy and ethical AI frameworks.
- Implementing observability and monitoring solutions to track model performance, drift, GPU utilization, inference latency, cost, and platform health using tools such as Prometheus, Grafana, OpenTelemetry, cloud-native monitoring services (e.g., Amazon CloudWatch, Azure Monitor, Google Cloud Monitoring), or specialized AI monitoring stacks, providing a unified view across on-premises and cloud platforms.
- Designing and operating GPU and accelerator platforms on-premises and in the cloud (e.g., NVIDIA, AMD, cloud GPU and TPU instances, or emerging accelerators), including capacity planning, scheduling strategies, burst-to-cloud patterns, and lifecycle management.
- Analyzing platform usage and cost metrics across on-premises and cloud platforms to optimize token consumption, GPU allocation, cloud spend (FinOps), and overall cost efficiency while maintaining performance and reliability.
- Partnering with AI Developers & Integration Specialists to define platform abstractions, deployment patterns, and service interfaces that enable rapid innovation without compromising security, supportability, or portability across on-premises and cloud environments.
- Producing and maintaining architectural documentation, disaster recovery and business continuity plans (including cross-cloud and cloud-to-on-premises recovery strategies), and technical guidance for researchers and platform users.
Essential Qualifications:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or an acceptable combination of education and equivalent experience.
- Eight or more years of experience in on-premises and public cloud infrastructure management.
- Three to five+ years of direct AI infrastructure or MLOps experience, recognizing the rapid evolution of the field, with demonstrated exposure to LLM platforms, RAG pipelines, or large-scale ML systems.
- Deep expertise in Infrastructure as Code (IaC) and automation (e.g., Terraform/OpenTofu, Ansible, and cloud-native tooling such as CloudFormation or Bicep) for managing complex, multi-environment and multi-cloud platforms.
- Advanced Kubernetes and container orchestration knowledge across self-managed and cloud-managed distributions (e.g., EKS, AKS, GKE), including GPU scheduling, operators, and container runtimes (Docker, Podman).
- MLOps and model lifecycle tooling experience, such as MLflow, Kubeflow, Weights & Biases, and model serving frameworks like Triton or vLLM, as well as their cloud-managed equivalents (e.g., Amazon SageMaker, Azure Machine Learning, Vertex AI).
- Strong understanding of high-performance and AI-optimized networking, including high-throughput, low-latency designs (e.g., InfiniBand, RDMA) and hybrid and multi-cloud connectivity (e.g., AWS Direct Connect, Azure ExpressRoute,Google Cloud Interconnect, private endpoints, and cross-cloud networking).
- Demonstrated experience designing and operating hybrid and multi-cloud architectures, with hands-on expertise in at least two of the three major public cloud platforms (AWS, Microsoft Azure, Google Cloud) and their AI/ML and GPU compute services, including identity federation, networking, security, and cost management (FinOps) considerations.
Assets (Nonessential):
- Experience with vector databases and retrieval systems supporting RAG-based architectures.
- Hands-on experience deploying open-weights AI models (e.g., Llama, Mistral) in on-premises, air-gapped, or tightly governed environments, as well as on cloud GPU infrastructure or through cloud model catalogues.
- Knowledge of AI security practices, including adversarial ML considerations, secure AI framework implementations, or cloud security posture management (CSPM) across multiple providers.
- Familiarity with Canadian data sovereignty, privacy, and research compliance requirements, particularly within higher education, including the data residency options offered by Canadian cloud regions.
- Contributions to open-source AI, infrastructure, or MLOps projects, or active participation in the AI engineering community.
- Experience with ITSM platforms (e.g., ServiceNow) and operational workflow automation.
- Professional-level cloud architecture certification (e.g., AWS Certified Solutions Architect – Professional, Microsoft Certified: Azure Solutions Architect Expert, Google Professional Cloud Architect) or equivalent demonstrated experience.
- Experience with FinOps practices and multi-cloud cost management tooling, including cost allocation and showback/chargeback models for shared research and administrative platforms.
- Experience with enterprise private cloud and virtualization platforms (e.g., VMware vSphere) and integrating them with public cloud services in a hybrid model.
To be successful in this role you will be:
- Decisive
- Efficient
- Organized
- Proactive
- Self-driven
- Innovative
Closing Date: 10/20/2026, 11:59PM ET
Employee Group: USW
Appointment Type : Budget - Continuing
Schedule: Full-Time
Pay Scale Group & Hiring Zone:
USW Pay Band 19 -- $128,706 with an annual step progression to a maximum of $164,586. Pay scale and job class assignment is subject to determination pursuant to the Job Evaluation/Pay Equity Maintenance Protocol.
Job Category: Information Technology (IT)
Divisional HR Office Email:
Lived Experience Statement
Candidates who are members of Indigenous, Black, racialized and 2SLGBTQ+ communities, persons with disabilities, and other equity deserving groups are encouraged to apply, and their lived experience shall be taken into consideration as applicable to the posted position.
Job descriptions are available upon request for internal applicants.