About this role
Job Description & Summary
The opportunity
Design and evolve the reusable platform capabilities required to build, deploy, secure, observe and scale AI products across hybrid environments.
What you will be doing
· Define reference platforms, environment topology, network patterns, identity, secrets and access controls.
· Design model access, compute, storage, vector services, orchestration, gateways and shared platform services.
· Support cloud, sovereign, edge and on-premises deployment choices where client constraints require them.
· Define platform non-functional requirements for availability, performance, resilience, scalability and cost.
· Create reusable landing patterns, templates and platform guardrails for delivery squads.
· Partner with the MLOps Engineer on environment automation, observability and release management.
What we need from you
· 8+ years in cloud platforms, infrastructure, platform engineering or architecture.
· Deep understanding of containers, Kubernetes, networking, identity, automation, data services and AI runtimes.
· Experience with at least one major cloud platform and hybrid integration patterns.
· Strong Infrastructure as Code, DevSecOps and platform-governance orientation.
Relevant AI technologies and tooling
· Deep experience with an enterprise AI platform such as Azure AI Foundry and Azure OpenAI, AWS Bedrock, Google Vertex AI, Databricks, or an equivalent platform, plus the ability to integrate alternative model providers.
· Hands-on knowledge of containers and Kubernetes platforms such as AKS or OpenShift; Infrastructure as Code using Terraform, Bicep or equivalent; and GitHub Actions or Azure DevOps pipelines.
· Experience exposing and securing model, agent and tool services through API gateways, private endpoints, service identities, secrets platforms and policy enforcement, for example Microsoft Entra ID, API Management and Key Vault.
· Ability to design hybrid inference patterns using managed endpoints and self-hosted runtimes such as vLLM, Hugging Face tooling, NVIDIA inference components, Ollama or equivalent technologies where appropriate.
· Knowledge of vector, search and state services such as Azure AI Search, PostgreSQL with pgvector, Elasticsearch, Pinecone, Weaviate, Milvus, Redis, Cosmos DB or equivalent.
· Experience integrating platform telemetry with OpenTelemetry and enterprise monitoring stacks, and designing for model or framework portability rather than unnecessary vendor lock-in.
Measures of success
· Platform readiness and reliability
· Time required to onboard new use cases
· Reuse of platform components
· Cost, performance and capacity transparency
· Compliance with approved architecture guardrails
Key interfaces
· Other members of the AI Transformation & Agentic Systems Practice
· PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists
· Client business owners, product owners, technology teams and operational users
· Technology alliance and implementation partners where relevant
Contribution to the practice
· Support proposals, client workshops and market development appropriate to seniority.
· Contribute reusable methods, patterns, code, assets and lessons learned.
· Coach colleagues and participate in the capability’s continuous learning agenda.
· Uphold PwC quality, independence, confidentiality and risk-management requirements.
#LI-BS1 #LI-Hybrid