Director of Platform Engineering

Gen Digital Inc.Prague, PragueOn-siteFull-timePrincipal, 12–15+ yearsListed 12 hours ago

Apply now

About this role

About Us:

Gen is a global company dedicated to powering Digital Freedom through its trusted consumer brands including Norton, Avast, LifeLock, MoneyLion and more. Our combined heritage is rooted in financial empowerment and cyber safety for the first digital generations, and today we deliver award-winning cybersecurity, online privacy, identity protection and financial wellness solutions to nearly 500 million users in more than 150 countries.

Together, we share a collective passion and vision to protect consumers and help them grow, manage and secure their digital and financial lives. We’re always looking for smart, fearless and high-impact talent who see AI as a teammate – leveraging it to move faster and deliver meaningful results.

When you’re part of Gen, you’ll have the flexibility, tools and support to do your best work and grow your career – from flexible working options and time off to competitive pay, benefits and well-being programs.

At Gen, we are scrappy and relentlessly customer driven. We create room for healthy debate, experimentation and continuous learning, and we seek out people with different experiences, identities and ideas to join our team. You’ll work with people who back each other, respect each other and understand that our differences are a competitive advantage.

If this sounds like you, we’d love you to be part of Gen.

About The Role:

The Director, Platform Engineering is a senior technical leadership role within Gen Digital's Production Operations organization, accountable for the platforms, infrastructure, and tooling that the entire development and technology organization builds on. You will lead three platform teams — Cloud Platform, Productivity Platform, and AI Platform — and the engineering managers who run them, setting a unified technical strategy that spans multi-cloud infrastructure, the internal developer platform, enterprise engineering tooling, and the emerging AI platform. You will treat these platforms as products: defining golden paths and paved roads, measuring adoption and developer experience, and reducing the cognitive load required for product teams to ship safely and quickly.

This is a highly visible role that partners directly with engineering, security, data, and business leadership. You are accountable for the availability, scalability, performance, cost efficiency, and compliance posture of the shared platform estate across AWS, Google Cloud, and Azure, including 24x7 production ownership, incident and problem management, and the operational readiness of AI and machine learning workloads running in production. You will build and develop a high-performing, multi-disciplinary organization and represent platform strategy to executive stakeholders.

In this role, you will:

Platform Strategy and Execution:

- Own the multi-year technical strategy and roadmap for the Cloud Platform, Productivity Platform, and AI Platform teams, aligned to the goals of the Production Operations organization and the broader technology roadmap.
- Operate the platforms as products: establish clear product ownership, publish golden paths and paved roads, maintain a self-service catalog, and manage a prioritized backlog driven by internal customer demand rather than ticket queues.
- Define and instrument the metrics that prove platform value — platform adoption, DORA delivery metrics, change failure rate, time-to-first-deploy, developer satisfaction, and unit cost per workload — and report them to executive stakeholders.
- Reduce cognitive load across the engineering organization by abstracting infrastructure complexity behind well-documented interfaces, reusable templates, and opinionated defaults that are secure and compliant by construction.
- Set standards for infrastructure as code, policy as code, service catalogs, and platform APIs, and drive consistent adoption across the development and technology organization.

Cloud Platform:

- Lead the design, delivery, and operation of resilient multi-cloud infrastructure across AWS, Google Cloud, and Azure, including compute, container orchestration, networking, identity, storage, and observability foundations.
- Drive Kubernetes and container platform strategy, progressive delivery, service mesh and traffic management, secrets management, and zero-trust network patterns.
- Own cloud financial management (FinOps) for the platform estate — capacity planning, commitment and reservation strategy, showback and chargeback, and continuous cost optimization.
- Partner with Security, Risk, and Compliance to embed guardrails, automated control validation, and audit evidence generation directly into the platform.

Productivity Platform:

- Own the engineering productivity and developer experience toolchain — source control, CI/CD, artifact and package management, build and test infrastructure, environment provisioning, and developer portal experience.
- Continuously improve build, test, and release cycle times and pipeline reliability, treating slow or flaky developer workflows as production incidents.
- Standardize the software development lifecycle tooling and templates used across the technology organization, including supply chain security, SBOM generation, and artifact provenance.

AI Platform:

- Build and operate the shared AI platform that enables product and engineering teams to develop, evaluate, deploy, and monitor AI and machine learning capabilities safely and at scale.
- Deliver the core AI platform services: model gateway and routing across commercial and open-weight models, inference and model serving, retrieval and vector data services, feature and embedding pipelines, prompt and artifact versioning, and agent orchestration and tool integration.
- Establish MLOps and LLMOps practice — reproducible training and fine-tuning workflows, automated evaluation and regression testing, observability and tracing for non-deterministic systems, drift and quality monitoring, and safe rollback.
- Manage GPU and accelerator capacity, quota, and token consumption across clouds, with transparent cost attribution and guardrails against runaway spend.
- Partner with Security, Legal, and Privacy to operationalize responsible AI governance: data residency and retention controls, access and tenancy isolation, content and output safeguards, and auditable model usage.
- Track the AI infrastructure landscape and make deliberate build, buy, and adopt decisions as standards and tooling mature.

Production Operations and Reliability:

- Own the production health of the platform estate, including 24x7 support coverage, incident command, problem management, blameless postmortems, and corrective action follow-through.
- Define and defend SLOs, SLIs, and error budgets for platform services, and hold the organization accountable to them.
- Lead resilience engineering practice — capacity and failure-mode analysis, disaster recovery, chaos and game day exercises, and business continuity validation.
- Participate in and lead escalation for a 24x7 on-call rotation supporting critical production systems.

Leadership:

- Lead and develop a multi-team organization of engineering managers, staff and principal engineers, and platform engineers; own hiring, performance management, career development, and succession planning.
- Build a culture of engineering excellence, operational ownership, automation, measurement, and knowledge sharing across the platform organization.
- Translate complex technical strategy and risk into clear, relevant, and digestible insight for executive decision makers and non-technical stakeholders.
- Manage budget, vendor relationships, and contract negotiation for platform and AI infrastructure spend.
- Build consensus and drive change across organizational boundaries, influencing teams you do not manage.

About You:

- Bachelor's degree in Computer Science, Engineering, Management Information Systems, or equivalent practical experience.
- Minimum 12 years of experience in software engineering, infrastructure, or technology operations.
- Minimum 6 years of experience leading engineering teams, including at least 3 years leading managers or leading multiple teams simultaneously.
- Demonstrated experience building, operating, and scaling production cloud infrastructure in all three major public clouds — AWS, Google Cloud, and Azure — in a medium or large enterprise. Depth in at least one and working fluency in the other two is required.
- Proven track record building an internal developer platform or shared platform capability adopted at scale, with measurable improvements to delivery velocity, reliability, or cost.
- Deep hands-on grounding in platform engineering practice: infrastructure as code, Kubernetes and container platforms, CI/CD at scale, observability, policy as code, and self-service developer tooling.
- Working knowledge of modern AI and machine learning infrastructure — model serving and inference, retrieval-augmented generation, vector data stores, accelerator capacity management, evaluation frameworks, and AI governance and safety controls.
- Experience owning a 24x7 production environment, including incident command, SLO and error budget management, and driving systemic reliability improvement.
- Experience with cloud financial management and multi-million-dollar infrastructure budget ownership.
- Strong executive communication skills, with a demonstrated ability to build consensus, influence change across organizational boundaries, and drive results.
- Demonstrated passion for building and leading cohesive technical teams, developing leaders, and taking pride in helping individuals achieve their best.
- Willingness to participate in and lead escalation for a 24x7 on-call rotating support schedule for production systems.

Nice to Have:

- Advanced degree in a technical discipline.
- Experience standing up an AI or machine learning platform from early stage through broad enterprise adoption.
- Experience with agentic systems, tool and context protocols, and LLM application observability in production.
- Background in cybersecurity, financial services, or another regulated or real-time transaction processing industry.
- Experience operating under formal compliance regimes such as SOC 2, PCI DSS, ISO 27001, FedRAMP, or GDPR.
- Experience leading distributed or globally dispersed engineering organizations.

What's Next: