Sr. Director of Engineering, Data Center Infrastructure

AMDSeattle, WashingtonOn-siteFull-timeStaff, 8–12 yearsListed 2 hours ago

Apply now

About this role

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.

THE ROLE

AMD is seeking a visionary Senior Director of AI Data Center Infrastructure Engineering to lead the strategy, architecture, deployment, and operation of next-generation AI data centers and large-scale compute environments.

This highly visible leadership role will be responsible for defining AMD's AI infrastructure architecture across rack-scale and cluster-scale deployments, spanning facility infrastructure, compute systems, networking, software operations, and reliability engineering. The successful candidate will lead a world-class organization of engineers and architects responsible for building scalable AI infrastructure that supports AMD's internal development, validation, and customer enablement initiatives.

This position requires a unique combination of executive leadership, systems thinking, and deep technical expertise across data center architecture, networking, distributed systems, cluster operations, and large-scale infrastructure deployment.

THE PERSON

The ideal candidate is a highly experienced technology executive with a proven track record of building and operating large-scale data center environments. You thrive in complex, cross-functional organizations and have demonstrated success leading large engineering teams while influencing strategy across hardware, software, systems, and operations.

You possess deep technical expertise in AI infrastructure and can translate business objectives into scalable engineering solutions. You are equally comfortable discussing system architecture with technical experts and presenting strategic recommendations to executive leadership, customers, and industry partners.

KEY RESPONSIBILITIES

AI Data Center Architecture & Infrastructure Strategy

- Define AMD's end-to-end AI data center infrastructure strategy, including compute architecture, facility requirements, operational models, and scalability planning.
- Lead architecture and deployment decisions for large-scale AI and HPC data center environments spanning thousands of servers and accelerators.
- Drive infrastructure planning across power distribution, cooling strategies, space utilization, capacity forecasting, and operational readiness.
- Analyze performance, efficiency, reliability, and total cost of ownership (TCO) tradeoffs to optimize infrastructure investments.
- Partner with internal engineering organizations, facilities teams, and external vendors to deliver world-class AI infrastructure platforms.

Systems Architecture & Platform Engineering

- Own architecture decisions spanning silicon, servers, rack-scale solutions, and full data center deployments.
- Collaborate closely with hardware, software, networking, systems design, platform engineering, and product organizations to align infrastructure capabilities with AMD technology roadmaps.
- Drive architecture reviews and technical decision-making across mechanical, electrical, thermal, firmware, software, and systems domains.
- Establish scalable infrastructure standards and design methodologies for future AI platforms.

Cluster Management, Operations & Reliability

- Lead the design, deployment, and lifecycle management of large-scale AI training and inference clusters.
- Establish operational excellence through Site Reliability Engineering (SRE) principles, observability frameworks, incident management processes, and capacity planning.
- Develop automation strategies that improve deployment speed, infrastructure efficiency, system reliability, and operational scalability.
- Drive continuous improvement initiatives focused on availability, resilience, performance, and cost optimization.
- Define reliability and service-level objectives (SLOs) for production infrastructure environments.

Data Center Networking & Interconnect Architecture

- Define and optimize networking architectures supporting large-scale AI and HPC environments.
- Lead strategy across network topology, high-speed fabrics, switching architectures, routing infrastructure, and accelerator interconnect technologies.
- Evaluate and implement next-generation networking technologies to maximize cluster performance and scalability.
- Champion software-defined networking (SDN), network automation, and observability solutions across global infrastructure deployments.
- Partner with AMD product and ecosystem teams to influence future networking roadmaps.

Executive Leadership & Organizational Development

- Lead, mentor, and grow a global organization of 50-75 engineers, architects, and technical leaders.
- Establish a high-performance engineering culture focused on innovation, accountability, collaboration, and execution excellence.
- Develop organizational strategy, succession planning, workforce development, and talent acquisition initiatives.
- Serve as a key technical advisor to executive leadership on AI infrastructure strategy, investment priorities, and industry trends.
- Collaborate with strategic customers, system integrators, technology partners, and suppliers to drive alignment across the AI ecosystem.
- Monitor competitive technologies, emerging market trends, and industry innovation to influence AMD's long-term infrastructure vision.

PREFERRED EXPERIENCE

- Experience designing, deploying, and operating large-scale data center environments.
- Experience leading engineering organizations, including ownership of large multidisciplinary teams.
- Demonstrated success building and scaling AI, cloud, hyperscale, HPC, or enterprise data center infrastructure.
- Extensive experience with data center architecture, power delivery, thermal management, capacity planning, and facility operations.
- Deep knowledge of AI infrastructure, server architectures, accelerator-based computing platforms, and rack-scale system design.
- Proven expertise designing high-performance data center networks, including switching, routing, Ethernet fabrics, RDMA, InfiniBand, and next-generation interconnect technologies.
- Experience leading compute cluster deployment, lifecycle management, and large-scale distributed system operations.
- Strong understanding of Site Reliability Engineering (SRE), Infrastructure as Code (IaC), automation frameworks, monitoring, observability, and operational best practices.
- Proven ability to lead technical organizations through significant technology transformations and rapid growth.
- Experience influencing cross-functional roadmaps involving hardware, software, networking, and infrastructure teams.
- Strong executive communication skills with demonstrated ability to influence senior leadership, customers, and external partners.

ACADEMIC CREDENTIALS

- Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, Mechanical Engineering, or related technical field.
- Master's degree preferred.
- PhD considered a plus.

This role is not eligible for visa sponsorship.

#LI-CJ1

Benefits offered are described:  AMD benefits at a glance .

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.