Rack-Scale AI Hardware Architect

MythicAustin, TexasOn-siteFull-timeMid level, 2–5 yearsListed 2 hours ago

Apply now

About this role

What You’ll Own

- Rack-level reference architecture — board, chassis, backplane, rack topology, mechanical envelope, physical constraints, and serviceability — from pathfinding through a working reference implementation.

- Scale-up and scale-out interconnect: PCIe hierarchy and switching, high-speed Ethernet and SerDes, die-to-die and board-level links, copper versus optical at each tier, retimers, connectors and cable plant, topology, oversubscription, and the bandwidth and latency budget for each hop.

- Thermal architecture for high-density racks: direct-to-chip and cold-plate cooling versus immersion, CDUs and manifolds, flow and pressure budgets, component thermal limits, and the rack-to-facility interface.

- Power delivery architecture, including rack distribution and busbar design, high-voltage DC and 48V-class conversion and efficiency, transient behavior across very large device counts, redundancy, telemetry, and RAS.

- Joint modeling with the software team of the full LLM solution — partitioning strategy, inter-layer traffic, collectives — translated into fabric, memory, and topology requirements and backed by quantitative rack- and cluster-level models of performance, power, thermal, and cost.

What We’re Looking For

- Minimum Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Computer Engineering, or a related field, with relevant experience in server, datacenter, HPC, or AI system design.

- Demonstrated architecture and design ownership of dense, rack-based server systems taken to production, including authoring the specifications ODM and OEM partners build against.

- Expert command of state-of-the-art copper and optical interconnect: PCIe, high-speed Ethernet and SerDes, DAC and twinax, backplane and cabled channels, AOCs, and pluggable or co-packaged optics, with the signal-integrity judgment to know where each approach stops working.

- Expert command of removing heat from racks, cooling architecture, cold plates, manifolds, CDUs, flow and thermal budgeting, and the facility-side interface.

- Ability to reason across interconnect, power, thermal, mechanical, and software boundaries rather than within one, including working directly with software on model partitioning, with a track record of modeling and trade studies driving architectural decisions.

Preferred Qualifications

- Master’s degree or PhD in Electrical Engineering, Mechanical Engineering, Computer Engineering, or a related field.

- Hyperscale or OCP experience, including ORv3 and contributions to relevant standards bodies.

- Experience with systems built from many small accelerators, dataflow, or wafer-scale architectures rather than only GPU-tray designs.

- Experience with LLM training or inference at scale and how parallelism strategies stress a fabric, plus familiarity with chiplets, die-to-die interfaces, and advanced packaging.

- Experience carrying a reference design through ODM or CM partners into production and into customer datacenters, including EMC, structural, and shock and vibration qualification.

What Success Looks Like

- A rack reference implementation exists, is buildable, and meets its performance, power, and thermal targets in hardware rather than in a spreadsheet.

- The partitioning and fabric story is coherent: software can map frontier-scale models across the system, and the interconnect carries the resulting traffic without being over-provisioned or starved.

- Thermal and power architectures are validated with measured data, hold headroom for the next device generation, and match a system model trusted for the next round of decisions.

- Customers and partners can deploy the design in real facilities against clearly documented power, cooling, weight, and service requirements.

- Hardware, mechanical, and software teams work from one architecture instead of three.

Why This Role Matters

- Mythic's advantage is energy per operation at the device. Delivering that advantage at scale depends on the rack around it — how it powers, cools, and moves data between a very large number of devices.

- At this density the binding constraints are rack-level: fabric topology, connector and cable plant, power distribution, and thermal. Getting them right is what turns device-level efficiency into rack-level performance.

- Model partitioning and fabric design are the same decision viewed from two directions. This role is where that decision gets made.

- If you want to define what an AI rack looks like when it is not built around a handful of kilowatt-class GPUs, we should talk.

Why Mythic

- Define the rack architecture for a fundamentally different approach to AI compute.

- Own foundational decisions with direct influence on silicon, packaging, systems, software, and product strategy.

- Work across the full stack, from model partitioning and interconnect down to busbars, cold plates, and connectors.

- Join a highly collaborative team solving difficult engineering problems across silicon, package, board, firmware, software, and systems.

- Have outsized technical impact in a senior individual-contributor role with broad organizational visibility.