Production Systems Engineer

MetaDublin, LeinsterOn-siteFull-timeMid level, 2–5 yearsListed 1 hour ago

Apply now

About this role

Meta is seeking a Production Systems Engineer to join our Release to Production (Release to Production) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services. The RTP team is responsible for the end-to-end Hardware Lifecycle of all Meta servers, from exploration and development to production health. RTP Engineers work closely with Production Engineering teams, Enterprise Networking, Hardware Designers, Networking Teams, Manufacturers, Vendors, Datacenter Operation teams and New Product Introduction teams to ensure the smooth operation of systems across the planet.

We encounter problems from the very smallest of scales (errors occurring at the microscopic scale, within single registers of a CPU) up to the very largest - deploying solutions to our entire millions-wide fleet. We look for people with proven experience of finding solutions to complex issues, navigating undefined problem spaces and delivering measurable outcomes, who want to tackle the hardest problems in the domain.

Typically we will hire engineers from backgrounds such as Site Reliability Engineer (SRE), Software Engineer, Systems Engineer, Systems Development Engineer, DevOps Engineer, Systems Administrator, or similar. You will have a history of driving projects to successful business outcomes. Your previous experience will always be less important than demonstrated problem solving abilities.

Responsibilities

Build and develop tooling solutions to automate business critical processes in service of managing the health of the Meta production fleet
Troubleshoot, diagnose and root cause system failures, working with key partners to identify and deliver solutions
Proactively identify opportunities to fix or enhance tooling, hardware and processes
Build subject matter expertise in one or more specialist areas including Firmware Deployment; Edge/CDN hardware; or Diagnostic Testing

Qualifications

Bachelor's degree in Computer Science, related technical discipline, or equivalent work experience
3+ years of experience coding in a higher-level language (Python, PHP, Java, Go, Rust, C++)
Experience building, maintaining and debugging production services or platforms - usually (but not necessarily) in a Linux/Unix environment
Knowledge of server architecture and components across Compute/Storage/AI Systems/Networking
Scientific approach to troubleshooting, root-cause analysis and investigation
Demonstrated track record of crafting clear and concise reporting for a broad range of stakeholders
Demonstrated track record of productive collaboration across organization boundaries Experience with configuration management or infrastructure-as-code tools (e.g., Ansible, Puppet, Chef, Terraform)
Experience with large-scale fleet management, including managing systems at scale across data centers