Site Reliability Engineer, Principal

SynopsysCalifornia, United StatesOn-siteFull-timePrincipal, 12–15+ yearsListed 9 hours ago

Apply now

About this role

Job Description and Requirements

We Are

Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with our customers across a wide range of industries to maximize their R&D capability and productivity, powering innovation today that ignites the ingenuity of tomorrow.

You Are

You are someone who naturally looks ahead. You do not wait for infrastructure problems to become urgent before paying attention to them you look for patterns, question assumption and think carefully about what may happen next. You are comfortable working in complex environments where priorities can shift, information may be incomplete and there is rarely a single obvious answer.

You approach challenges with curiosity and structure. You enjoy understanding how systems behave, why issues occur and where small improvements can make a meaningful difference over time. You think about reliability not only in terms of fixing what is broken, but in creating environments that are easier to operate, easier to scale and less likely to surprise the teams relying on them. You are thoughtful about trade offs . You understand that the technically perfect solution is not always the right business decision, and you are comfortable balancing reliability, performance, cost, speed, and long term sustainability. You communicate these trade offs clearly, helping others understand both the immediate impact and the longer-term consequences of decisions.

You enjoy bringing people together around complex problems. You can work independently and take ownership, but you also value collaboration and seek input from engineering, operations, architecture, and business partners. You ask thoughtful questions, challenge ideas constructively, and build trust through transparency and follow-through. Above all, you are motivated by making things better. You look for ways to simplify complexity, remove recurring pain, improve how teams work, and create systems and processes that remain strong as the organization grows.

What You'll Be Doing

- Own the enterprise capacity strategy and roadmap across compute, storage, network , and cloud to ensure scalable, resilient growth.
- Lead long-range demand forecasting and scenario planning, translating business direction into clear infrastructure investment recommendations.
- Build executive-ready dashboards and narratives that connect utilization, performance, availability, and cost trends to actionable decisions.
- Architect and deliver automation that improves capacity planning accuracy, reduces operational toil, and prevents recurring reliability issues.
- Establish governance mechanisms—standards, planning cadences, and prioritization workflows—that align teams on resource allocation and service commitments.
- Drive reliability improvements by defining and operationalizing SLOs, strengthening observability, and proactively mitigating scalability and capacity risks.
- Lead response and follow-through for major infrastructure events, ensuring rapid stabilization and durable root-cause elimination.

The Impact You Will Have

- Enable predictable, scalable growth by ensuring the right capacity is available ahead of demand across on-prem and cloud environments.

- Reduce unplanned downtime and performance degradation by surfacing capacity and reliability risks early and driving preventative action.
- Improve infrastructure cost efficiency by increasing utilization, eliminating waste, and guiding smarter investment decisions with data-backed scenarios.
- Accelerate decision-making for senior leaders by providing clear, trusted forecasts, tradeoffs, and risk assessments tied to business priorities.
- Increase operational maturity by embedding SLO-driven reliability practices, stronger observability, and consistent governance across teams.
- Shorten recovery times and prevent repeat incidents by driving durable root-cause elimination and automation that reduces human dependency.
- Strengthen cross-functional alignment by creating shared visibility and accountability for capacity, availability, and service commitments.

What You'll Need

- You have deep, hands-on understanding of enterprise infrastructure across compute, storage, networking, and cloud, and you can reason about how changes in one domain affect reliability in another.
- You bring a strong quantitative mindset, using statistical thinking, forecasting, and structured analysis to turn messy operational data into clear recommendations.
- You have a track record of defining meaningful SLIs/SLOs, instrumenting services for observability, and building alerting that drives action rather than noise.
- You bring a proactive reliability approach, designing automation and guardrails that prevent recurring issues and reduce operational toil.
- You have led large, cross-functional infrastructure initiatives end-to-end, aligning stakeholders and delivering outcomes through ambiguity and competing priorities.
- You communicate with clarity and influence, tailoring technical depth to the audience and building executive-ready narratives that support investment decisions.
- You bring a continuous-improvement mindset, challenging the status quo and simplifying processes to raise service maturity over time.

Who You Are

- You are at your best in ambiguity, turning incomplete inputs into clear options, tradeoffs, and next steps without waiting for perfect data.
- You approach problems with disciplined curiosity, digging past symptoms to understand system behavior and prevent repeat failures.
- You are skilled at aligning diverse stakeholders, translating between engineering details and business priorities to drive timely decisions.

- You communicate with precision and calm during high-pressure moments, keeping teams focused on restoration, learning, and follow-through.
- You are intentional about simplification, consistently reducing toil and process friction so teams can operate reliably at scale.

Rewards and Benefits

We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings. Your recruiter will provide more details about the salary range and benefits during the hiring process.

At Synopsys, we want talented people of every background to feel valued and supported to do their best work. Synopsys considers all applicants for employment without regard to race, color, religion, national origin, gender, sexual orientation, age, military veteran status, or disability.

In addition to the base salary, this role may be eligible for an annual bonus, equity, and other discretionary bonuses. Synopsys offers comprehensive health, wellness, and financial benefits as part of a competitive total rewards package. The actual compensation offered will be based on a number of job-related factors, including location, skills, experience, and education. Your recruiter can share more specific details on the total rewards package upon request. The base salary range for this role is across the U.S.