About this role
Position description
We are still at the beginning of our growth journey, so we are putting new processes, technologies, and tools in place on a continuous basis. Your role is a pivotal engineering contributor to the tooling and services to automate and enhance the software development lifecycle, empowering our fellow SonarSourcers to deliver with speed, confidence, and security. You would be a member of a team that delivers solutions across all of our 5 offices: Austin (Texas, US), Geneva (Switzerland), Bochum (Germany) and Singapore.
As a Major Incident Manager, you use and create automation tools to monitor and observe production infrastructure services both on premises and in the cloud. You are allergic to repetitive tasks, preferring to maximize automation and reliability. You are expert in change management, infrastructure management, system support, and configuration management.
What you will do
- Cloud platform engineering through code: Design and evolve secure, reliable AWS foundations using reusable infrastructure-as-code, automated deployment pipelines, and tested change controls.
- Automation and self-service: Develop pragmatic tooling, integrations, and glue code that solve operational problems for Systems Engineering and make internal customers more self-sufficient.
- AI-enabled platform engineering: Use and develop AI-enabled tooling and automation to improve infrastructure delivery, operational triage, knowledge retrieval, and self-service, with appropriate human review, security, and governance.
- Reliability and observability: Define and improve monitoring, alerting, dashboards, runbooks, SLIs/SLOs, capacity planning, and service-health reporting for Systems Engineering services.
- Security engineering: Design, architect, implement, and operate secure cloud controls — including least-privilege access, policy-as-code, secrets management, auditability, configuration baselines, and automated evidence collection — in partnership with Information Security.
- Incident response: Participate in the Systems Engineering on-call and escalation model for services and tools owned by Systems Engineering. Lead or contribute to incident response, root-cause analysis, and preventative follow-up work.
- Resilience engineering: Lead or facilitate FMEAs and recovery exercises for critical cloud services and material changes; translate findings into funded, owned engineering actions.
- Technical leadership: Produce lightweight designs, establish standards, review complex changes, mentor peers, and influence service-owning teams to adopt reliable platform patterns.
- Collaboration: Work closely with Systems Engineering, S/NOC, Information Security, IT Ops, Product Engineering, and external providers to define clean service boundaries and effective escalation paths.
Experience and qualifications
- 8+ years of hands-on experience in cloud, platform, site reliability, infrastructure, or DevOps engineering in a SaaS, cloud-native, or complex enterprise environment.
- Strong AWS experience, including networking, IAM, DNS, logging/monitoring, compute, storage, security controls, and multi-account or multi-environment design.
- Deep infrastructure-as-code experience with AWS CDK, Terraform, CloudFormation, or equivalent, including reusable modules, automated testing, code review, and CI/CD delivery.
- Strong software engineering and automation skills in at least one of Python, TypeScript, Go, Bash, or a comparable language with a heavy emphasis on AI skillsets.
- Experience designing production-grade observability: metrics, logs, traces, actionable alerting, dashboards, and operational runbooks.
- Demonstrated ability to automate operational controls and remove manual work while preserving appropriate approvals, auditability, and safety.
- Experience with incident response, post-incident engineering, root-cause analysis, and translating lessons into measurable reliability improvements.
- Sound knowledge of cloud security principles: identity and access management, least privilege, encryption, secrets handling, secure networking, vulnerability management, and logging.
- Excellent written and verbal English communication; able to explain technical decisions, risks, and trade-offs to both technical and non-technical stakeholders.
In-office culture
We're intentional about this. We believe the best teams are built in the room together. Three anchor days — Mondays, Tuesdays, and Thursdays — create the collaboration rhythm that makes a hub office worth having.
Candidates need to be genuinely based in the location the role is posted — if that's not where you are today, we're happy to support relocation for the right person.
We value diversity, equity, and inclusion
At Sonar, we believe that our diversity is our strength. We are a global company that values and respects different backgrounds, perspectives, and cultures. We are committed to fostering a diverse and inclusive work environment where everyone feels valued and empowered to contribute their best. We are proud to be an equal opportunity employer and welcome all qualified applicants, regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status.
If you need any accommodation, please reach out to us at [email protected].
All offers of employment at Sonar are contingent upon the results of a comprehensive background check and reference verification conducted before the start date.
Applications that are submitted through agencies or third party recruiters will not be considered.