Sr. Manager, SRE, Operations & Product Support

American Bureau of ShippingHouston, TexasHybridFull-timePrincipal, 12–15+ yearsListed 13 hours ago

Apply now

About this role

ABS Group Digital Solutions is seeking a Lead, Site Reliability Engineering, Operations & Product Support to establish the production reliability and support operating model for its next-generation Fleet Management System (FMS). The role will help take FMS from development through beta, customer migration, and scaled SaaS operations. It combines hands-on reliability engineering with leadership of incident response, operational readiness, technical support escalation, and post-launch improvement.

Working with the Platform Engineering & Cloud Architecture lead, application engineering, QA, security, and customer support teams, this person will ensure the product is observable, recoverable, and supportable. The lead will use conventional and AI-assisted automation to streamline operations while establishing clear boundaries between technical SRE ownership and customer-facing support ownership.

What You Will Do:

- Define service health measures and reliability objectives for FMS, including availability, latency, error rates, recovery expectations, and appropriate service-level indicators and objectives.
- Establish end-to-end observability across applications, infrastructure, integrations, and AI-enabled services through actionable logs, metrics, traces, dashboards, health checks, and alerts; assess where AI-assisted anomaly detection can improve signal quality.
- Lead the technical incident-response model, including severity definitions, on-call and escalation practices, incident coordination, recovery procedures, and post-incident reviews; use AI-assisted summarization and evidence gathering where it improves response without replacing human judgment.
- Work with engineering and platform teams to design for resilience, performance, capacity, backup and recovery, and safe operation under expected customer and data growth.
- Define and implement operational-readiness criteria for beta and production releases, including monitoring, runbooks, rollback plans, ownership, support handoffs, and post-release validation.
- Establish the technical support escalation model and partner with customer support, product, and engineering to resolve issues and turn recurring incidents and tickets into permanent fixes; evaluate AI-assisted ticket categorization and knowledge retrieval to speed technical triage.
- Support customer migrations, go-lives, and post-launch stabilization by preparing technical monitoring and response plans, triaging production issues, and incorporating lessons into repeatable procedures.
- Automate routine operational tasks, health checks, deployment verification, incident triage, and recovery; use AI where it demonstrates improved speed or accuracy, with access controls, auditability, and human approval for production-impacting actions.
- Track reliability, incident, supportability, and operational-efficiency trends; communicate risks, corrective actions, and progress to engineering and program leadership, including evidence of whether AI-assisted workflows reduce toil or improve outcomes.
- Help build and mentor an SRE/production-operations capability as FMS moves from initial releases to scaled customer use.

What You Will Need:

Education and Experience

- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent relevant experience.
- 10+ years of relevant experience in site reliability engineering, production engineering, cloud operations, or software operations, including experience leading incident response or operational improvement across teams.
- Demonstrated experience operating production web or SaaS services and improving their reliability through software engineering and automation.
- Experience establishing observability, on-call practices, runbooks, and service-health or reliability measures for production systems.
- Experience partnering with software engineering, platform engineering, and customer-facing support teams during releases, incidents, and customer go-lives.
- Experience applying AI-assisted tools or workflows to technical operations, incident triage, monitoring analysis, support knowledge retrieval, or operational automation, with an understanding of how to validate results before use in production.
- Experience with Azure cloud preferred.

Knowledge, Skills, and Abilities

- Strong command of SRE practices, including SLIs/SLOs, incident response, root-cause analysis, performance, capacity, resilience, and disaster recovery.
- Experience operating production SaaS applications on Azure, including compute, networking, identity, storage, containers, databases, integrations, and security.
- Ability to build effective observability and on-call practices using telemetry, logs, metrics, traces, and tools such as Azure Monitor, Application Insights, and Log Analytics-without creating unnecessary alert noise.
- Ability to automate secure deployments and operational workflows using scripting, APIs, CI/CD, infrastructure as code, managed identities, and Key Vault.
- Sound judgment on release risk, rollback, customer impact, and the responsible use of AI-assisted operations, including data protection and human oversight of production-impacting actions.
- Clear communication and collaborative leadership across engineering, support, and business stakeholders, including the ability to drive improvements without direct ownership of every team or system.

Reporting Relationships:

Reports to the Sr Director, Platform Engineering & Cloud Architecture. The role will initially lead cross-functional operational practices; direct-report scope will be determined as the SRE and production-operations capability scales. Customer-facing support teams retain ownership of routine customer communications and frontline support.