Production Services Lead

Chubb ExternalColombiaOn-siteFull-timeSenior, 5–8 yearsListed 1 hour ago

Apply now

About this role

. Job Summary

We are seeking a leadership role to run NA Commercial Insurance Production Services and Site Reliability Engineering (SRE) function out of engineering center in Bogota. This role will be responsible for the reliability, availability, and performance of critical production systems. In this role, you will bridge software engineering and operations, driving the cultural and technical transformation toward engineering-based operations. You will collaborate closely with SREs, development, testing and business teams to ensure our systems are robust and scalable effectively providing 99.9% availability. While the focus is on technical engineering, you will also mentor junior team members and share best practices with regional teams.

---

Key Responsibilities

Leadership & Collaboration

- Lead a team of SREs; mentor engineers on reliability practices
- Team building and performance evaluation of production services/SRE staff in CECC
- Partner with development, testing and business teams to embed reliability requirements into the SDLC
- Partner with development, infrastructure teams to ensure lower environment stability across CI footprint
- Act as the escalation point for major incidents and production crises Communicate effectively with business and operations partners, especially during critical system outages, client escalations

Reliability & Availability

- Partner with technology and business partners to define and own SLAs, SLOs, SLIs, and error budgets across CI portfolio
- Lead incident response, blameless post-mortems, and remediation follow-through
- Drive reduction in MTTR and MTTD across the production estate

Process & Governance

- Enforce change management, release gating, and production readiness reviews Establish on-call practices, runbooks, and operational playbooks
- Report on production health metrics to senior stakeholders
- Partner with Bogota leadership team in providing oversight into AMS MSM vendor engagement. This may include weekly/monthly governance calls, site visits, performance evaluation etc.

Platform Engineering

- Design and implement observability frameworks (metrics, logs, traces)
- Build self-healing automation to reduce toil and manual intervention
- Own capacity planning, performance baselining, and scalability initiatives

---

Skills & Experience

Required:

- 10+ years in SRE, platform engineering, or production operations
- 5+ years in a lead or senior individual contributor capacity
- Strong reasoning, analytical thinking and troubleshooting skills for applications, including RCA & memory debugging.
- Experience with observability and monitoring tools (e.g., ELK Stack, Application Insights, Splunk, Kibana, AppDynamics, DynaTrace).
- Basic/intermediate knowledge of databases such as MS SQL Server.
- Strong communication skills and ability to work under pressure.

Nice to Have:

- Experience with AI technologies (e.g. Claude) and their application in enhancing system reliability and performance.
- Experience with application production support/SRE management
- Background in regulated or high-compliance industries.
- Familiarity with chaos engineering, performance optimization or fault injection.
- Familiarity with Azure cloud infrastructure and services (e.g. PaaS and identity management such as Active Directory, Azure AD).

Soft Skills:

- Excellent verbal and written communication skills; must have strong experience in working with senior technology and business stakeholders
- Proactive, detail-oriented, and able to handle production-critical issues.
- Collaborative mindset and willingness to mentor others.

---

What Success Looks Like

- Rapid, effective resolution of incidents and performance issues.
- High uptime and reliability for all critical applications.
- Continuous improvement in automation and operational efficiency.
- A culture of reliability and technical excellence within the team.