Executive Director - Site Reliability Engineering (Network Domain) & Argentina Site Lead

JPMorgan Chase & Co.Buenos Aires, Buenos Aires F.D.On-siteFull-timeStaff, 8–12 yearsListed 56 minutes ago

Apply now

About this role

The Executive Director, Site Reliability Engineering owns reliability for a major product domain of the global network estate, as a senior member of the Network Services SRE leadership team. This leader drives the definition and adoption of SLOs for the domain, owns and prioritises its reliability backlog, runs its reliability and post-incident reviews, and directly leads the SREs embedded in the domain alongside domain engineers aligned to the SRE practice through a dotted line. Working with the SRE leads of the other domains, they define and drive SRE culture, standards, and ways of working across the whole of Network Services.

In parallel, the Executive Director serves as the Argentina (Buenos Aires) Regional Lead for Infrastructure Platform Foundational Services (Network, Storage, DataCentre, Data Protection and Recovery). Accountable for site strategy, talent development, governance, executive presence, and operational alignment across the region. The role establishes Argentina as a strategic engineering hub by building and sustaining high-performing engineering teams aligned to global priorities and outcomes.

This leader partners closely with network engineering, Network Rapid Response (NRR), Product Engineering, and the automation platform and data engineering teams to ensure strong operational alignment, effective escalation, end-to-end service ownership, and continuous reliability improvement.

Job responsibilities

Functional Responsibilities - Site Reliability Engineering

- Own the reliability outcomes for a product domain of the network estate, and build deep understanding of its architecture, failure modes, and risk profile
- Drive the definition, instrumentation, and adoption of SLIs and SLOs across the domain's services, working through the domain's engineering teams rather than defining them in isolation, and use error-budget consumption as a live prioritisation input
- Own the domain's reliability backlog: identify the engineering work that materially improves availability, detection, and recovery, keep it prioritised, and hold it visible to domain leadership
- Run the domain's reliability reviews and lead its post-incident practice, ensuring blameless review, credible systemic analysis, and that resulting engineering fixes are tracked through to landing
- Partner with the network engineering leads in the domain to prioritise reliability work against product engineering and service delivery demand, and make the trade-offs explicit to stakeholders
- Directly manage the SREs assigned to the domain, and lead domain engineers aligned to the SRE practice on a dotted line, holding both populations to the same standards, tooling, and career framework
- Work with the SRE leads of the other domains to define, evolve, and drive SRE culture, engineering standards, common tooling, the hiring bar, and the career framework across Network Services
- Share patterns, tooling, and lessons learned across domains so reliability improvements compound rather than being rebuilt in isolation
- Translate the domain's reliability needs into requirements for the automation platform and data engineering teams, and drive adoption of the resulting capability in the domain
- Drive reduction of toil in the domain: treat recurring manual work and repeat failure as engineering defects, with measurable reduction targets
- Drive adoption of AI and agentic capability across the reliability lifecycle (triage, diagnosis, root-cause analysis, remediation, AI-assisted development) with clear validation standards, so speed never compromises correctness, security, or risk
- Hold a code-first engineering bar: production-quality code, testing, code review, and CI/CD, so reliability work ships as software rather than scripts
- Deliver measurable outcomes for the domain: improved availability, reduced detection and recovery times, fewer repeat incidents, reduced manual touch, and improved change success rate

Regional Responsibilities - Argentina IP Foundational Services Site Lead

- Represent the Argentina hub in global Infrastructure Platforms Foundational Services forums
- Drive site culture, retention, and consistent engineering standards across functional silos
- Ensure effective cross-functional collaboration with security, infrastructure, application teams, service management, and business partners to deliver integrated outcomes
- Manage regional resource planning and budget inputs: capacity forecasting, skills coverage, on-call sustainability, and investment recommendations tied to measurable service improvements
- Hire, develop, and retain engineering talent across the Argentina footprint
- Ensure governance, risk, and audit compliance for in-country operations

Leadership Expectations

- Strong ownership of reliability outcomes, with the technical depth to lead engineers rather than only manage them
- Ability to build and scale engineering teams across functional and matrixed reporting lines, including dotted-line reports
- Influences peers and partner teams without direct authority, and drives change from within a leadership team rather than from the top of one
- Drives the transformation from an operations model to an engineering model
- Strong executive communication skills, including calm, credible communication during high-severity events
- Maintains compliance and control discipline
- Credible senior technology presence in-region, able to represent JPMC externally with regulators, universities, and partners

Required Qualifications, Capabilities, and Skills

- 10+ years in infrastructure, production, or reliability engineering, including leadership roles running systems at scale
- Demonstrated ownership of SLI/SLO/error-budget practice, incident and post-incident leadership, and measurable toil reduction in a production environment
- A code-first foundation: credible software engineering background and the judgment to hold a code-first SRE bar rather than an operations-only one
- Deep observability expertise: white-box and black-box monitoring, SLO-based alerting, and telemetry
- Proven track record adopting agentic AI and LLM-driven capability in production engineering and operations environments
- Experience operating within formal risk and control frameworks: audit engagement, evidence quality standards, and remediation governance
- Demonstrated experience in incident management, problem management, and change governance, including executive communications during high-severity events
- Proven ability to hire, develop, and retain strong engineers, and to raise the engineering bar of an existing team
- Strong global collaboration skills across distributed, matrixed organizations
- Prior site, country, or regional engineering hub leadership experience
- BS/BA degree or equivalent practical experience in technology, engineering, or a related discipline

Preferred Qualifications, Capabilities, and Skills

- Networking depth (routing, switching, security, packet and flow analysis), or experience leading reliability for network or network-adjacent platforms. A strong plus, not a gate
- Experience running an embedded SRE model, driving reliability into engineering teams from within rather than from a central operations silo
- Experience establishing or growing an SRE practice as part of a leadership team, including influencing peer domains
- Financial services or other regulated environment experience
- Experience with large-scale network automation (Python, APIs, config management, validation, and remediation at scale)
- Demonstrated use of AI to redesign engineering and operational workflows for measurable impact, and to build organizational AI fluency
- Proven operational governance track record: incident, problem, and change management, and measurable service improvement plans

Team Scope

- Direct: a domain-aligned SRE team embedded with the network engineering and product teams for the domain, growing
- Dotted line: domain engineers aligned to the SRE practice, held to the same standards, tooling, and career framework
- Reporting line: into the Head of SRE for Network Services, as a peer of the other domain SRE leads on the SRE leadership team
- Partnership: strong partnership with network engineering leads in the domain, the NRR organization, global Product Engineering, the automation platform and data engineering teams, and regional Infrastructure leadership
- Regional (matrix): site leadership for all IP Foundational Services engineers in Argentina