About this role
Guide and shape the future of technology at a globally recognized firm, driven by pride in ownership.
As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Infrastructure Platforms team, you are the non-functional requirement owner and champion for the applications in your remit. You are a key influencer in your team’s strategic planning, driving continual improvement in customer experience, resiliency, security, scalability, monitoring, instrumentation, and automation of the software in your area. You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.
Job responsibilities
- Shows expertise in site reliability principles and demonstrates what it means to strike the balance between features, efficiency, and stability, negotiating with peers and executive partners to ensure optimal outcomes
- Drives the adoption of site reliability practices throughout the organization and ensures teams demonstrate site reliability best practices empirically through stability and reliability metrics
- Manage a small team of enthusiastic SREs with varying experiences.
- Accountable for their Annual performance reviews, 1:1 and career development.
- Drives a culture of continual improvement by encouraging real-time feedback to improve the customer’s experience
- Ensures your team collaborates with other teams within your specialization and avoids duplication of work where possible
- Follows an objective, data-driven, post-mortem strategy by conducting regular team debriefs that enable team members to learn from successes and mistakes
- Coaches and develops entry to mid-level team members through tailored feedback and growth plans
- Ensures your team documents and shares their knowledge and innovations via internal forums, communities of practice, guilds, and conferences
- Establishes team standards for AI-assisted reliability workflows across automation and delivery practices, ensuring traceability/auditability, resiliency, and security controls.
- Drives reuse-first adoption of enterprise-authorized AI capabilities within the work environment to improve reliability operations and customer experience outcomes, with human-in-the-loop validation and appropriate handling of sensitive data.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience . I
- Demonstrates advanced proficiency in site reliability culture and principles and can demonstrate how to implement site reliability across application and platform teams while avoiding common pitfalls
- Experience leading teams in the safe use of enterprise-authorized AI capabilities within the work environment for reliability engineering workflows, including validation habits and awareness of data sensitivity.
- Ability to set and reinforce organization-level practices for reviewing AI-assisted recommendations and escalating uncertain decisions while maintaining resiliency, security, and auditability outcomes.
- Experience leading technologists to manage and solve complex technological issues on an organizational level
- Influences the team’s culture by championing innovation and change for success
- Experience hiring, developing, and recognizing talent
- Proficient in at least one programming language such as Python, Java etc., and Web frameworks like Django Flask etc.,
- Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI, etc.)
- Proficient with CI/CD practices and related tooling, container/container orchestration, and troubleshooting common networking technologies and issues
Preferred qualifications, capabilities, and skills
- Ability to initiate and implement ideas to solve business problems
- Passion for learning new technologies and driving innovative solutions.