Senior Lead Software Engineer

JPMorgan Chase & Co.Palo Alto, CaliforniaOn-siteFull-timeSenior, 5–8 yearsListed 54 minutes ago

Apply now

About this role

We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.

As a Senior Lead Software Engineer at JPMorgan Chase within the Infrastructure Platforms and Foundational Services (IPFS) organization, you are an integral part of a technical team that works to enhance, build, and ensure resiliency in a secure, stable, and scalable way. As a core technical contributor, you will design and build the tools and automation that improve the reliability and operability of critical Problem management function and you will work closely with our Problem Management (SRE) and Central SRE team to deliver Automations, tools and reliability initiatives in partnership with our Financial Services (FS) domain partners.

Job responsibilities

- Executes creative software solutions across design, development, and technical troubleshooting, breaking down complex problems and thinking beyond conventional approaches.

- Builds and maintains tooling and automation that reduce operational toil while improving the stability and reliability of infrastructure platforms and services.

- Leads enterprise-authorized AI capabilities for RCA generation, incident analysis, log analytics, pattern discovery, trend analysis, and corrective-action recommendations, prioritizing elimination/automation of recurring issues.

- Facilitates deep-dive, evidence-based RCA challenge sessions with domain leads and internal teams, validating findings, contributing factors, and causal chains and driving outcomes-oriented investigation.

- Independently assesses and ensures RCA quality (sound, complete, defensible), driving blameless accountability and thorough evaluation of detection gaps, observability/monitoring weaknesses, resilience issues, process breakdowns, and human factors while leading communities of practice to advance adoption of leading-edge technologies.

- Drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within the SDLC/TLM toolchain.

- Applies knowledge of tools within the Software Development Life Cycle toolchain, including approved AI-assisted development and automation capabilities, to improve the value realized by automation at scale.

Required qualifications, capabilities, and skills

- Formal training or certification on software engineering concepts and 5+ years applied experience

- Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security

- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching senior engineers/leads on compliant usage patterns and controls.

- 7+ years leading or supporting Service Management functions, including Major Incident Management, incident response, and problem management investigations from triage through prevention.

- Proven ability to conduct deep technical RCAs in enterprise environments using disciplined methodologies (e.g., Five Whys, Fault Tree Analysis, event correlation, human factors, systemic cause analysis) and translate findings into durable corrective actions.

- Experience operating in highly regulated, mission-critical environments, with strong governance habits and attention to risk, auditability, security, and operational controls.

- Advanced expertise across multiple infrastructure and architecture domains such as networking, cloud infrastructure, Linux/Windows platforms, middleware, databases, storage, distributed systems, DevOps toolchains, enterprise monitoring, and application architecture.

- Strong hands-on SRE observability and reliability engineering: SLO/SLI design, distributed tracing, telemetry design, reliability metrics, error budget management, and leveraging AIOps platforms to reduce MTTR and improve resiliency.

- Practical proficiency with modern monitoring and telemetry tooling such as Splunk, Dynatrace, Grafana, Datadog, Prometheus, AppDynamics, Elastic, and OpenTelemetry, including instrumentation and operationalization at scale.

- Advanced programming and automation capability (e.g., Java, Go, or Python), strong CI/CD and continuous delivery practices, and demonstrated leadership in safe, compliant adoption of approved AI-assisted development tools (validation expectations, data sensitivity, and engineer coaching)

Preferred qualifications, capabilities, and skills

- Experience developing observability solutions or integrating with tools such as Grafana, Prometheus, Splunk, and infrastructure/network monitoring (e.g., SolarWinds, SCOM, SNMP/NetFlow, and storage/backup platform monitoring)

- Familiarity with site reliability engineering principles (SLIs/SLOs, error budgets, toil reduction) and partnering with SRE teams

- Experience with infrastructure as code (Terraform) and GitOps workflows

- Exposure to on-premises data-center infrastructure, networking, storage, or data protection/replication

- Experience building RESTful APIs and event-driven integrations (Kafka, RabbitMQ, SQS)

- Contributions to open-source projects or engineering communities of practice

This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries.