SRE Software Engineer III

JPMorgan Chase & Co.Ohio, United StatesOn-siteFull-timeSenior, 5–8 yearsListed 1 hour ago

Apply now

About this role

You will balance hands-on development with technical resiliency.

As a Site Reliability Engineer (Software Engineer III) at JPMorganChase within Consumer & Community Banking , you will help build and run resilient, scalable services by combining software engineering with operational excellence. You will improve production stability through automation, reliability engineering, and disciplined incident management, partnering across engineering and product teams to reduce risk and accelerate safe delivery.

Job Responsibilities

- Engineer and improve reliability for large-scale, distributed services by applying software and systems engineering best practices to production operations.
- Design and deliver automation that reduces manual operational work, improves recovery time, and strengthens production stability and service resilience.
- Lead response and coordination for high-priority incidents, drive issue triage to resolution, and facilitate blameless post-incident reviews that result in measurable fixes.
- Partner with application development teams throughout the software delivery lifecycle to enable sustainable releases, safer change practices, and repeatable deployments.
- Implement self-healing patterns, resilience strategies, and capacity management practices to improve availability and reduce operational toil.
- Build and evolve end-to-end observability (metrics, logs, traces) and alerting standards to enable actionable monitoring and reduce noise.
- Apply data-driven analysis of incidents and usage patterns to anticipate reliability risks and proactively prevent customer-impacting issues.
- Support a balanced operating model that includes both engineering delivery and operational responsibilities, including participation in on-call rotations as needed.
- Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Required Qualifications, Capabilities, and Skills

- Formal training or certification on software engineering concepts and 3+ years applied experience
- Hands-on software development experience in one or more general-purpose languages (e.g., Python, Java, shell scripting) supporting production-grade systems.
- Experience operating and improving reliability of distributed systems on Linux/Unix environments, including troubleshooting across application, infrastructure, and data layers.
- Experience with cloud and virtualization concepts, APIs, and modern engineering practices for scalable and fault-tolerant services.
- Experience building observability and operational diagnostics using monitoring/logging platforms (e.g., Dynatrace, Splunk, Grafana, cloud-native telemetry tools).
- Working knowledge of version control and modern delivery practices (e.g., Git-based workflows, continuous integration/continuous delivery) with a focus on quality and secure engineering.
- Strong critical thinking, incident leadership, and communication skills, with the ability to partner effectively across engineering, product, and operations stakeholders.
- Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
- Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.

Preferred Qualifications, Capabilities, and Skills

- Experience with Kubernetes and containerized workloads, including deployment patterns and operational troubleshooting.
- Experience designing frameworks that improve developer experience, release velocity, code health, and engineering standards.
- Experience with performance testing, bottleneck analysis, and capacity planning for high-throughput services.
- Experience administering application servers, web servers, and databases (e.g., Tomcat, Nginx, Oracle, MySQL) in production contexts.
- Certifications in cloud architecture or data platforms (e.g., AWS Solutions Architect or equivalent).
- 5+ years of related industry experience across application development, site reliability engineering, or DevOps in large-scale environments.