Lead Software Engineer - Observability & Platform Health Monitoring

JPMorgan Chase & Co.Columbus, OhioOn-siteFull-timeSenior, 5–8 yearsListed 2 hours ago

Apply now

About this role

We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.

As a Lead Software Engineer at JPMorgan Chase within the Corporate Technology, Identity & Access Management organization, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm's business objectives.

Job responsibilities
- Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or breakdown technical problems
- Develops secure and high-quality production code, and reviews and debugs code written by others
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
- Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
- Owns the architecture and multi-quarter technical roadmap for a multi-service observability platform running on Kubernetes, including the shared internal library contracts every service consumes and the versioning and migration strategy across consumers
- Designs the data layer end to end: schema and query design against distributed SQL and large-scale Oracle datasets, plus the migration tooling and rollback strategy that keeps changes safe in production
- Defines the observability model, owning how the platform emits and publishes metrics, logs, and traces, and setting the standard for what constitutes a meaningful signal versus noise
- Serves as the senior technical voice to leadership, control partners, and internal audit, translating platform evidence into decisions and control requirements into engineering work
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Hands-on practical experience delivering system design, application development, testing, and operational stability
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Proficient in all aspects of the Software Development Life Cycle
- Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
- In-depth knowledge of the financial services industry and their IT systems
- Practical cloud native experience
Preferred qualifications, capabilities, and skills
- Hands-on depth with observability tooling at an advanced level: OpenTelemetry instrumentation, Prometheus-compatible metric backends, Splunk, and Grafana, including alert rule design, datasource configuration, and custom dashboarding beyond out-of-the-box panels.
- Production experience with Kubernetes and Helm in a regulated enterprise environment, distributed SQL (CockroachDB or another Postgres-compatible distributed database), advanced SQL against large relational datasets, and schema migration tooling such as Liquibase.
- Working knowledge of Identity and Access Management concepts including entitlements, certifications, provisioning and revocation, and SCIM, paired with experience building systems whose output is consumed by internal audit, regulators, or a formal control framework.