Software Engineer III – MLOps

JPMorgan Chase & Co.Plano, TexasOn-siteFull-timeSenior, 5–8 yearsListed 3 hours ago

Apply now

About this role

As a Software Engineer III – MLOps at JPMorgan Chase as a part of Banking and Wealth Management Technology, you build and operate a machine learning operations (MLOps) platform on Amazon Web Services (AWS) that supports model development, CI/CD, deployment, monitoring, and governance across multiple environments

We are building a secure, resilient machine learning delivery ecosystem that helps teams move faster without compromising controls. In this role, you will develop platform foundations—Infrastructure as Code, Kubernetes patterns, multi-environment governance, and operational excellence—so product and data teams can deploy and operate models with confidence. You will collaborate across engineering, data, and risk partners to standardize best practices, strengthen reliability, and improve developer experience. The work spans build, run, and continuous improvement across multiple environments.

You will contribute to platform resiliency through high-availability designs, multi-Availability Zone patterns, capacity planning, and incident readiness. You will define observability standards to ensure services are measurable, diagnosable, and supportable in production. Your work will help teams deliver batch and online inference safely with standardized deployment and rollback approaches.

Job responsibilities

- Five years of experience in platform engineering, software engineering, or machine learning operations roles.
- Hands-on experience building and operating production platforms on AWS.
- Hands-on experience with Kubernetes, including Amazon EKS, Helm or Kustomize, ingress, autoscaling, and workload scheduling.
- Hands-on experience using Terraform for repeatable environment provisioning and controlled change management.
- Demonstrated experience using approved AI-assisted software development tools to support coding, code review, testing, troubleshooting, and documentation, with clear validation practices for correctness, performance, and security.
- Knowledge of responsible AI use in engineering workflows, including data sensitivity awareness and secure handling of inputs and outputs.
- Experience implementing security controls for cloud and Kubernetes platforms (for example: identity and access management, role-based access control, secrets management, network policies, and vulnerability management).
- Experience designing and operating highly available systems, including on-call readiness, incident response, and post-incident reviews.
- Experience establishing or operating observability practices across metrics, logs, traces, dashboards, alerting, and service-level objectives (where defined).
- Experience building and operating streaming pipelines with Apache Kafka and Apache Flink.
- Proficiency in one or more programming languages used in platform or streaming engineering (for example: Python, Java, Scala, or Go), plus CI/CD automation experience.

Required Qualifications, Capabilities, and Skills

- 5+ years (or equivalent) in platform engineering, Java, or MLOps roles.
- Hands-on experience building production platforms on AWS.
- Strong Kubernetes experience, including Amazon EKS, Helm/Kustomize, ingress, autoscaling, and workload scheduling.
- Strong Terraform experience for repeatable environment provisioning and change management.
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Experience implementing security controls for cloud/Kubernetes platforms (IAM, RBAC, secrets, network policies, vulnerability management).
- Proven ability to deliver resilient, highly available systems and to operate them (on-call readiness, incident response, postmortems).
- Strong observability experience (metrics/logs/traces, dashboards, alerting, SLOs).
- Streaming experience with Kafka and real-time data processing; experience building/operating Flink pipelines is required.
- Solid programming/scripting skills (commonly Python, Java/Scala for Flink, and/or Go) and CI/CD automation experience.

Preferred Qualifications

- Experience enabling GPU workloads on Kubernetes (for example: device plugins, node pools, scheduling, and performance tuning).
- Familiarity with machine learning platform components such as model registries, experiment tracking, feature stores, and artifact management.
- Experience with service mesh, policy as code, and container supply chain security.
- Experience integrating with data lake or data warehouse platforms and implementing data quality monitoring.
- Experience working in regulated environments with strong engineering governance.