About this role
Standard Job Description
Your Mission:
The AI Factory team is seeking a Site Reliability Engineer to improve the reliability of the platforms and services that support AI development and deployment. In this role, you will use software development and automation to identify operational problems, reduce repetitive work, and help teams deliver dependable systems.
You will work with engineers and other partners across the AI Factory to improve observability, incident response, release validation, and platform health. The role is broader than any single tool or product: you may contribute to reliability platforms such as Cluster Concierge, but your focus will be on solving reliability problems across the environment.
Key Responsibilities:
· Improve the scalability, resilience, and reliability of existing platform services and middleware to ensure they remain dependable as usage and demand grow.
· Develop and maintain automation and operational tooling that improve platform reliability and reduce recurring manual work.
· Help teams investigate incidents, identify contributing factors, and implement fixes that prevent repeat issues.
· Build and improve dashboards, alerts, and other observability capabilities using metrics, logs, and traces.
· Create automated tests and validation workflows for upgrades, releases, and changes to platform services.
· Contribute to CI/CD and GitOps workflows that support consistent, reliable deployments.
· Assess platform health, document findings, and work with partner teams on practical reliability improvements.
· Participate in design reviews, code reviews, testing, and incident reviews.
· Contribute to reliability improvements for AI Factory services, including AIF Up and tools that support health validation and investigation.
Responsible for autonomy hardware and software system integration and testing including verification and validation.Translates customer requirements into product and systems specifications while addressing technical, schedule, and cost considerations; Establishes functional and technical specifications and standards for autonomous systems; Determines sensing hardware components; Recommends and selects appropriate control systems; Integrates and optimizes the autonomous compute and sensing hardware and software; Solves hardware/software interface problems; Develops plan(s) to integrate autonomous functionality into product(s) and platform(s) and other system(s); Develops test plans for validation and verification and procedures for test and evaluation requirements in collaboration with designers and developers; Performs integration testing and coordinates subsystem and/or system testing activities for autonomous programs; Documents and conducts analysis of test results and recommends fixes to software, hardware components, subsystems and systems; Interfaces with other teams involved the development lifecycle for perception
Basic Qualifications
· Experience designing, developing, and maintaining production software or automation using Go, Python, or a comparable language.
· Experience operating or engineering Kubernetes-based platforms, including troubleshooting complex service or infrastructure issues.
· Experience building or improving CI/CD, GitOps, infrastructure automation, or deployment workflows.
· Experience using observability data—including metrics, logs, or traces—to diagnose problems and improve system health.
Desired Skills
· Experience with OpenShift, GitLab CI/CD, Argo CD, Argo Rollouts, or similar GitOps and progressive-delivery tooling.
· Experience improving incident response, reducing operational toil, or defining actionable service-health measures.
· Familiarity with Prometheus, Grafana, OpenTelemetry, or comparable observability tools.
· Familiarity with designing automated reliability tests, upgrade validation, resilience tests, or failure-mode analyses.
· Familiarity with AI/ML platforms, GPU-based infrastructure, or deployments in disconnected environments.
· Strong oral and written communication skills, and ability to collaborate with cross-functional partners
· Creative and resourceful when it comes to problem-solving
· Ability to work with internal stakeholders to collect feedback, prioritize tasks, and manage the engineering backlog
· Self-motivated, self-directed, and the ability to thrive in a fast-paced environment in an industry that constantly changes
Pay Information
GeoZone Definition: GeoZones are geographic groupings created by Lockheed Martin to align compensation ranges with regional labor markets and cost-of-labor differences across the United States. Locations are assigned a Geo Zone based on the primary work location of the role.
* Full-time salary range (GEOZONE 1): $101400.00 - $188200.00
* Includes metropolitan areas such as Sunnyvale CA; Pal Alto, CA; New York City metropolitan area; Newark, New Jersey; etc.
* Full-time salary range (GEOZONE 2): $91200.00 - $169400.00
* Includes metropolitan areas such as Denver, CO; King of Prussia, PA; Stratford, CT; Moorestown, NJ; etc.
* Full-time salary range (GEOZONE 3): $81100.00 - $150500.00
* Includes metropolitan areas such as Dallas–Fort Worth, TX; Orlando, FL; Grand Prairie, TX; Marietta, GA; etc.
* Full-time salary range (GEOZONE 4): $72900.00 - $135500.00
* Includes metropolitan areas such as Camden, AR; Lexington, KY; Lufkin, TX; etc.
At Lockheed Martin, we know mission success starts with taking care of our people. Our Total Rewards program is designed to attract top talent, support your well-being, and help you grow—both professionally and personally.
The salary range for this position is as listed on the requisition. Please note that the salary information listed is a general guideline only.
Lockheed Martin considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, education/ training, key skills as well as market(work location) and business considerations when extending an offer.
Benefits offered: Medical, Dental, Vision, Flexible work arrangements and schedules (e.g., 4x10), 401(k) match, Paid time off, Holidays, Parental Leave, EAP, Flexible Spending Accounts, Education Assistance, Life Insurance, Short-Term Disability, and Long-Term Disability.
* Annual short-term and/or long-term incentive compensation programs may be offered depending on the position. Payments under these annual programs are not guaranteed and can vary from year to year and are tied to a range of performance metrics.
* For (Washington state applicants only) Non-represented full-time employees: accrue at least 10 hours per month of Paid Time Off (PTO) to be used for incidental absences and other reasons; receive at least 90 hours for holidays. Represented full time employees accrue 6.67 hours of Vacation per month; accrue up to 52 hours of sick leave annually; receive at least 96 hours for holidays. PTO, Vacation, sick leave, and holiday hours are prorated based on start date during the calendar year.