Senior Platform Engineer (Observability)

OneMain FinancialEvansville, IndianaOn-siteFull-timeSenior, 5–8 yearsListed 2 days ago

Apply now

About this role

We’re seeking a Senior Monitoring Engineer  to join a high‑performing Monitoring Engineering team in a fast‑paced finance technology organization. You’ll design, develop, and maintain monitoring and observability solutions that keep core applications and infrastructure healthy and visible. In close partnership with application, platform, and development teams, you will implement alerting systems, dashboards, correlations, and automation—driving reliability, reducing MTTR, and elevating operational awareness.

Critical thinking, system analysis, and proactive troubleshooting are essential to success in this role.

Key Responsibilities

Design, Build, and Maintain Monitoring & Observability Solutions

- Develop and maintain instrumentation, telemetry, and alerting  for the Enterprise Monitoring Center using industry‑leading tools, such as: Grafana
- OpsRamp
- AppDynamics
- Elastic Stack
- BigPanda
- AWS CloudWatch
- Azure Monitor

- Implement Observability best practices , ensuring comprehensive coverage of metrics, logs, and traces  across critical systems.
- Integrate and manage OpenTelemetry  for distributed tracing and telemetry data collection, enabling end‑to‑end visibility of business‑critical transactions.

Collaboration & Project Participation

- Collaborate with application development teams to define and document observability requirements  for each project or release.
- Participate in complex initiatives, ensuring accurate and actionable monitoring and tracing are in place for every step of business‑critical workflows.

Alerting & Escalation Process

- Define and maintain standardized alert payloads  per engineering guidelines, ensuring alerts are actionable .
- Partner with Level 2 and Level 3 support teams to reflect process changes in monitoring dashboards.
- Maintain and optimize thresholds , ensuring seamless escalations  via BigPanda  as the central alert hub.

Dashboard Creation & Maintenance

- Create and maintain intuitive, actionable dashboards  for the Enterprise Monitoring Center and other finance teams.
- Ensure dashboards are effectively monitored by Level 1 teams , presenting clear, actionable data that reduces MTTR .

System Validation, Documentation & Automation

- Develop and maintain automation scripts  to enhance monitoring efficiency and improve team quality of life.
- Proactively identify process improvements  and learning opportunities; drive continuous improvement .

Automation & Quality‑of‑Life Improvements

- Contribute to the automation of monitoring, alerting, and operational tasks  to streamline workflows and improve overall system reliability.

Qualifications

Education

Bachelor’s in Computer Science, IT, or related field.

Experience

- Minimum 4 years  in a technology organization, with ≥1 year  hands‑on engineering experience  in monitoring or production operations .

Required Skills

- Strong experience developing instrumentation and alerting  for large, complex  environments.
- Expertise in ≥4  of the following: OpsRamp, Grafana, AppDynamics, Elastic Stack, InfluxDB, BigPanda , and other monitoring solutions.
- Hands-on experience with Observability concepts and frameworks , including metrics, logs, and traces .
- Working knowledge of OpenTelemetry  for distributed tracing and telemetry data collection.
- Experience with dashboard creation , alert management , and tool configuration .
- Excellent verbal and written communication —able to present complex technical issues to both technical and non‑technical stakeholders.
- Strong problem‑solving and troubleshooting  in high‑pressure  environments.
- Ability to prioritize and manage multiple tasks  in a deadline‑driven  setting.
- Proven collaboration with cross‑functional teams  in large, complex IT environments .
- Experience with scripting  (e.g., Bash , PowerShell ) and proficiency in one programming language  (e.g., Python , C family , JavaScript ).
- Experience designing and implementing scalable, reliable  monitoring solutions.
- Experience with agile software development methodologies
- Familiar with problem diagnosis; performance tuning; capacity planning and configuration management across the stack via continuous improvement.

Preferred Qualifications

- Experience querying, manipulating, and visualizing time‑series  data.
- Familiarity with Infrastructure as Code  tools (e.g., Ansible , Terraform ).
- Strong understanding of how to create actionable, digestible visualizations  for Level 1 monitoring  teams.
- Working knowledge of REST APIs , JSON , and ServiceNow .
- Experience with cloud monitoring —particularly AWS  or Azure .

OneMain Holdings, Inc. is an Equal Employment Opportunity (EEO) employer. Qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship status, color, creed, culture, disability, ethnicity, gender, gender identity or expression, genetic information or history, marital status, military status, national origin, nationality, pregnancy, race, religion, sex, sexual orientation, socioeconomic status, transgender or on any other basis protected by law.