Senior Observability Engineer

StarHub LtdSingaporeOn-siteFull-timeSenior, 5–8 yearsListed 2 days ago

Apply now

About this role

Job Description

Responsibilities:

- Administer the log and search platform (currently Splunk Cloud): indexes, data onboarding, roles and access, apps and knowledge objects, retention, and search and dashboard performance. Monitor ingest latency, skipped searches and licence consumption, and tune to keep cost under control.

- Administer the APM, RUM and infrastructure monitoring platform (currently Splunk Observability Cloud): teams, tokens, integrations, detectors, alert routing and dashboards. Manage metric cardinality and usage against entitlement.

- Roll out telemetry to all IS applications: OpenTelemetry Collector (currently the Splunk distribution) on AWS EKS and EC2, and APM instrumentation for Java, Go, .NET, Node.js and Python. Keep instrumentation vendor-neutral where possible so the backend can change without re-instrumenting. Maintain onboarding guides and report coverage per application.

- Set up RUM for customer-facing web and mobile frontends and link browser sessions to backend traces, logs and infrastructure metrics, so an incident can be followed from the user to the data store.

- Build and maintain observability as code (Terraform or equivalent) in Git with peer review and CI/CD, so detectors, dashboards and configuration are versioned, repeatable and free of manual drift.

- Plan and run platform and agent upgrades with vendors: define the test plan, validate in non-production, sign off before release, and manage support cases and escalations. Run structured evaluations and proofs of concept when StarHub considers a new or replacement tool.

- Design, build and operate AI and agent-assisted observability workflows, such as alert noise reduction, automated triage, root-cause summaries and natural-language querying, with human review and guardrails.

- Apply controls aligned with MAS-TRM and CSA best practices to telemetry: retention, access control, audit logging, and masking of PII and sensitive data before ingest.

- Perform User Access Reviews (UAR) on the observability platforms at least twice a year, covering user accounts, roles and tokens, and follow up on removals and exceptions. Participate in company-wide audits when required, providing evidence and closing findings on time.

Qualifications

Minimum Profile/ Track Record:

Desired Background

- Experience in medium-to-large technology, telecommunications or financial services organizations with complex hybrid-cloud environments and many application teams; has owned an enterprise observability platform, not only used one.

Seniority, Skills, Certifications (must-haves)

- Bachelor's degree in Computer Science, Information Technology, Engineering or a related field.

- 5+ years of relevant experience in observability, monitoring, SRE or platform engineering, including hands-on administration, optimization and health monitoring of an enterprise observability or log analytics platform. Splunk (Cloud or Enterprise) is strongly preferred.

- Hands-on experience with an enterprise APM/RUM/infrastructure monitoring platform across RUM, APM, infrastructure and data-storage monitoring. Splunk Observability Cloud is preferred; Datadog, Dynatrace, New Relic, Elastic or similar is acceptable. Clear understanding of distributed tracing from frontend to backend.

- Experience installing and operating OpenTelemetry-based telemetry on AWS (EKS, EC2) for Java, Go, .NET, Node.js and Python applications.

- Experience managing observability configuration as code (Terraform or equivalent, Git, CI/CD).

- Prior AI delivery experience in observability or IT operations: has taken an AI or agent-based capability into production use by operations or engineering teams.

- Familiarity with security and compliance standards (MAS-TRM, CSA), including telemetry data governance.

- Certifications are a plus: Splunk Core Certified Admin or Power User, Splunk Observability Cloud certification, or equivalent vendor certifications on another observability platform.

Ideal track record #1

- Owned an enterprise observability or log analytics platform at scale (Splunk preferred) as administrator and platform owner, improving search performance, stability or ingest cost, and ran vendor-led upgrades with test and validation plans and no unplanned outage.

Ideal track record #2

- Rolled out APM and infrastructure telemetry (OpenTelemetry) across a large application estate with RUM-to-backend tracing in use during real incidents, measurably reducing time to detect or resolve.

Ideal track record #3

- Delivered an AI or agent-based capability in production for monitoring, triage or root-cause analysis, with measurable impact such as less alert noise or faster resolution.