About this role
Job Description
We are looking for an experienced Platform Engineer to design, build, and implement a centralized observability platform across a large enterprise technology environment.
This is a 6-month renewable contract.
The role will focus on improving visibility across critical applications, APIs, integration services, and infrastructure using Grafana, Prometheus, Grafana Loki, and OpenTelemetry .
The successful candidate will work across Microsoft Azure and OpenShift/Kubernetes environments to establish centralized logging, metrics, tracing, dashboards, and automated alerting.
Key responsibilities include:
- Design, deploy, and maintain a centralized observability platform using Grafana, Prometheus, and Grafana Loki .
- Implement and manage OpenTelemetry collectors for logs, metrics, and distributed tracing.
- Build telemetry pipelines for applications, APIs, infrastructure, integration services, and security-related logs.
- Configure operational and management dashboards to monitor system health, performance, availability, and incidents.
- Set up automated alerts for application errors, API latency, service timeouts, infrastructure issues, and abnormal system behaviour.
- Define and maintain standards for structured logging, error codes, trace IDs, correlation IDs, and application telemetry.
- Ensure internally developed and third-party applications comply with agreed observability and logging standards.
- Develop troubleshooting guides and operational runbooks for support teams.
- Support L1 and L2 teams in using dashboards, logs, alerts, and traces to diagnose incidents.
- Collaborate with software engineering, infrastructure, security, QA, and operations teams.
- Integrate synthetic monitoring and application health checks into the wider monitoring platform.
- Help improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) across critical technology services.
Requirements
- 3 - 5+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering, Cloud Engineering, or a similar role.
- Strong hands-on production experience with Grafana, Prometheus, and Grafana Loki .
- Strong experience with OpenTelemetry , including collectors, instrumentation, distributed tracing, and trace propagation.
- Experience designing and operating centralized logging, monitoring, metrics, and alerting platforms .
- Strong experience with Kubernetes and/or OpenShift .
- Hands-on experience with Microsoft Azure infrastructure and services.
- Experience working in hybrid cloud and on-premise environments .
- Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, GitHub Actions, Azure DevOps, or similar .
- Good understanding of application and infrastructure logging, log parsing, and structured logging.
- Experience monitoring APIs, microservices, and enterprise applications.
- Ability to troubleshoot complex application and infrastructure issues using logs, metrics, and traces.
- Experience defining technical standards and ensuring engineering teams follow them.
- Strong communication skills and the ability to work with engineering, infrastructure, security, and support teams.
Experience with Java/Spring Boot, Node.js, enterprise integration platforms, WAF/security logging, or regulated enterprise environments would be an added advantage.