About this role
skillset : Java , Observability (ELF , Grafan , Splunk) , Github , Service now(ticketing) , AWS/GCP
Replacement of Harikaran.
Roles & Responsibilities:
• Provide hands-on support for the runtime operation of our applications, ensuring high availability and performance.
• Collaborate with software engineering and infrastructure teams to troubleshoot and resolve runtime issues, including performance bottlenecks, scalability challenges, and system failures.
• Contribute to the design and implementation of monitoring, alerting, and logging solutions to proactively identify and address potential runtime issues.
• Participate in incident response and root cause analysis efforts to ensure the stability and resilience of the applications.
• Work closely with cross-functional teams to understand application requirements and provide input on runtime and operational considerations during the software development lifecycle.
• Contribute to the development and maintenance of runtime automation and tooling to streamline operational processes and improve efficiency.
• Develop common framework components (to be leveraged by enterprise applications), define standards for configuration, monitoring, reliability, and performance engineering
• Create automation and ensure automated tests are completed for new features.
• Good attitude, communication, willingness to learn and collaborate.
• Continuously improve automated remediation tasks to ensure the highest levels of availability.
• Cloud: Manage secure, scalable, and highly available cloud infrastructure.
• Kubernetes & Containers: Deploy, operate, and troubleshoot containerized workloads.
• Observability: Implement monitoring, logging, tracing, dashboards, and actionable alerts.
• Reliability: Define SLOs/SLIs, manage error budgets, and improve service availability.
• Networking: Troubleshoot DNS, TCP/IP, HTTP/S, TLS, routing, and load balancing.
• Linux & Systems: Administer and troubleshoot Linux systems and performance issues.
• Programming/Scripting: Automate operational tasks using Python, Go, or Shell.
• Infrastructure as Code: Provision and manage infrastructure using Terraform or equivalent IaC tools.
• CI/CD: Build and maintain automated, reliable deployment pipelines.
• Incident Management: Respond to incidents, perform RCA, and implement preventive actions.