Senior Site Reliability Engineer

FinastraPune, MaharashtraHybridFull-timeSenior, 5–8 yearsListed 3 hours ago

Apply now

About this role

Who are we?

At Finastra, we’re a global leader in financial services software, dedicated to expanding access to financial services and shaping what’s next for the industry. Our technology powers mission‑critical solutions across Lending, Payments and Universal Banking, supporting over 7,000 customers, including 80% of the world’s top 50 banks, in more than 110 countries.

What Success Looks Like  
The successful candidate will help move the organization further toward an SRE operating model by reducing manual operations, improving observability, and engineering automation for recurring operational activities.

The engineer should be proactive rather than waiting for every task to be assigned. We are looking for someone who identifies repetitive work, reliability risks, monitoring gaps, and opportunities for automation and then takes ownership of developing the solution.

During production incidents, the engineer should actively investigate, communicate findings, recommend actions, and remain engaged through resolution.

The role will primarily focus on SRE, automation, DevOps, monitoring, and reliability engineering, while maintaining enough infrastructure and Windows expertise to contribute to the broader operational responsibilities of the team

Role Summary

We are seeking a hands-on Senior Site Reliability Engineer to improve the reliability, availability, automation, and operational efficiency of business-critical platforms.

The role focuses on Site Reliability Engineering practices, with a strong emphasis on automation, proactive monitoring, incident ownership, disaster recovery, and continuous improvement.

The position supports production infrastructure across Azure, and on-premises environments. The engineer will work across multiple technologies and platforms, including Windows infrastructure, with a focus on reducing manual operational work and improving service reliability.

This is a hands-on individual contributor role.

Key Responsibilities:

Site Reliability Engineering & Automation

- Design and develop automation to reduce repetitive manual operational tasks.

- Develop automation using PowerShell, Python, Ansible, APIs, and other appropriate technologies.

- Integrate infrastructure automation into enterprise CI/CD and DevOps pipelines.

- Implement infrastructure as code using Terraform and related technologies.

- Develop reusable automation rather than one-off scripts.

- Identify reliability risks and opportunities for automation and proactively drive improvements.

- Develop automated health checks, validation, remediation, and reporting.

Disaster Recovery & Resilience

- Design and implement automation for disaster recovery activities across multiple products.

- Automate failover, failback, infrastructure validation, application validation, and customer validation.

- Reduce manual intervention during DR exercises and recovery events.

- Develop reusable DR automation that can be adopted across multiple products.

- Improve system availability, resiliency, and recoverability.

Monitoring & Observability  
Improve monitoring, observability, and alerting using Grafana and other enterprise monitoring platforms.

- Develop and automate monitoring plugins and health checks.

- Automate monitoring configuration and deployment.

- Improve proactive detection of infrastructure and application issues.

- Reduce unnecessary alerts and improve actionable monitoring.

- Develop automated remediation where appropriate.

Infrastructure & Cloud Engineering

- Support production infrastructure across Azure, virtualized, and on-premises environments.

- Automate infrastructure provisioning, configuration, maintenance, and validation.

- Support high availability and resilient infrastructure architectures.

- Automate patching, vulnerability remediation, security hardening, and compliance validation.

- Integrate existing Ansible-based infrastructure and security automation into enterprise DevOps pipelines.
- Collaborate with cloud, network, database, middleware, security, application, and infrastructure teams.

Windows & Production Operations

- Provide hands-on support for Windows Server infrastructure as part of the broader production environment.

- Troubleshoot Windows services, Active Directory, Group Policy, DNS, networking, certificates, authentication, and operating system issues.

- Support Windows patching, security hardening, and vulnerability remediation.

- Develop automation to reduce manual Windows administration.

- Participate in production maintenance and disaster recovery activities

Incident Management & Reliability

- Participate in Sev1/2/3 production incidents and take technical ownership through service restoration.

- Troubleshoot complex infrastructure and application-related issues.

- Perform root cause analysis and implement permanent engineering solutions for recurring problems.

- Remain actively engaged during incidents and communicate technical findings and recommended actions.

- Identify patterns from incidents and convert them into monitoring, automation, or reliability improvements.

Embed AI into project delivery to drive productivity, automate routine activities, strengthen decision quality through trusted data, and deliver continuous process improvement while applying appropriate human judgement

Required Qualifications

Experience : 5+ years of experience in Site Reliability Engineering, DevOps, cloud engineering, infrastructure engineering, or systems engineering.

- Strong hands-on scripting and automation experience using PowerShell, Python, Bash, or similar languages.

- Experience with CI/CD and enterprise DevOps platforms such as Azure DevOps, Jenkins, GitHub Actions, GitLab CI, or equivalent.

- Experience with Ansible or similar configuration-management technologies.

- Experience with Terraform or other infrastructure-as-code technologies.

- Experience supporting cloud infrastructure in Azure, or comparable platforms.

- Experience with monitoring and observability platforms such as Grafana, Prometheus, or equivalent.

- Experience supporting business-critical production environments.

- Understanding of high availability, disaster recovery, failover, and failback.

- Working knowledge of Windows Server administration and troubleshooting.

- Strong production troubleshooting and incident-management skills.

- Ability to independently identify opportunities for improvement and drive engineering solutions through completion.

Preferred Qualifications

- Strong PowerShell and/or Python experience.

- Microsoft Azure infrastructure experience.

- Azure DevOps.

- Azure Site Recovery.

- Kubernetes or AKS.

- Grafana and Prometheus.

- ServiceNow or similar ITSM platforms.

- Experience with vulnerability-management and privileged-access technologies.

- Experience automating patching, security hardening, vulnerability remediation, or compliance controls.

- Experience supporting financial services, payments, or other highly available regulated environments.

- Familiarity with SRE practices including SLI/SLO concepts and reducing operational toil.

We are proud to offer a range of incentives to our employees worldwide. These benefits are available to everyone, regardless of grade, and reflect the values we stand for:

Flexibility: Enjoy unlimited vacation, subject to local regulations and business priorities. Benefit from hybrid working arrangements and inclusive policies such as paid time off for voting, bereavement, and sick leave.

Well‑being: Access confidential one‑to‑one support through our Employee Assistance Program, connect with our network of Wellbeing Champions and Gather Groups, and take part in monthly events and initiatives designed to help you thrive—inside and outside of work.

Health & Financial Security: Medical, life and disability insurance, retirement plans, lifestyle, and other benefits.*

Sustainability: Paid time off for volunteering and donation‑matching opportunities to support causes that matter to you.

Inclusion: Get involved in our inclusion communities, such as Count Me In, Culture@Finastra, Proud@Finastra, Disabilities@Finastra, and Women@Finastra—open to everyone who wants to participate and contribute.

Career Development: Access online learning and accredited courses through our Skills & Career Navigator tool.

Recognition: Take part in our global recognition program, Finastra Celebrates, and share your voice through regular employee surveys that help shape our culture and ways of working.

*Specific benefits may vary by location.

At Finastra, each individual is unique—bringing their own ideas, perspectives, cultural backgrounds, and experiences. We learn from one another, value what makes us different, and create an environment where everyone feels included, supported, and able to be their authentic selves.

Be unique. Be exceptional. Help us make a difference at Finastra.