Senior Site Reliability Engineer

Public Sector Resourcing (PSR)On-siteFull-timeSenior, 5–8 yearsListed 16 hours ago

Apply now

About this role

On behalf of DWP, we are looking for a Senior Site Reliability Engineer for a 5 Months (Inside IR35) contract based Hybrid/ in Leeds/Manchester/Sheffield/Newcastle/Blackpool .

The Department for Work and Pensions (DWP) is responsible for welfare, pensions, and child maintenance policy. As the UK’s biggest public service department, it administers the State Pension and a range of working age, disability and ill health benefits to around 20 million claimants and customers. As such, we operate on a scale that is almost unmatched anywhere in Europe and most people in Britain come into contact with us at some point in their lives.

Working with DWP, you will be helping us to drive our priorities to:

- Help people to move into work and support those already in work to progress, with the aim of increasing overall workforce participation
- Help people to plan and save for later life, while providing a safety net for those who need it now
- Provide effective, efficient, and innovative services to the millions of claimants who rely on us every day, including the most vulnerable in society
- Improve experience of our services while maximising value for money for the taxpayer.

As a Senior Site Reliability Engineer your main responsibilities will be:

- As a Senior Site Reliability Engineer, you will play a pivotal role in ensuring the reliability and performance of our applications and infrastructure.
- The SRE team will put you in the position to work with application teams across the department on developing reliable and secure solutions to provide to citizens across the UK.
- You will lead by example, providing technical direction and supporting other SREs within your team.
- You will work with development teams from the design phase to help them use good practice and department standards when building their application infrastructure.
- Design and develop the techniques for improving application reliability, run books, knowledge transfer across teams, and ongoing SRE strategy within your Functional and Professional Communities.
- Work collaboratively with development teams and provide guidance around best practice and ensure monitoring of applications is enabled.
- Push a mindset change within the organisation to foster engineering ownership , SRE best practice and the importance of the integrity and maintenance of the Live Service.
- Manage the error budget agreed with the product owner for the application and ensure that work is balanced in alignment with it.
- Act as the focal point for the investigation and resolution of major or complex incidents for the service, ensuring people with the right skills and expertise are proactively available to respond effectively.
- Assess the impact of change requests in consultation with stakeholders, providing technical expertise and authorising the implementation of subsequent changes.
- Coach and mentor application development and operations engineers in the practice and techniques of SRE.
- Conduct reviews for all high priority and major incidents ensuring they are done quickly and published.
- Routinely seek views and capture ideas from stakeholders and team members for improvements and encourage collaboration and innovation.
- Provide on-call support to help restore services, through dedicated run books or technical experience.
- Help to reduce toil and increase automation; by developing reliability to ensure we have a reduction of the time to live, and cost spend on repetitive tasks.

PLEASE NOTE: SC Clearance is an essential requirement for this role, as a minimum you must be willing & eligible to undergo checks. Please note, due to the exceptional requirements of this position (short-term nature of this role and speed at which we require a postholder in situ) preference may be given to candidates who meet all of the essential criteria and hold active security clearance.

Essential:

- Demonstrable experience using automation to remove toil with scripting, infrastructure, and configuration as code.
- Demonstrable experience of reliability engineering including capacity and performance management through monitoring, logging, and alerting.
- Demonstrable experience of supporting a Live Service.
- Demonstrable experience of developing and supporting cloud-based applications in AWS.
- Demonstrable experience of building and maintaining CI/CD pipelines.
- Demonstrable experience communicating effectively with stakeholders at multiple levels to provide feedback and support.

Please be aware that this role can only be worked within the UK and not Overseas.

Disability Confident

As a Disability Confident employer, DWP encourage applications from disabled and neurodivergent candidates . Applications from candidates who identify as having a disability and meet the essential criteria for the role will be prioritised for review. Where application volumes are high, additional role-specific and desirable criteria may be applied as part of the shortlisting process which may include holding active security clearance.

Armed Forces Covenant

As a signatory of the Armed Forces Covenant, DWP welcome applications from veterans, service leavers, reservists and military spouses or partners. Applications from eligible candidates who meet the essential criteria for the role will be prioritised for review. Where application volumes are high, additional role-specific and desirable criteria may be applied as part of the shortlisting process which may include holding active security clearance.

In applying for this role, you acknowledge the following "this role falls in scope of the Off Payroll Working in the Public Sector legislation. Any rates of payment quoted will reflect the gross rate per day for the assignment and will be subject to appropriate taxes and statutory costs. As such the payment to the intermediary and your income resulting from this contract will be different".