About this role
About the Role
We are looking for a Cloud Operational Engineer (L3) to become the most senior technical authority in our Cloud Operations team and the final escalation point for our client estate.
This is a hands-on role with real authority. You will hold the access you need to resolve work end to end, in Entra ID, Azure, Intune and across our clients' SaaS estate, rather than diagnosing a problem and passing it to somebody else. If you have spent time in roles where you knew exactly what the fix was but could not action it, this is the opposite of that.
You will spend your time on the work that genuinely needs senior judgement: complex incidents, identity and access operations, platform troubleshooting, and root cause analysis on the issues that keep coming back. Alongside that, you will turn what you know into runbooks and standards the wider team can use, so that expertise lives in the system rather than in one person's head.
This is an operations role, not a project role. You will not be pulled onto client implementations or infrastructure builds. Our Engineering (Professional Services) team owns design, build and delivery. You own how the platform runs once it is live, and you own the standards it runs to.
What You Will Do
Monitoring and operational response
• Triage and resolve platform alerts across the monitoring estate, covering availability, performance, capacity and configuration drift
• Judge quickly which alerts represent genuine service impact and which do not, and act accordingly
• Turn recurring alert patterns into problem records and automation candidates rather than working them repeatedly
• Own operational response for managed file transfer and scheduled job failures, including rerun, validation and client communication
Identity and access operations
• Entra ID: user and guest lifecycle, group membership, cross-tenant collaboration, attribute and licensing changes
• Azure RBAC and PIM: role assignment within an approved scope catalog, and time-bound privileged access within a defined role set
• App registrations and service principals: creation against standard patterns, secret and certificate lifecycle, API permissions and admin consent
• Credential hygiene: proactive rotation and renewal of secrets and certificates ahead of expiry, handled as scheduled work rather than as incidents
• Conditional Access and MFA: investigate authentication failures, analyze policy behavior and impact, and support security investigations
Platform operations
• Microsoft 365: Exchange Online mail flow, shared mailboxes and distribution lists, quarantine and anti-spam handling, SharePoint and OneDrive access models, Teams policy and governance
• DNS and domains: record changes across managed zones, new domain onboarding, and mail authentication records including SPF, DKIM and DMARC
• Data platform access: access and permission operations across Microsoft Fabric, Databricks, Snowflake and Azure SQL
• Federation and SSO: configure and operate enterprise applications, SAML federation and SCIM provisioning for third-party SaaS
• Endpoint and virtual desktop: Intune configuration and application deployment across Windows, macOS, iOS and Android, plus Azure Virtual Desktop and Cloud PC session and profile troubleshooting
• Compute and infrastructure: availability, performance, access and capacity issues across existing virtual machine estates
Incidents, problems and root cause
• Act as technical lead during major incidents, including client-facing technical communication under pressure
• Own root cause analysis on significant and recurring incidents, and drive preventive actions through to closure
• Identify systemic risk and recurring failure patterns, and convert them into problem records with named owners
• Escalate to Engineering (Professional Services) with a clear handoff: symptom, diagnostics performed, hypothesis and specific ask
Standards, runbooks and knowledge
• Author and maintain runbooks for recurring work, to the standard that any task performed twice becomes a documented procedure
• Own operational standards, troubleshooting guidance and escalation criteria
• Review and validate the runbooks used by our Tier 1 and Tier 2 engineers, and identify work that can safely move down a tier
• Maintain documentation in Confluence as the system of record, validated against the live environment
• Record your own work clearly, so that what was requested, what was done and what was verified is captured before closure
Automation and continuous improvement
• Identify and eliminate repetitive operational work through scripting and automation, with a bias toward removing work rather than doing it faster
• Build and maintain PowerShell automation that is idempotent, commented, re-runnable, and safe to test before it runs
• Propose self-service and workflow automation for high-volume, low-complexity requests
• Recommend monitoring and alerting improvements, supported by evidence
Multi-client service delivery
• Operate consistently across multiple client tenants while maintaining strict tenant isolation
• Work within delegated administration models, privileged access workflows and break-glass procedures, and understand the audit implications of each
• Apply standards uniformly across the client estate, and flag where an environment has drifted from standard
• Handle client data with the care our sector requires, and support our SOC 2, ISO 27001 and UK GDPR obligations
Where This Role Starts and Stops
We are specific about this because unclear boundaries are what make senior operations roles frustrating.
You own
Operating, troubleshooting, configuring and maintaining platforms that are already live. Access and entitlement work within an approved scope. Alert response and runbook execution. Root cause analysis and problem management. The standards and documentation the rest of the team works to. Automation that removes repetitive work.
You do not own
Client implementations, onboardings and migrations. Infrastructure built through Terraform or other infrastructure-as-code pipelines. Network and firewall architecture. Monitoring design. First-time integration architecture for a new platform. Those sit with Engineering (Professional Services), and you will work closely with them without being pulled into their delivery schedule.
A simple test: if the work needs a repository commit, a pull request or a design decision, it belongs to Engineering (Professional Services). If it operates something that already exists, it belongs to you.
Requirements
What We Are Looking For
Experience
• 6 to 10+ years in cloud, identity or workplace engineering, including demonstrable time operating at L3 or subject matter expert level
• Proven experience as a final escalation point in a production environment with real service commitments
• Degree or equivalent practical experience in IT or a related field
Technical skills
• Expert-level Microsoft 365 administration, including Exchange Online mail flow, SharePoint, OneDrive and Teams governance
• Deep, hands-on Entra ID expertise: authentication and authorization, Conditional Access, MFA, SSO, PIM, app registrations, service principals and Microsoft Graph permissions
• Strong Azure operational capability: RBAC, subscription and resource group management, virtual machines, storage, Key Vault and Log Analytics
• Expert Microsoft Intune and endpoint management across Windows, macOS, iOS and Android
• Practical experience with Azure Virtual Desktop and Windows 365 Cloud PC, including profile and session troubleshooting
• Identity integration experience: SAML configuration, SCIM provisioning, and claims and attribute mapping
• DNS administration and mail authentication records
• PowerShell automation to a production standard, including Microsoft Graph PowerShell
• Identity security and Zero Trust principles, with least-privilege applied by default
• Advanced troubleshooting and structured root cause analysis
Service management and tooling
Hands-on ServiceNow experience is required. You should be comfortable with incident, request and catalog task fulfilment, correct categorization and assignment, work note and closure discipline, SLA awareness, and how catalog items and request workflows drive fulfilment. We also want someone who reads ticket data to spot recurring problems, not only to close individual tickets.
• ITIL-aligned service management understanding across incident, problem, change and knowledge management
• Confluence or an equivalent knowledge platform, as an active author rather than an occasional reader
• Monitoring and alerting platform experience, such as Zabbix or Azure Monitor
• Azure DevOps familiarity, sufficient to manage access, repositories and pipelines at an entitlement level
Managed service provider experience
Prior experience in an MSP, managed service or multi-tenant environment is required. Operating across many client tenants is fundamental to this role rather than incidental to it.
• Comfort working across multiple client tenants at once, with the discipline to keep them properly separated
• Understanding of delegated administration, privileged and break-glass access models, and their audit implications
• Awareness of the commercial side of managed services: SLAs, contracted scope, and why accurate time and ticket records matter
• Client-facing communication skill: able to explain technical detail credibly to a non-technical stakeholder, and to stay calm in a difficult conversation
How you work
• You prefer to close work rather than route it
• You reach for the standard, repeatable fix rather than the clever one-off
• You write things down, and treat documentation as part of the job rather than overhead
• You stay structured under pressure during major incidents
• You challenge assumptions and push for clarity without making it personal
• You enjoy raising the capability of the engineers around you
Certifications
Preferred
• Microsoft 365 Certified: Enterprise Administrator Expert (MS-102)
• Microsoft Certified: Identity and Access Administrator Associate (SC-300)
• Microsoft Certified: Azure Administrator Associate (AZ-104)
• Microsoft Certified: Endpoint Administrator Associate (MD-102)
Advantageous
• Microsoft Certified: Azure Security Engineer Associate (AZ-500)
• Microsoft Certified: Security Operations Analyst Associate (SC-200)
• ITIL 4 Foundation
• Experience in a regulated sector such as insurance, financial services or healthcare
Benefits
- Free lunch meal, fruits, snacks, and drinks
- Onsite gym with a free professional instructor
- Weekly fitness activity and an annual fitness challenge where you can win up to 70,000 PHP
- Weekly engagement activities with prizes that are up to 3,000 PHP
- Free upskilling academy to improve your performance and skillset
- State-of-the-art facilities from toilets to your workstation
- Amenities such as sleeping quarters, game area, chat room, shower room