About this role
Site Reliability Engineer - ED&A
The Site Reliability Engineer – EDA is a technical contributor responsible for Site Reliability Engineering practices supporting the Enterprise Data & Analytics platform. This role performs reliability and operational engineering for data and analytics platforms, integrations, pipelines, and related services, including platforms such as Google BigQuery, Microsoft Fabric, Power BI, Power Platform, and associated Azure and Google Cloud services. The Site Reliability Engineer establishes service-level indicators, service-level objectives, error budgets, monitoring, alerting, dashboards, and runbooks while driving incident response, root-cause analysis, automation, capacity planning, performance tuning, resilience, and production readiness. This role provides technical guidance, engineering standards, and operational best practices for engineers supporting EDA services, and partners closely with data engineering, application, cloud, security, and governance teams to improve reliability, supportability, and business outcomes.
JOB DUTIES
- Performs reliability and operational engineering for the Enterprise Data & Analytics platform, including Google BigQuery, Microsoft Fabric, Power BI, Power Platform, associated data pipelines, integrations, automation, reporting services, and dependent cloud services.
- Establishes and governs service-level indicators, service-level objectives, error budgets, availability targets, performance baselines, monitoring standards, alerting practices, dashboards, runbooks, and operational health metrics for EDA services.
- Analyzes telemetry from monitoring, logging, tracing, platform administration, pipeline execution, query performance, capacity, consumption, and cost-management tools to identify reliability, performance, security, scalability, and efficiency improvements.
- Partners with data engineering, application, cloud, security, governance, analytics, and infrastructure teams to improve platform design, integration patterns, deployment practices, release readiness, supportability, resilience, and production operations.
- Drives automation for provisioning, deployment, remediation, monitoring configuration, environment validation, job and pipeline health checks, alert enrichment, access reviews, incident response workflows, and operational reporting.
- Guides capacity planning, performance tuning, resilience engineering, disaster recovery planning, backup and restore validation, service continuity planning, architecture reviews, and production readiness assessments for EDA services.
- Troubleshoots and resolves incidents involving data and analytics platforms, workloads, integrations, pipelines, APIs, automation flows, connectors, workspaces, gateways, permissions, queries, semantic models, and platform dependencies.
- Drives incident response, root-cause analysis, post-incident reviews, corrective actions, and reliability improvement plans to reduce recurrence, improve operational maturity, and strengthen customer experience.
- Provides technical guidance, engineering standards, implementation patterns, peer support, operational reviews, documentation practices, and reliability expectations for engineers supporting the EDA platform.
- Promotes secure, compliant, cost-effective, and well-governed data and analytics operations by supporting access controls, data protection practices, platform governance, resource utilization reviews, lifecycle management, and operational reporting.
EDUCATION & EXPERIENCE
Typically requires a bachelor's degree and five (5) to seven (7) years of experience in a technology and/or software engineering role or an equivalent combination.
KNOWLEDGE, SKILLS, ABILITIES
- Advanced understanding of SRE principles, including reliability engineering, observability, automation, incident response, root-cause analysis, post-incident improvement, service-level indicators, service-level objectives, error budgets, and production readiness.
- Experience performing reliability, operations, or engineering support for enterprise data and analytics platforms used for reporting, data engineering, integration, automation, business intelligence, and business productivity workloads.
- Experience with Google Cloud data services such as BigQuery, Cloud Storage, Cloud Logging, Cloud Monitoring, IAM, networking concepts, and workload or job performance troubleshooting.
- Experience with Microsoft Azure services and operational capabilities, including Azure Monitor, Log Analytics, Azure networking, identity and access management, resource management, and cloud-native administration.
- Experience with Microsoft Fabric, Power BI, Power Platform, Power Automate, Power Apps, gateways, connectors, workspaces, data pipelines, semantic models, and platform administration concepts.
- Ability to monitor, troubleshoot, and tune platform performance, capacity, reliability, availability, jobs, queries, pipelines, APIs, integrations, automation flows, and dependent services.
- Knowledge of identity, access, security, data governance, compliance, data protection, backup, recovery, and change-management practices for enterprise data and analytics platforms.
- Experience creating and governing operational dashboards, alerting standards, runbooks, production support documentation, platform standards, reliability scorecards, and continuous improvement plans.
- Experience with automation, scripting, infrastructure as code, configuration management, and CI/CD tooling such as Azure DevOps, Terraform, PowerShell, Python, or similar tools.
- Experience with observability and monitoring platforms such as Azure Monitor, Google Cloud Monitoring, Grafana, Datadog, Dynatrace, or similar tools.
- Ability to provide technical guidance across infrastructure, security, data, application, cloud, governance, and business technology teams to improve reliability, supportability, standards adoption, operational maturity, and customer experience.
- Strong troubleshooting skills across cloud services, network dependencies, APIs, databases, operating systems, authentication, authorization, and enterprise integration patterns.
- A strong mix of software engineering, systems engineering, data platform operations, automation, production support, technical guidance, and cross-functional collaboration skills.
PHYSICAL DEMANDS:
LICENSES & CERTIFICATIONS:
SUPERVISORY RESPONSIBILITY:
BUDGET RESPONSIBILITY: No
COMPANY INFORMATION: Motion offers an excellent benefits package which includes options for healthcare coverage, 401(k), tuition reimbursement, vacation, sick, and holiday pay.
DISCLAIMER: This job description illustrates the general nature and level of work performed by employees within this job classification. It is not intended to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and skills required. Management retains the right to add or modify duties at any time.
Not the right fit? Let us know you're interested in a future opportunity by joining our Talent Community on jobs.genpt.com or create an account to set up email alerts as new job postings become available that meet your interest!
GPC conducts its business without regard to sex, race, creed, color, religion, marital status, national origin, citizenship status, age, pregnancy, sexual orientation, gender identity or expression, genetic information, disability, military status, status as a veteran, or any other protected characteristic. GPC's policy is to recruit, hire, train, promote, assign, transfer and terminate employees based on their own ability, achievement, experience and conduct and other legitimate business reasons.