Senior HPC Scheduler Operations Engineer

Qualcomm IncorporatedMexicoOn-siteFull-timeSenior, 5–8 yearsListed 6 days ago

Apply now

About this role

## Company:
Qualcomm Intl Inc., Mexico Branch Office

## Job Area:
Information Technology Group, Information Technology Group > IT Engineering

General Summary:

Position Summary

Qualcomm's Grid Solution team is seeking a Senior HPC Scheduler Operations Engineer to   operate , scale, and continuously   optimize   large-scale job scheduling infrastructure supporting Electronic Design Automation (EDA) and compute-intensive engineering workloads. This role is embedded within the   Engineering IT   Hardware Infrastructure — EDA Compute organization and requires deep   expertise   in IBM Spectrum LSF as the primary scheduler, with a working knowledge of   Slurm   as an emerging priority. The ideal candidate brings strong Linux systems administration skills, a reliability engineering mindset, and the ability to translate complex scheduler behavior into actionable insights for engineering and management stakeholders.

Key Responsibilities

Scheduler Operations & Optimization

- Manage, scale, and   optimize   IBM Spectrum   LSF   job scheduling systems for EDA and HPC compute-intensive workloads across multi-site environments.

- Administer scheduler configuration including queue policies, resource limits,   fairshare , job arrays, and preemption strategies.

- Analyze scheduler and infrastructure performance data to   identify   bottlenecks and improve   utilization , throughput, and job turnaround time.

- Perform capacity planning and workload characterization to support SKU-level packing models and tiered job scheduling strategies.

- Implement and   maintain   scheduler tuning parameters to minimize dispatch latency and maximize grid efficiency under high-saturation conditions.

Reliability & Incident Management

- Troubleshoot and resolve service-impacting issues across scheduler, OS, and workload layers with minimal time-to-resolution.

- Define and track SLOs for service performance and reliability; partner with customer teams to set and communicate realistic expectations.

- Implement automation and process improvements to reduce manual   toil   and prevent recurring incidents.

- Develop and enforce operational standards, runbooks, and best practices ensuring consistency across all sites.

Observability, Metrics & Automation

- Build and   maintain   observability systems including   metrics   pipelines, dashboards, and alerting frameworks for scheduler and compute infrastructure health.

- Leverage accounting and telemetry data (e.g., LSF stream logs,   bjobs ,   bacct ) to quantify workload coverage, grid efficiency, and tier performance.

- Develop automation tooling (Python, Shell, APIs) to streamline operations, enforce policy guardrails, and surface actionable insights.

- Produce operational reports and management summaries that combine automated metrics with human operational narratives.

Stakeholder Collaboration

- Collaborate directly with   CAD and   engineering teams to clarify workload requirements, translate technical tradeoffs, and drive issues to closure.

- Communicate scheduler performance metrics, reliability posture, and capacity status clearly to engineering and management audiences.

- Partner with peer infrastructure teams to coordinate cross-functional changes and   align on   shared operational standards.

Required Qualifications
Education  
Bachelor's degree in Computer Science , Electrical Engineering, or related field, or equivalent practical experience.

Experience

5+   years   operating   and supporting large-scale Linux-based   compute   infrastructure in HPC or silicon design environments.

Scheduler Expertise

Strong hands-on experience with IBM Spectrum LSF, including queue configuration, policy tuning,   fairshare , job arrays, and multi-cluster operations.

Linux Systems

Proficiency   in   SLES   administration including OS-level troubleshooting,   workload   profiling   and user environment management.

Problem Solving

Ability to independently analyze complex system behavior   under load ; experience with root cause analysis across scheduler, OS, and hardware layers.

Scripting & Automation

Proficiency   in Python and/or Shell scripting for operational automation, data processing, and   tooling   development.

Communication

Demonstrated ability to clearly articulate technical tradeoffs,   reliability   metrics, and operational status to both engineering and management audiences.

Preferred Qualifications

- Working knowledge of   Slurm   (Simple Linux Utility for Resource Management); familiarity with   Slurm   configuration, partition management, and job scheduling concepts is a strong plus as the team evaluates multi-scheduler environments.

- Deep knowledge of scheduler internals, configuration tuning, and advanced troubleshooting for LSF and/or   Slurm .

- Familiarity with container technologies (Docker, Singularity/ Apptainer ,   Podman ) in HPC or EDA compute environments; experience deploying or managing containerized workloads on HPC schedulers is   advantageous   as the team evolves toward container-native compute delivery.

- Background influencing adoption of new infrastructure standards and operational practices across distributed engineering teams and multi-site organizations.

- Exposure to   hardware-level telemetry and performance monitoring (e.g., PMU counters, IPC, memory bandwidth) for workload analysis.

Minimum Qualifications:
• 3+ years of IT-related work experience with a Bachelor's degree.
OR
5+ years of IT-related work experience without a Bachelor’s degree.

*Completed advanced degrees in a relevant field may be substituted for up to two years (Master’s = one year, Doctorate = two years) of work experience.

Applicants : Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail  [email protected]  or call Qualcomm's toll-free number found  here . Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).

Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.

To all Staffing and Recruiting Agencies :   Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.

If you would like more information about this role, please contact Qualcomm Careers .