Systems Administrator

HPC Systems Administrator

Company
Santa Clara University
Posted
Posted 5 days ago
Email me new Systems Administrator jobs in Santa Clara, CAFree account. One email a night at most, only when something new lands.

The job

Position Title

HPC Systems Administrator

Position Type

Regular

Hiring Range

$129,000.00 - $161,265.00 annually; Compensation will be based on education, experience, skills relevant to the role, and internal equity.

Pay Frequency

Annual

A. POSITION PURPOSE

The High-Performance Computing (HPC) System Administrator is a hands-on role responsible for the day-to-day configuration, monitoring, optimization, and operation of the organization's high-performance computing infrastructure.

This individual will focus on system optimization, troubleshooting, and general planning for future infrastructure enhancements across compute, storage, and high-speed interconnects (InfiniBand). A key responsibility is to collaborate with existing system administrators, building the team’s collective HPC expertise, strengthening shared support capabilities, and ensuring long-term operational resilience and efficiency.

The HPC Systems Administrator is a member of the Enterprise Systems team within the Cyberinfrastructure Technolog i es department. The incumbent works with the other Cyberinfrastructure teams - Network and Telecommunications, Enterprise Applications, and the Information Security Office - and other campus divisions in coordinating services, providing support and providing appropriate guidance. This incumbent will also work with University vendors and partners.

The HPC Systems Administrator will have a passion for providing excellent customer serv i ce, and a focus on continual improvement; a commitment to supporting innovative infrastructure technologies; and a desire to learn and grow their technical expertise to meet the needs of the campus community.

B. ESSENTIAL DUTIES AND RESPONSIBILITIES

1. HPC Infrastructure Management and Optimization

  • Compute: Supports the lifecycle of all compute nodes, including installing, configuring, and maintaining hardware, operating systems, and core system software to ensure optimal performance, stability, and resource utilization for scientific workloads.
  • Storage: Assists in the management of the high-performance parallel file systems (e.g., Lustre, GPFS, CephFS), NAS, and backup solutions, executing capacity planning, performance tuning, and integrity checks to guarantee secure, high-speed, and reliable data access for all users.
  • InfiniBand: Designs, deploys, and provides troubleshooting and maintenance for the InfiniBand high-speed interconnect fabric, ensuring low-latency, high-bandwidth inter-node communication essential for scalable HPC application performance.

2. Workload Management and System Deployment

  • Slurm: Supports day-to-day administration and configuration of the Slurm Workload Manager, monitoring job queues, partitions, and resource allocation policies to enforce fair-share scheduling, maximize cluster utilization, and meet diverse research computational needs.
  • System Imaging: Helps develop, maintain, and update standardized, optimized system images for all compute nodes, utilizing automation tools to facilitate rapid, consistent deployment, efficient patching, and streamlined upgrades across the cluster environment.
  • Software Licenses: Supports administration of all commercial scientific software licenses, helping track compliance with vendor agreements and managing license servers and usage policies to optimize utilization and accessibility for the HPC user base.

3. Professional Development and Knowledge Building

  • Actively participates in cross-training and knowledge-sharing sessions for existing system administrators to enhance the team's collective expertise in HPC-specific technologies (Slurm, InfiniBand, parallel file systems).
  • Documents routine procedures and troubleshooting steps to support team knowledge sharing and consistency.
  • Stays current on emerging HPC technologies and relevant industry trends, evaluating vendor solutions, and providing recommendations with the team.

4. Coordination and Collaboration

  • Participates and plays an active role in the planning and implementation phases of new technologies, contributing ideas in architecture brainstorming and design discussions with the team.
  • Provides guidance on infrastructure questions, escalating complex architecture decisions to senior team members.
  • Supports team members and contributes to a collaborative, solutions-oriented team culture.
  • Models and supports openness and honesty to enhance positive relationships based on trust, predictability, and communication.

5. Resource Planning

  • Provides input on Enterprise Systems and CIT goals, objectives, and strategies based on the University's mission, goals and strategic plan.
  • Provides input in technology planning processes to develop cost-effective customer-focused solutions.
  • Uses strong technical and organizational knowledge to plan and support projects and working groups.

6. Service Delivery

  • Work closely with the ES Manager in the planning, maintenance, and secure expansion of SCU's computing infrastructure.  This includes, but is not limited to, local and hosted servers, virtual appliances and devices, and storage.
  • Helps ensure that architecture principles and standards are consistently applied across the data center compute and storage services.
  • Collaborates with the Information Security Office (ISO) to support a secure and compliant enterprise environment.
  • Supports planning for future security needs and threats, with the guidance of the ISO and senior staff.
  • Ensures the appropriate distribution of infrastructure services to faculty, staff, and students .
  • Follows and helps document standards and practices regarding data center, compute and storage services for use across the University.
  • Supports the creation and performance of infrastructure production and test environments.
  • Assists in implementing scalable, interoperable, and flexible infrastructure solutions.
  • Supports assigned systems with on-call availability and respond within agreed upon timeframes.
  • Follows existing, and helps establish new processes  for applying patches/updates to operat i ng systems, applications, and hardware and firmware to ensure all physical, virtual, and hosted systems are appropriately secure and current .
  • Participates as necessary in backup operations , ensuring all required file systems and system data are successfully backed up to the appropriate media and are available off site.
  • Participates in disaster recovery and business continuity planning.
  • Performs daily system monitoring, verifying i ntegrity and availability of all hardware, server resources, systems, and key processes . Checks for potential problems, resource availability, capacity, performance and load characteristics, network integrity, and security threats. Monitors systems activity and usage to maintain a secure environment. Develops related solutions as warranted .
  • Collaborates with the CISO and system stakeholders to establish upgrade and update schedules, and maintenance windows .
  • Keeps abreast of software releases and updates, keeping all systems at current release levels as appropriate for the successful operation of the data center in support of the University.
  • Assists senior staff in monitoring service level agreements with hosted platform and third party providers.

7. Service Optimization

  • Assists senior staff in reviewing existing architecture frameworks and identifying opportunities for simplification and improvement .
  • Assists in the design, planning, and implementation of infrastructure systems optimization and process improvement projects.
  • Tests and assesses existing infrastructure against industry standard benchmarks to ensure optimal performance and service delivery.
  • Participates in IT and information security audits and helps implement corrective actions .

Posted by Santa Clara University. Some listings are shortened, so open the full posting for the rest.

Applying is the easy part. Keeping track is not.

JobQuill is the job tracker that keeps you going. Every application in one pipeline, a nudge when one goes quiet, and daily quests and streaks that count every step, rejections included.

  • Save jobs from this board to your queue
  • Unlimited applications on the free plan
  • One free AI resume review
  • Free forever plan, no card needed