HPC Systems Engineer

Remote · Full Time

About the role

We’re hiring an HPC Systems Engineer to own the reliability, performance, and day-to-day operability of high-performance computing environments. This role is for someone who’s comfortable deep in Linux, understands how researchers and engineers actually use clusters, and can translate that into stable scheduling, predictable throughput, and clean automation.

You’ll focus on building and operating Slurm-based clusters end-to-end: provisioning and configuration, user onboarding, queue and partition design, fair-share policies, and troubleshooting jobs from “why is my node down?” to “why is this MPI run stalling?”. You’ll also help standardize how environments are delivered—modules, containers, images, and repeatable configuration—so the cluster stays maintainable as demand grows.

What success looks like

Clusters are available, observable, and predictable: users can submit jobs with confidence, scheduling behavior matches policy, and incidents are resolved quickly with clear root cause and follow-up improvements. You’ll reduce toil through automation, keep upgrades and changes low-risk, and ensure capacity is used efficiently without sacrificing fairness.

How you’ll work

You’ll collaborate closely with infrastructure and platform stakeholders, partnering with researchers, data/ML teams, and application owners to understand workload patterns and remove bottlenecks. Expect hands-on work, pragmatic trade-offs, and a strong bias toward documentation and operational clarity.

  • Slurm operations: partitions, QoS, accounting, fair-share, and troubleshooting
  • Linux cluster engineering: provisioning, configuration management, patching, and hardening
  • Performance & reliability: monitoring, capacity planning, and incident response

What you'll do

Success in this role means researchers and engineers get reliable, fast access to compute—without having to think about the cluster.

  1. Keep Slurm-backed compute capacity available, predictable, and right-sized for demand.
  2. Reduce time-to-resolution for incidents through clear runbooks, metrics, and automation.
  3. Improve job throughput and user experience by tuning scheduling, storage, and network performance.
  4. Ship secure, repeatable cluster changes with minimal downtime.

From day one

Get hands-on with existing Linux clusters, Slurm configuration, and operational tooling. Triage current pain points, validate monitoring and alerting, and establish a clear change process for maintenance windows, upgrades, and user-impacting work.

What you’ll own

  • Operate and evolve Slurm (partitions/QoS, fairshare, job accounting, reservations, preemption) to balance throughput, priority, and cost.
  • Administer Linux cluster nodes: provisioning, patching, kernel/driver management, and lifecycle maintenance across heterogeneous hardware.
  • Automate infrastructure and configuration management (e.g., Ansible) and standardize golden images and node bring-up/replace workflows.
  • Troubleshoot performance and reliability issues across compute, storage, and networking; drive root-cause analysis and corrective actions.
  • Own monitoring, logging, and alerting (Prometheus/Grafana, syslog, Slurm metrics); define SLOs and on-call playbooks.
  • Manage shared filesystems and data paths (e.g., Lustre/GPFS/NFS) and tune for parallel I/O workloads.
  • Implement security hardening: least privilege, SSH/key management, audit trails, vulnerability remediation, and segmentation.
  • Collaborate with research teams to translate workload needs into capacity plans, queue policies, and documentation.
  • Plan and execute upgrades (Slurm, OS, drivers, firmware) with tested rollbacks and clear stakeholder communication.

What we're looking for

Requirements

  • 3–7 years experience administering Linux systems in production, including multi-node environments.
  • Hands-on experience operating and supporting Slurm in an HPC or research/engineering compute setting.
  • Comfort working in a ticketed/on-call environment and owning incidents through root-cause analysis and follow-up actions.

Technical

  • Strong Linux fundamentals: systemd, networking, storage, kernel/user limits, performance troubleshooting, and security hardening.
  • Cluster operations: provisioning, configuration management, patching, lifecycle management, and documentation for repeatable builds.
  • Slurm administration: partitions/QoS, accounting, fairshare, job priorities, reservations, node states, and troubleshooting scheduling issues.
  • Scripting and automation using Bash and/or Python; ability to build reliable tooling for day-to-day operations.
  • Monitoring and observability: deploying/maintaining metrics, logging, and alerting to support capacity planning and uptime.

Experience

  • Supporting heterogeneous workloads (MPI, GPU, and high-throughput batch) and working with users to optimize job submissions.
  • Collaborating with researchers/engineers to translate workload needs into scheduler and cluster configuration changes.
  • Clear written documentation for runbooks, change records, and post-incident reviews.

Bonus Points

  • Experience with InfiniBand/RDMA, parallel filesystems (e.g., Lustre/BeeGFS), and performance tuning for low-latency interconnects.
  • Experience with containers in HPC (e.g., Apptainer/Singularity) and integrating container workflows with Slurm.
  • Infrastructure-as-code and CI practices (e.g., Ansible, Terraform, Git-based change management).

Apply for HPC Systems Engineer

Upload your CV and links. A person reviews every application and we'll get back to you either way.

* Required

Frequently Asked Questions

01

What are the working hours and expectations for this job?

Since this is a remote role, the working hours are flexible. The work model in some cases could be hourly-based, and in other cases project-based.
02

How does the hiring process work?

1. Initial Screening of your application and CV 2. An introductory interview to assess fundamental qualifications and key criteria. 3. Technical interview/s covering DevOps, Platform Engineering, and SRE principles, as well as proficiency and problem-solving skills. 4. CEO interview assessing both advanced technical skills and personal character.
03

What types of projects will I be working on?

We work on a wide variety of projects across multiple industries and company sizes, each involving different technologies. Our clients are located in different countries worldwide, providing opportunities to work on international projects. Projects are matched based on the candidate's expertise, ensuring the best fit for both the project requirements and your skillset.
04

How is the payment structured?

We offer hourly contracts and you get paid for the hours worked every month.
05

Is there potential for long-term collaboration?

Absolutely! Most of our engineers stay with us long-term. Even after a project concludes, we usually have new opportunities available, allowing us to continue our collaboration.