HPC Platform Engineer

RemoteFull Time

About the role

We鈥檙e hiring an HPC Platform Engineer to own the on-prem GPU and HPC platform as an integrated whole. Where our Network, Storage, and Systems engineers go deep in their domains, this role owns the layers that hold the platform together鈥攑rovisioning, GPU orchestration, and the cross-domain integration work that turns a stack of components into a dependable service.

You鈥檒l work hands-on across bare-metal lifecycle (PXE/iPXE, MAAS, BMC/Redfish), Kubernetes with the NVIDIA GPU Operator (MIG/MPS, device plugins, topology-aware scheduling), the GPU stack itself (drivers, CUDA, NCCL, container runtimes), and the integration glue that makes scheduling, networking, storage, and compute behave predictably under real workloads.

What you鈥檒l focus on

  • Provisioning & lifecycle: bare-metal automation, OS imaging, firmware/driver management, and predictable bring-up
  • GPU orchestration: Kubernetes with the NVIDIA GPU Operator, MIG/MPS, and workload scheduling
  • Cross-domain integration: stitching together scheduling, networking, storage, and compute into a coherent platform
  • Operational excellence: upgrades, capacity planning, runbooks, and incident response with measurable improvements

You鈥檒l partner closely with the Network, Storage, and Systems engineers鈥攁ligning on architecture, escalating cross-domain issues, and making sure each layer鈥檚 behavior contributes to a platform users can trust. Success looks like fast, boring bring-ups, calm upgrades, and a platform that holds together under real GPU workloads鈥攖raining, fine-tuning, and inference.

What you'll do

Success in this role means the on-prem HPC and GPU platform is delivered as a coherent, dependable service鈥攁cross provisioning, scheduling, networking, storage, and GPU operations鈥攔ather than a stack of disconnected components.

  1. Keep the platform available and predictable for compute-intensive workloads at scale.
  2. Reduce operational toil through automation, repeatable processes, and clear standards.
  3. Shorten time-to-resolution for incidents with strong observability and disciplined root-cause analysis.
  4. Ship changes safely with tested procedures, validation, and clean rollback plans.

From day one

Get hands-on with the existing environment: review provisioning, scheduler configuration, GPU stack versions, networking, and storage. Validate observability and incident history, then prioritize the highest-impact reliability and automation work.

What you鈥檒l own

  • Bare-metal provisioning and lifecycle: PXE/iPXE, MAAS or equivalent, golden images, and BMC/Redfish-based automation across heterogeneous hardware.
  • Linux cluster operations: OS, kernel, drivers, systemd, security hardening, and configuration management at scale.
  • Scheduling and orchestration: Slurm (partitions/QoS, fairshare, accounting) and/or Kubernetes (GPU Operator, device plugins, MIG/MPS, topology-aware scheduling).
  • GPU stack health: firmware, CUDA, NVIDIA Container Toolkit, DCGM telemetry, NCCL validation, and known-good state across the fleet.
  • High-performance networking awareness: InfiniBand/RoCE behavior, MTU/PFC/ECN/QoS impact on workloads, and partnering with network engineers on fabric design.
  • Storage integration: parallel filesystems (Lustre/GPFS/BeeGFS) and shared filesystems tuned for HPC I/O patterns.
  • Observability: metrics, logs, alerts, and SLOs tied to availability, utilization, job throughput, and time-to-provision.
  • Upgrades and migrations (Kubernetes, Slurm, OS, drivers, firmware) with tested rollbacks and minimal user impact.
  • Documentation: architecture decisions, runbooks, change records, and post-incident reviews with concrete follow-ups.

What we're looking for

Requirements

  • 5+ years operating production Linux infrastructure at scale, including hands-on HPC, GPU, or performance-sensitive environments.
  • Demonstrated breadth: comfortable working across provisioning, scheduling, networking, storage, and GPU operations rather than a single silo.
  • Strong Linux fundamentals: systemd, networking, storage, kernel/driver troubleshooting, and performance debugging.
  • Comfortable with on-call rotations, change windows, and disciplined incident response.

Technical

  • Bare-metal provisioning and lifecycle automation: PXE/iPXE, MAAS or similar, image build pipelines, BMC/IPMI/Redfish, firmware/driver management.
  • Working experience with at least one HPC scheduler (e.g., Slurm) and/or Kubernetes with GPU workloads (NVIDIA GPU Operator, device plugins, MIG/MPS).
  • GPU operations: drivers, CUDA, NVIDIA Container Toolkit, DCGM-based observability, and NCCL validation/troubleshooting.
  • High-performance networking awareness: InfiniBand and/or RoCE fundamentals and how fabric behavior affects real workloads.
  • Parallel storage exposure (e.g., Lustre, GPFS, BeeGFS) and how I/O patterns interact with compute performance.
  • Infrastructure-as-code and configuration management (Ansible, Terraform) plus scripting/automation in Python and/or Go.

Experience

  • Operating GPU or HPC clusters supporting real workloads (training, fine-tuning, inference, MPI/OpenMP, scientific computing).
  • Driving upgrades and migrations safely, with measurable outcomes and clear stakeholder communication.
  • Building automation that makes the platform easier to operate over time鈥攏ot just one-off scripts.
  • Fluent English for documentation, change planning, and cross-team coordination.

Bonus Points

  • Experience designing or operating multi-tenant GPU platforms or GPU-as-a-Service environments.
  • Familiarity with hybrid Slurm + Kubernetes patterns and converged HPC/AI workflows.
  • Low-level diagnostics across NUMA, IRQ affinity, PCIe topology, and NVIDIA tools (nvidia-smi, nvbandwidth, DCGM, NCCL tests).
  • Contributions to open-source HPC, Kubernetes, or GPU tooling.

Apply for HPC Platform Engineer

Upload your CV and links. A person reviews every application and we'll get back to you either way.

* Required

Frequently Asked Questions

01

What are the working hours and expectations for this job?

Since this is a remote role, the working hours are flexible. The work model in some cases could be hourly-based, and in other cases project-based.
02

How does the hiring process work?

1. Initial Screening of your application and CV 2. An introductory interview to assess fundamental qualifications and key criteria. 3. Technical interview/s covering DevOps, Platform Engineering, and SRE principles, as well as proficiency and problem-solving skills. 4. CEO interview assessing both advanced technical skills and personal character.
03

What types of projects will I be working on?

We work on a wide variety of projects across multiple industries and company sizes, each involving different technologies. Our clients are located in different countries worldwide, providing opportunities to work on international projects. Projects are matched based on the candidate's expertise, ensuring the best fit for both the project requirements and your skillset.
04

How is the payment structured?

We offer hourly contracts and you get paid for the hours worked every month.
05

Is there potential for long-term collaboration?

Absolutely! Most of our engineers stay with us long-term. Even after a project concludes, we usually have new opportunities available, allowing us to continue our collaboration.