← Back to jobs

HPC Engineer

  • Remote
  • Sweden
  • English
  • Posted 08.09.26 18:24

At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.

We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.

Join Verda while it’s still being built - not once it’s finished.

Why Verda

Cash and equity compensation along with various fringe benefits (healthcare, lunch, wellbeing, and more).Profitable operations with rapid, sustained growth.30+ nationalities, with 6 different ones on the management team.A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.

Practicalities

Work mode: Remote (EU)Level: Mid / Senior Employment type: Full time and permanent

About The Role

GPUs only deliver value once they're wired into a cluster that researchers can actually train on. As our HPC Engineer, you'll own the baremetal and virtualized clusters behind our AI cloud - from the InfiniBand fabric and shared filesystems up through the workload orchestration layer. You'll be the person who keeps large-scale GPU clusters healthy, performant, and ready for the next workload.

Your Responsibilities

Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operationsAdminister virtualized clusters, including the hypervisor and GPU virtualization stack underneath themDesign, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validationDeploy and operate shared/parallel filesystems supporting training and inference workloads, balancing performance, capacity, and reliabilityTroubleshoot and resolve issues across the full fabric stack: fibers, transceivers, NICs, switches, drivers, and firmwarePartner with remote-hands and data center teams to diagnose hardware faults and execute physical-layer fixes and cluster expansionsOperate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teamsKeep issue tracking, IPAM, and DCIM records accurate as clusters are built, changed, and decommissionedParticipate in on-call rotations and incident response for cluster-level issuesCollaborate with platform, network, and storage teams to integrate new clusters into the broader AI cloud

Your key competencies

Solid Linux skills, with specialization in memory management, PCIE topologies and virtualization being a bonusDeep Infiniband knowledge, including fabric design, subnet management, and performance tuningSolid experience with IB clustering, troubleshooting/debugging, understanding the ecosystem of fibers+transceivers+NICs+switches and all the things that could possibly fail in themExperience about shared filesystems (e.g. Lustre, GPFS/Spectrum Scale, WekaFS, or similar)Ability to work with remote hands teams to diagnose and resolve hardware issues remotelyKnowledge about NCCL, CUDA, DOCA and the Nvidia stackKnowledge about Slurm and/or other workload scheduling solutions, bonus points for Slinky/slurm-bridgeUnderstanding the importance of keeping issue tracking/IPAM/DCIM up to dateComfort operating production clusters where uptime and performance directly affect customer workloadsScripting/automation ability (e.g. Python, Bash, Ansible) for repeatable cluster operations

Nice to have

RoCEv2 knowledge (and/or Spectrum-X)Understanding of agentic guardrails, especially when applied to administration of complex systemsAbility to think beyond what is needed right now vs. some given trajectory or roadmapExperience with GPU health-checking and diagnostics tooling (e.g. DCGM, field diagnostics)Experience with baremetal provisioning/orchestration tooling (e.g. MAAS, Foreman, custom PXE/iPXE pipelines)Familiarity with GPU-aware virtualization or containerization (Kubernetes device plugins, KVM/QEMU with GPU passthrough, SR-IOV)Exposure to observability stacks (Prometheus, Grafana, Loki) for cluster-level monitoring

What's Next

We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now.

Please submit your application through our Careers page. We don't accept applications sent by email.