HPC / Slurm Infrastructure Engineer
HPC / Slurm Infrastructure EngineerContract | Remote | Part-time or Full-time | Rate DOE I'm working with a technology company looking to improve and scale its engineering compute / HPC environment, with a particular focus on designing and deploying Slurm. This isn't a requirement for someone who has simply administered an existing Slurm environment or submitted workloads to a cluster. They need someone who has been involved in the architecture, build and deployment of Slurm-based compute infrastructure and can take ownership of the environment from an early stage. The requirement is still being shaped, so we're looking for someone senior enough to assess the existing infrastructure, recommend the appropriate architecture and then help lead the implementation. Likely areas of responsibility include:Architecture and deployment of Slurm-based HPC / compute clustersSlurm controller and compute-node architectureScheduling, queues/partitions and resource managementQoS, Fairshare and workload prioritisationSlurm accounting and monitoringLinux compute infrastructureCluster provisioning and configurationInfrastructure automation and scriptingPerformance, utilisation and capacity optimisationWorking with engineering teams to understand compute requirements and translate these into infrastructure Experience within semiconductor / ASIC engineering environments would be particularly interesting. Ideally, you'll have supported compute-intensive EDA workloads, such as:RTL simulation and verificationRegression workloadsSynthesisPhysical design / implementationCadence, Synopsys or Siemens/Mentor toolchains The key requirement, however, is genuine Slurm architecture and deployment experience rather than semiconductor experience alone. If you've previously designed and built Slurm clusters rather than simply used or supported them, I'd be interested in speaking. Please message me directly for more information.
