← Takaisin työpaikkoihin

GPU Platform Engineer

  • Etätyö
  • Ruotsi
  • Englanti
  • Julkaistu 02.09.26 13:06

GPU Platform EngineerRemote — European UnionAI Infrastructure | GPU Computing | Distributed Systems | ML Platforms Our client, a growing AI Infrastructure company based in Paris, is looking for a GPU Platform Engineer to build and optimise the shared compute platform supporting large-scale model training and high-volume inference. You'll work deep in the infrastructure layer, combining Kubernetes, NVIDIA GPUs, distributed computing, and platform automation to help AI teams get maximum performance from expensive compute resources. What You'll Work On• Build and operate GPU-enabled Kubernetes infrastructure• Support distributed model training and large-scale inference workloads• Optimise GPU scheduling, allocation, utilisation, and performance• Build distributed compute environments using Ray• Develop platform tooling and automation in Python• Provision infrastructure using Terraform• Improve observability across GPU, Kubernetes, and workload layers• Troubleshoot performance bottlenecks across compute, memory, networking, and storage• Build self-service capabilities for ML and AI engineering teams• Improve platform reliability, scalability, and compute efficiency Core Skills• 4+ years in ML Infrastructure, Platform Engineering, MLOps, HPC, or similar roles• Kubernetes• NVIDIA GPU infrastructure• CUDA ecosystem• Python• Terraform• Ray or comparable distributed compute technologies• Monitoring and observability• Strong understanding of Linux and distributed systems Nice to HaveNVIDIA GPU OperatorNCCL and distributed GPU communicationPyTorch distributed trainingTriton Inference Server / vLLMPrometheus / Grafana / OpenTelemetryAWS, GCP, or Azure GPU infrastructureSlurm or HPC environmentsGPU capacity planning and cost optimisationLarge-scale LLM training or inference