Principal Solutions Architect
Job Title: Principal Solutions Architect, Post-Sales (EMEA) – AI/ML Infrastructure Role:A pioneering Kubernetes-native platform provider, enabling enterprises and neoclouds to orchestrate large-scale GPU infrastructure for AI and ML workloads, is hiring a Principal Solutions Architect to anchor its EMEA post-sales function. This is a senior, customer-facing position sitting at the intersection of distributed systems, GPU infrastructure and MLOps, working with some of the most demanding compute environments in the region. The successful candidate will operate with genuine autonomy, acting as the primary technical authority for strategic accounts across EMEA. Direct exposure to cutting-edge GPU fabric and large-scale distributed training environments, working hands-on with the infrastructure underpinning frontier AI workloads. Join at a moment of explosive AI compute demand, with genuine investment and scale behind the business, and a seat that feeds directly into the product and engineering roadmap. A small, senior Solutions Architecture function built on trust rather than micromanagement, with real autonomy, direct access to product and engineering leadership, and a mentorship remit over junior team members. Responsibilities: Design end-to-end AI/ML platform architectures spanning inference, training and data pipelinesDevelop reference architectures for GPU cluster deployment, LLM serving and multi-tenant ML infrastructureAdvise on GPU fabric topology (NVLink, InfiniBand, RoCEv2) for distributed training workloadsAct as the primary technical advisor and escalation point, leading root cause analysis on complex production issuesDeliver technical workshops, proofs of concept and executive-level presentations to senior stakeholdersDesign observability strategies using DCGM, OpenTelemetry, eBPF and GPU metrics pipelinesPartner with customer platform, MLOps and data science stakeholders to translate requirements into architectureFeed customer insights back into the product and engineering roadmap, and mentor junior Solutions Architects Skills/Must have: Experience: 8+ years in infrastructure, platform or solutions engineering, including 3+ years specifically in AI/ML infrastructure or MLOpsCore Tech/Domain: Deep hands-on Kubernetes expertise (cluster lifecycle, workloads, operators, RBAC) and direct NVIDIA GPU infrastructure experience (H100/H200/B200 preferred)Methodology/Protocols: Distributed training knowledge (NCCL, tensor/pipeline parallelism), LLM inference serving and optimisation (vLLM, NIM, TGI), and GPU networking (GPU Operator, MIG, SR-IOV, IB/RoCEv2)Soft Skills: Strong customer-facing and advisory communication skills, credible at both engineer and executive level Nice to haves: Run:AI or Slurm scheduling experiencePyTorch/TensorFlow familiarityPublic cloud experience (AWS/Azure/GCP) including networking, IAM and managed KubernetesCKA, CKAD, AWS Solutions Architect, Azure Solutions Architect or GCP Professional Cloud Architect certificationMulti-tenant GPU isolation knowledge (SR-IOV VFs, DPU offload) Benefits: Performance-related bonusPrivate healthcarePension contributionTraining and certification allowanceRemote-first, flexible working across EMEA Salary:£150,000 to £180,000 base plus performance-related bonus
