Senior Site Reliability Engineer
Infrastructure Site Reliability Engineer – Linux | Network Reliability | Python / Go | Remote EU [€140K Total Compensation] [AI Cloud Infrastructure] Our client is building a full-stack AI cloud platform, supporting developers and enterprises from model training through to production deployment. Built by engineers, for engineers. From GPU orchestration at scale to inference optimisation, they own the hard problems across compute, storage, networking and applied AI. Publicly listed, headquartered in Europe, with R&D hubs across Europe, the UK, North America and Israel, and a team of 1,500+ including hundreds of engineers. We are seeking a Network SRE to help build and run the part everything else depends on for our client: the network. This is an engineering-first role. You define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where they make sense), and you drive improvements across the whole network, not only services but also site readiness, inter-site connectivity (DCI) and operational standards. You own incident response, lead investigations and postmortems, and turn failures into durable fixes instead of repeated firefighting. A large part of this job is building. You evolve observability into actionable metrics, logs, traces and alerting that shorten the debug loop during and after incidents. You design safer change workflows: automation, CI/CD, staging environments, canarying, rollbacks and auditability for network changes, working closely with network engineers to embed operability into designs. You will report to a team lead in Europe. Our client is adding nine SREs to the infrastructure organisation this year, seven in Europe and two in the US. We expect you to haveStrong production Linux fundamentals and a structured approach to debugging complex systemsSolid understanding of networking and how real networks fail: control plane versus data plane, latency and loss, failure domainsHands-on experience operating high-availability systems and improving them over time, not just keeping the lights onAbility to write and maintain software and automation in Python or GoExperience with modern infrastructure tooling such as IaC, CI/CD and container platforms, and comfort automating operational workflows Bonus points for high-throughput datapath work (load balancers, tunnelling/decap, NAT64), low-level networking performance debugging (eBPF/XDP, DPDK, perf/ftrace, kernel networking), or large-scale routing and flow telemetry. If you like networks where the failure modes are real and every packet has consequences, please apply. Location: Europe - 100% Remote.Total compensation: up to €140.000,- depending on experience Interested?Contact me at s.vanderiet@doghouse.nl
