← Takaisin työpaikkoihin

Senior Site Reliability Engineer

  • Etätyö
  • Ruotsi
  • Englanti
  • Julkaistu 24.09.26 21:21

Staff Site Reliability Engineer – Bare Metal Linux | Kubernetes | Python / Go | Remote EU [€140K Total Compensation] [AI Cloud Infrastructure] Our client is building a cloud platform for high-throughput, compute-heavy AI workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not “handled”. We are seeking a Senior or Staff SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You will build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise. This is a bare-metal environment: think Linux, data centres, physical fleets and real hardware constraints, not managed services. You will work close to the metal across Kubernetes internals (scheduling, autoscaling behaviour, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/IO contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it. You will sit in our client’s Hardware Infrastructure team and report to a team lead in Europe. They are adding nine SREs to the infrastructure organisation this year, seven in Europe and two in the US. Must requirementsExtensive Production Engineering experience running bare metal / on-prem / data centre infrastructure (not public cloud only)Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO, kernel-level behaviours)Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting under load)Strong Kubernetes experience beyond manifests: scheduler behaviour, autoscaling edge cases, kubelet pressure/evictions, etcd/control planeExperience with Terraform, Docker, Helm and modern CI/CD practicesStrong coding skills in Python and/or Go, beyond automation scripting. Real engineering capability is a mustExperience in low latency environments If you are looking for complexity and a new place to nerd out on infrastructure optimisation, we would love to hear from you. Location: Europe - 100% Remote.Total compensation: up to €140.000,- depending on experience Interested?Contact me at s.vanderiet@doghouse.nl