← Back to jobs

Senior Site Reliability Engineer

  • Remote
  • Sweden
  • English
  • Posted 04.09.26 21:51

About Distillery Distillery is a software development partner that helps organizations design, build, and scale modern technology solutions. Leading companies work with us to accelerate digital initiatives, solve complex technical challenges, and bring innovative products to market faster. Working as an extension of our clients' teams, we deliver lasting business value.Distillery is committed to diversity and inclusion. We embrace a dynamic and inclusive culture where you have the opportunity to continuously learn and grow. About the RoleDistillery is seeking a Senior Site Reliability Engineer to improve the reliability of its production platform. The platform supports an open-source AI runtime and a product website, so website availability is a core product requirement. The engineer will own reliability improvements across infrastructure, observability, incident response, and production change controls, working with application engineers to diagnose failures across the browser, API, runtime, and cloud layers.The environment includes AWS, Amazon EKS, Terraform, GPU infrastructure, and a C++ runtime, plus Python and TypeScript extensions and a JavaScript frontend. Current ops stack uses PagerDuty; a Grafana-based observability approach is under consideration.We're looking for someone who can stabilize an existing system before introducing larger platform changes, and who works well in an early-stage environment with incomplete documentation and changing priorities. ResponsibilitiesOperate and improve production workloads on AWS and Amazon EKS.Maintain infrastructure as code with Terraform.Design actionable monitoring for infrastructure and user-facing services.Build useful dashboards and alerts from service-level indicators.Implement synthetic, browser-level, and API-level health checks.Participate in a shared on-call rotation across distributed time zones.Lead incident triage, mitigation, communication, and follow-up.Write clear incident reviews focused on system and process improvements.Create runbooks, escalation paths, and service ownership records.Improve deployment safeguards, rollback procedures, and configuration validation.Automate repetitive operational work with tested scripts and tools.Diagnose Linux, container, Kubernetes, network, DNS, TLS, and HTTP failures.Partner with developers on application monitoring and pre-production testing.Test recovery procedures and identify resilience gaps.Support secure access, secrets handling, and audit-ready operational practices.Communicate risks, trade-offs, and reliability recommendations clearly.Required Experience5–7 years of direct SRE work is ideal.A mix of SRE, DevOps, and platform engineering experience is also valuable, including 3+ years of direct SRE work.Related production, infrastructure, systems, or software engineering experience is valuable when it includes reliability ownership.4+ years operating AWS services in production.3+ years operating Kubernetes in production, including 2+ years with Amazon EKS.3+ years using Terraform, including modules, state, reviews, and safe changes.3+ years working with metrics, logs, traces, dashboards, and actionable alerts.2+ years defining and using service-level indicators and service-level objectives.3+ years responding to production incidents and leading root-cause reviews.5+ years troubleshooting Linux systems and production networks.3+ years using CI/CD controls, automated validation, and safe rollback methods.3+ years writing reliable automation in Python, Go, Bash, or a similar language (strong in one, working knowledge of another).Ability to work across infrastructure and application boundaries.Clear written and verbal communication in a remote environment.A record of ownership from problem detection through verified resolution.Preferred Experience2+ years with Grafana or a similar observability platform.1+ year improving on-call operations with PagerDuty or a similar platform, including alert noise and escalation paths.1+ year implementing synthetic monitoring or real-user monitoring.1+ year supporting production C++ services or other high-performance runtimes.2+ years supporting Python applications and TypeScript or JavaScript applications.1+ year supporting GPU infrastructure or AI inference workloads.1+ year supporting OCR, speech, audio, or vision workloads.1+ year supporting identity systems such as Zitadel or another OIDC provider.1+ year supporting SOC 2 or ISO-aligned operational controls, or one complete audit cycle.2+ years in an early-stage company or another fast-changing product environment.Success MeasuresMonitoring detects user-visible failures before customers report them.Alerts are actionable and have a clear owner and escalation path.Incident response is documented, consistent, and effective.Repeat incidents decrease because follow-up actions are completed.Production changes receive appropriate automated and pre-production validation.Recovery procedures are documented and tested.Service reliability can be measured against agreed objectives.Application and infrastructure teams share clear operational responsibilities.