← Takaisin työpaikkoihin

Weekend Site Reliability Engineer

  • Etätyö
  • Ruotsi
  • Englanti
  • Julkaistu 22.09.26 18:49

What You’ll Be Doing

Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SREImprove existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the futureContinuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practicesMonitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviewsDesign and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification floodingDefine and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisationTake ownership and responsibility for our cloud operation activitiesLiaise with external security agencies for annual audits as well as perform our own internal security sweepsAid in reconfiguring existing architecture to allow for rapid deployments to new countriesMentoring less experienced team members

What You’ll Bring

3+ years DevOps / platform engineering experienceMust be based in Europe or Asia or LatAMExperience independently leading the planning and deployment of a projectExperienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and productionStrong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valuedExperience with Infrastructure-as-Code, particularly TerraformProficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plusHands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetryExperience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plusProven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actionsAbility to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patternsExperience defining SLIs and SLOs and using them to inform reliability workFamiliarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environmentsSolid networking knowledge, especially the TCP / IP stack and HTTP protocolExperience handling high HTTP request volumes and designing systems for high availability and high traffic environmentsA strong understanding of cache, including CDN, HTTP cache, Redis / MemcachedExcellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous

Our stack

Languages: Java / Spring Boot, Node.js, Python, JavaScriptDatabase: Aurora MySQL & PostgreSQL, MongoDB, MySQL CommunityCache: ElastiCache, Redis, ValkeyMessaging: Apache RocketMQ, AutoMQ, KafkaNetworking & Proxy: Nginx, Kong, Cilium, eBPFOrchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, HelmComputing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3CI/CD: Jenkins, GitHub ActionsMetrics: Prometheus, Mimir, Grafana, AlertmanagerLogs: Loki, VectorTraces: Tempo, OpenTelemetry, AlloyProfiling: PyroscopeRUM: Grafana Faro, OpenTelemetry SDKInfrastructure as Code: TerraformCDN & Edge: Cloudflare, AWS CloudFrontAWS CloudWatch

What’s In It For You

Sporty is a remote first company in pursuit of sustainabilityA competitive salary + individual performance based bonuses every quarter28 days paid annual leaveOur core working hours are 10am-3pm in your local time zone with flexibility outside of thisReferral bonuses & flash bonusesTop of the line equipmentAnnual company retreats to provide great internal networking opportunities

Interview process

Remote video screening with our Talent Acquisition Team Online assessment via HackerrankRemote video interview with 3 x Team Members (45 mins each, not separate days)

If you're interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.