Observability Engineer
About the Company
We are looking for a talented and driven Site Reliability Engineer (Observability Engineer) to join our team and play a key role in building the next generation of our high-performance observability framework.
About the Role
In this role, you will be at the heart of designing, developing, and operating the systems that give our engineering organisation deep visibility into the health, performance, and reliability of our products and infrastructure. You will work closely with engineering teams across the business to ensure our observability stack is scalable, reliable, and fit for purpose in a fast-moving, high-stakes environment.
Responsibilities
Designing and developing a next-generation, high-performance observability platform using modern tooling and best practices.Driving the adoption and implementation of OpenTelemetry (OTEL/OTLP) for metrics and log ingestion across the organisation.Building and maintaining Infrastructure as Code (IaC) and configuration management pipelines to automate platform operations.Contributing to CI/CD pipelines and deployment workflows to support continuous delivery of observability components.Owning and improving the reliability, scalability, and performance of the observability stack.Collaborating with engineering teams to define and track Service Level Indicators (SLIs) and establish reliability standards.Participating in code reviews, architectural discussions, and technical planning sessions.Supporting capacity planning, release processes, and change management practices.Championing an automation-first mindset and contributing to a culture of operational excellence.
Qualifications
Proven experience with Infrastructure as Code, specifically Terraform.Configuration management experience using Ansible and/or Chef.
Required Skills
Strong scripting and development capabilities in Bash, Python, and Go.Experience with Rust and/or Java is a nice to have.Solid understanding and hands-on experience with Git and GitLab.Experience with ArgoCD and/or Jenkins for CI/CD pipeline management.Metrics platforms: Prometheus, Mimir, InfluxDB, VictoriaMetrics.Log aggregation: Splunk, Loki, Elasticsearch, VictoriaLogs.Visualisation: Grafana.Log forwarding: FluentBit, Vector.dev.OpenTelemetry (OTEL/OTLP): Metrics and log pipeline experience.Strong experience with Docker and/or Podman.Kubernetes (including workload management and cluster operations).Helm for Kubernetes package management.Hands-on experience with AWS.
Preferred Skills
Systems Architecture: ability to design and reason about complex distributed systems.Networking principles: TCP/IP, DNS, load balancing, proxies, and beyond.TSDB (Time Series Database) principles: understanding of storage, retention, cardinality, and query patterns.Logging principles: structured logging, log pipelines, and log management at scale.Messaging queues: pub/sub patterns, event-driven architectures.Clustering principles: high availability, fault tolerance, distributed coordination.Load balancing: strategies, implementations, and traffic management.API / JSON principles: RESTful APIs, data serialisation, and integration patterns.Caching mechanisms: In-memory caching, external caches, and cache invalidation strategies.Code reviews: Contributing to and leading thoughtful, constructive reviews.Testing and documentation: Writing quality tests and clear technical documentation.Agile mindset: Comfortable working in iterative, collaborative delivery environments.Automation mindset: Defaulting to automation and repeatability in everything you build.
Pay range and compensation package
Expected compensation details are not provided in the original description.
Equal Opportunity Statement
We are committed to diversity and inclusivity in our hiring practices.
