← Tillbaka till jobb

Site Reliability Engineer

  • Distans
  • Sverige
  • Engelska
  • Publicerad 25.09.26 12:37

OverviewWe are seeking a Site Reliability Engineer specialising in AI systems and platforms to help shape and deliver our internal AI strategy.Our SRE team operates primarily as a platform engineering function. The team manages the infrastructure, security, automation, and shared platforms that engineering and business teams rely on. We operate with a strong sense of responsibility and ownership, ensuring we are deeply accountable for the systems we build and maintain.In this role, you will build and manage secure, scalable, and cost-effective platforms for internal AI adoption. You will help bring organisational data into governed central environments, enable teams to develop and use AI tools safely, and establish common technical standards that prevent duplicated effort and uncontrolled technology sprawl.You will work across infrastructure, security, data, engineering, and business teams. Your goal will be to make approved AI capabilities easy to adopt while ensuring that company data, models, integrations, and infrastructure remain secure, observable, maintainable, and financially sustainable.This is a hands-on engineering role. It does not require you to be an AI researcher or data scientist, but you should understand how modern AI systems are deployed, integrated, secured, monitored, and operated at scale. The tools here change fast - we value the ability to learn quickly and think through unfamiliar problems over a fixed skill list.Roles and ResponsibilitiesAI Platform EngineeringDesign, build, and manage shared platforms for developing, deploying, and operating internal AI tools and services.Establish reusable platform capabilities for model access, prompt and agent execution, retrieval-augmented generation, data integration, evaluation, observability, and access control.Provide secure interfaces to commercial and open-source AI models through approved gateways, APIs, and development frameworks.Build self-service capabilities that allow teams to adopt AI without creating isolated infrastructure or duplicating common components.Evaluate when AI workloads should use managed cloud services, commercial APIs, self-hosted models, or hybrid infrastructure.Integrate AI platforms with existing cloud, Kubernetes, identity, security, monitoring, and developer tooling.AI Security and GovernanceDefine and implement security controls for internal AI systems, including identity, authorization, secrets management, encryption, network isolation, logging, and auditability.Protect sensitive company and customer data from unauthorised model training, retrieval, disclosure, or external transmission.Establish controls around model providers, plugins, agents, tools, data sources, and third-party AI services.Work with security, legal, privacy, and compliance stakeholders to turn AI policies into enforceable technical controls.Implement guardrails for data classification, prompt injection, insecure tool use, excessive permissions, data leakage, and other AI-specific risks.Maintain inventories of approved models, tools, data sources, integrations, and AI services.Support security reviews, threat modelling, risk assessments, and incident investigations relating to AI systems.Own administration of enterprise AI platforms (e.g. Claude Code, Claude apps, Gemini, and similar tools), including provisioning, SSO/identity integration, license and seat management, and policy configuration.Data EnablementHelp bring approved organisational data into governed, central environments where it can be securely used by AI applications.Design secure patterns for connecting AI tools to databases, document stores, APIs, vector databases, data lakes, and internal knowledge platforms.Implement access controls that preserve source-system permissions and prevent AI applications from exposing data to unauthorised users.Work with data owners to define data quality, ownership, retention, lineage, and lifecycle requirements.Build repeatable ingestion and integration patterns instead of one-off data pipelines.Ensure that AI-generated outputs can be traced to their models, prompts, source data, and execution context where required.Standardisation and Internal EnablementDefine and maintain organisation-wide standards for AI architecture, development, security, deployment, and operations.Create reusable templates, libraries, reference architectures, APIs, and infrastructure modules for internal AI projects.Identify overlapping AI initiatives and consolidate common requirements into shared services.Establish approved technology patterns that reduce fragmentation and make solutions easier to secure and support.Partner with engineering and business teams to assess AI use cases and guide them toward suitable shared platforms.Produce clear technical documentation and help teams adopt established standards.Build internal communities of practice and promote responsible, practical, and measurable AI adoption.Infrastructure, Automation, and ReliabilityManage the cloud, Kubernetes, compute, storage, networking, and supporting infrastructure required by AI platforms.Automate platform provisioning, configuration, security controls, deployment, upgrades, and lifecycle management.Apply infrastructure-as-code and GitOps practices to ensure environments are repeatable, reviewable, and auditable.Design platforms for appropriate availability, scalability, resilience, performance, and recoverability.Implement monitoring and observability across infrastructure, models, agents, data pipelines, APIs, and user-facing AI services.Troubleshoot complex problems across infrastructure, applications, integrations, data flows, and model services.Participate in architecture reviews and platform design decisions.Cost and Resource OptimisationMeasure and manage expenditure across AI APIs, cloud services, compute infrastructure, storage, data processing, and software licensing.Implement usage metering, budgets, quotas, rate limits, chargeback or showback, and cost alerts.Compare model and platform options based on quality, security, latency, portability, and total cost.Optimise GPU, CPU, storage, token, and API consumption without compromising required outcomes.Identify unused, duplicated, or unnecessarily expensive tools and recommend consolidation.Provide stakeholders with transparent reporting on AI adoption, utilisation, value, and cost.Technology EvaluationEvaluate emerging AI infrastructure, model platforms, agent frameworks, security tooling, and data technologies.Run structured proofs of concept and assess technologies against defined business, security, operational, and financial requirements.Separate practical enterprise capabilities from short-lived industry trends.Make clear build-versus-buy recommendations and support platform roadmap decisions.Continuously develop your knowledge of AI systems, platform engineering, cloud infrastructure, and security.RequirementsFive or more years of experience in platform engineering, site reliability engineering, systems engineering, cloud engineering, DevOps, infrastructure engineering, or a related discipline.Strong foundation in identity and access management, role-based access control, secrets management, encryption, network controls, and audit logging - with the judgment to enforce them and push back on stakeholders when risk exceeds acceptable limits.Strong and practical experience with AWS as the major cloud provider.Ability to evaluate and learn new tools, platforms and technologies quickly - the specific stack will keep changing.Strong Linux systems engineering and troubleshooting skills.Production experience deploying and managing Kubernetes and containerised workloads.Experience using infrastructure-as-code and configuration-management tools such as Terraform and Ansible.Strong automation and scripting skills using Python, Go, Bash, or a similar language.Experience building or operating shared internal platforms used by multiple engineering or business teams.Experience with CI/CD, GitOps, source control, and automated platform delivery.Experience with monitoring, logging, metrics, tracing, and observability platforms such as Prometheus, Thanos, Grafana, Loki, or the Elastic Stack.Good understanding of APIs, distributed systems, networking, storage, databases, and data integration.Ability to evaluate technologies and make decisions based on security, maintainability, performance, scalability, and cost.Strong communication skills and the ability to work across infrastructure, security, engineering, data, and business teams, ability to reason based on security, maintainability, performance, scalability and cost..A practical approach to standardisation: able to create common patterns without blocking responsible experimentation.Practical experience using AI as part of development workflows (e.g. coding agents), applied with a high degree of professional responsibility and accountability — owning the code produced, verifying output quality, and identifying and resolving errors.AI Systems ExperienceCandidates don't need deep AI/ML expertise, but should bring curiosity and a working understanding of the following:What large language models and generative AI systems are, how they're deployed and integrated — including via commercial APIs, managed platforms (e.g. Bedrock, Vertex AI), or self-hosted open-source models — and their practical capabilities and limitations.AI-specific security risks — prompt injection, unsafe tool or agent execution, data leakage, excessive permissions, and supply-chain risks.AI-specific privacy and governance concerns — how training and retrieval data is handled, and how model, tool, and agent access is governed, data retention and lineage.Working with AI gateways, model routing, prompt management, or agent frameworks.Implementing evaluation, tracing, monitoring, or quality controls for AI applications.Monitoring token usage, inference resources, and AI platform costs.Advantageous ExperienceExperience working in telecommunications or another highly distributed, security-sensitive environment.Experience with MLOps, LLMOps, machine learning platforms, or data platforms.Experience operating GPU-based workloads or shared compute environments.Knowledge of model lifecycle management, model registries, evaluation frameworks, and experiment tracking.Experience with data lakes, data warehouses, streaming platforms, vector databases, or enterprise search.Experience implementing policy as code or automated compliance controls.Familiarity with recognised AI risk and security frameworks.Experience consolidating fragmented tools or building an internal developer platform.Strong networking or network security knowledge.