← Back to jobs

AI Evaluation Engineer

  • Remote
  • Sweden
  • English
  • Posted 24.09.26 15:21

JOB DESCRIPTION Position: AI Evaluation EngineerLocation: EU – Remote Job Type: Contract Requirements· 5+ years of hands-on work in AI/ML evaluation, NLP, or applied data science,including recent delivery on LLM, RAG, or agent systems.· Has built evaluation frameworks and gold datasets for information extractionand retrieval/ranking systems. Please share examples.· Strong grounding in evaluation metrics and statistics: precision/recall/F1.Recall@K, MRR, nDCG, inter-rater reliability (e.g., Cohen’s kappa orKrippendorff’s alpha), and confidence intervals for small evaluation sets.· Experience designing and validating LLM-as-judge systems against human expertratings.· Experience running annotation projects with domain experts: guidelines,adjudication, and quality control.· Strong Python; experience with evaluation harnesses (e.g., Ragas, DeepEval,promptfoo, Inspect, or in-house equivalents), LLM APIs, and Git/CI/CD.· Can deliver against fixed milestones with little onboarding and workindependently.· Fluent professional English, written and spoken.