← Takaisin työpaikkoihin

AI Evaluation Engineer

  • Etätyö
  • Ruotsi
  • Englanti
  • Julkaistu 23.09.26 21:01

Required experience:5+ years of hands-on work in AI/ML evaluation, NLP, or applied data science, including recent delivery on LLM, RAG, or agent systems.Has built evaluation frameworks and gold datasets for information extraction and retrieval/ranking systems. Please share examples.Strong grounding in evaluation metrics and statistics: precision/recall/F1,Recall@K, MRR, nDCG, inter-rater reliability (e.g., Cohen’s kappa orKrippendorff’s alpha), and confidence intervals for small evaluation sets.Experience designing and validating LLM-as-judge systems against human expert ratings.Experience running annotation projects with domain experts: guidelines, adjudication, and quality control.Strong Python; experience with evaluation harnesses (e.g., Ragas, DeepEval, promptfoo, Inspect, or in-house equivalents), LLM APIs, and Git/CI/CD.Can deliver against fixed milestones with little onboarding and work independently.Fluent professional English, written and spoken. Nice to have:Experience with regulatory affairs, regulatory intelligence, or submissions:FDA/EMA guidance, CTD structure (e.g., Modules 2.5 and 2.7), Health Authority questions, and labeling.Experience evaluating AI in regulated (GxP) environments.Familiarity with clinical trial outputs (CSRs, TLFs, statistical outputs) and drug development.Experience evaluating document-comparison, claim-verification, or fact-checking systems.Publications or open-source work in AI evaluation, information retrieval, or biomedical NLP. 5+ years of hands-on work in AI/ML evaluation, NLP, or applied data science, including recent delivery on LLM, RAG, or agent systems.Has built evaluation frameworks and gold datasets for information extraction and retrieval/ranking systems. Please share examples.Strong grounding in evaluation metrics and statistics: precision/recall/F1,Recall@K, MRR, nDCG, inter-rater reliability (e.g., Cohen’s kappa orKrippendorff’s alpha), and confidence intervals for small evaluation sets.Experience designing and validating LLM-as-judge systems against human expert ratings.Experience running annotation projects with domain experts: guidelines, adjudication, and quality control.Strong Python; experience with evaluation harnesses (e.g., Ragas, DeepEval, promptfoo, Inspect, or in-house equivalents), LLM APIs, and Git/CI/CD.Can deliver against fixed milestones with little onboarding and work independently.Fluent professional English, written and spoken. Nice to haveExperience with regulatory affairs, regulatory intelligence, or submissions:FDA/EMA guidance, CTD structure (e.g., Modules 2.5 and 2.7), Health Authority questions, and labeling.Experience evaluating AI in regulated (GxP) environments.Familiarity with clinical trial outputs (CSRs, TLFs, statistical outputs) and drug development.Experience evaluating document-comparison, claim-verification, or fact-checking systems.Publications or open-source work in AI evaluation, information retrieval, or biomedical NLP.