← Back to jobs

AI Evaluation Engineer (LLM Quality Assurance)

  • Remote
  • Sweden
  • English
  • Posted 07.09.26 12:45

Job Title: AI Evaluation Engineer (LLM Quality Assurance)Engagement StatusThis is a freelance, project-based opportunity. Specific project details, including the scope, timeline, workload, and engagement arrangements, are currently being finalized with the client.At this stage, we are proactively connecting with qualified professionals who may be interested in participating in the project. Candidates who are a potential fit will be considered and contacted as the project details are confirmed and the engagement moves forward.Job SummaryWe are seeking an AI Evaluation Engineer to assess, validate, and improve AI task quality across end-to-end evaluation pipelines. The ideal candidate possesses strong software engineering skills, hands-on experience with Python-based environments, and the ability to identify root causes of issues within AI evaluation, benchmarking, and rollout processes.Key ResponsibilitiesTask & Pipeline Evaluation: Evaluate and analyze AI tasks, benchmarks, and execution traces to identify quality issues and performance gaps across evaluation pipelines.Root Cause Analysis: Investigate root causes of failures and determine whether issues originate from task design, evaluation metrics, reference solutions, execution environments, or rollout contamination.Methodology Validation: Validate evaluation methodologies, ensuring reproducibility, environment consistency, model discrimination, and proper visible/hidden evaluation isolation.Hands-on Debugging: Execute and troubleshoot tasks locally using Python, Shell, Docker, and AI coding assistants.Log & Trace Review: Review evaluation results, artifacts, logs, and traces, providing evidence-based recommendations and quality assessments.Decision & Reporting: Produce clear Pass, Conditional Pass, or Reject decisions with actionable improvement plans and comprehensive documentation.Basic RequirementsBachelor's degree in Computer Science, Engineering, Physics, Data Science, AI, or a related technical field, or equivalent practical experience.Strong hands-on programming experience in Python, Shell scripting, and Docker-based environments.Practical experience in AI model evaluation, benchmarking, testing, validation, quality assurance, or research operations.Hands-on experience using AI coding tools (such as Claude Code, Codex, or similar tools) to test, debug, and validate tasks locally.Proven ability to analyze execution traces, debugging logs, and evaluation outputs to identify root causes across complex workflows.Strong analytical thinking, problem-solving, and technical reporting skills.Preferred QualificationsExperience in AI/ML research, model evaluation, reinforcement learning, or benchmark development.Experience designing evaluation methodologies and ensuring reproducible experiment results.Deep knowledge of model comparison, testing frameworks, and AI quality assessment practices.What Success Looks LikeIssue Resolution: Consistently identify and resolve quality issues within AI evaluation pipelines.Accurate Assessments: Deliver accurate, evidence-backed evaluation conclusions.Quality & Reproducibility: Improve benchmark quality, reproducibility, and model assessment reliability.Standard Setting: Support the development of high-quality datasets and evaluation standards for AI systems.