← Takaisin työpaikkoihin

LLM Evaluation Engineer

  • Etätyö
  • Ruotsi
  • Englanti
  • Julkaistu 10.09.26 12:38

LLM Evaluation EngineerRemote — European Union. Company based in Paris, France.Enterprise AI | LLM Evaluation | RAG | AI Observability | Production AI Our client, a growing Enterprise AI company based in Paris, is looking for an LLM Evaluation Engineer to build the evaluation systems used to measure the quality, reliability, and behaviour of production LLM and RAG applications. You'll work at the intersection of AI Engineering, Evaluation, and Observability, creating automated frameworks that turn subjective AI quality into measurable engineering signals. What You'll Work On• Build automated evaluation pipelines for production LLM applications• Design evaluation frameworks for RAG and retrieval-based systems• Create datasets and test suites for repeatable AI evaluation• Measure answer quality, groundedness, relevance, and consistency• Evaluate retrieval quality and its impact on downstream LLM responses• Build regression testing for prompts, models, retrieval strategies, and application changes• Develop evaluation and experimentation tooling in Python• Integrate tracing and telemetry using OpenTelemetry• Analyse production behaviour using SQL and application data• Compare models and configurations through structured experimentation• Build monitoring and feedback loops to identify quality degradation• Work with AI Engineers to turn evaluation results into product improvements Core Skills• 3+ years in AI Engineering, ML Engineering, Applied ML, Data Science, or similar roles• Python• LLM evaluation• RAG architectures• LLM APIs• OpenTelemetry or comparable tracing/observability technologies• SQL• Experiment design and quantitative evaluation• Strong understanding of production LLM applications Nice to HaveRAG evaluation frameworksLangSmith / Langfuse / Arize PhoenixRagas / DeepEval or similar toolingRetrieval evaluation and ranking metricsLLM-as-a-judge techniquesHuman evaluation workflowsPrompt regression testingLangChain / LangGraphVector databasesStatistical analysis / A/B testingProduction AI observability