← Back to jobs

AI Safety Practitioner

  • Remote
  • Sweden
  • English
  • Posted 29.08.26 13:31

AI Safety Practitioner [$60-$70/hr]

AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics and assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback

Role Responsibilities

Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall qualityReview content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domainsApply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarkingIdentify unsafe outputs, hallucinations, reasoning failures, and policy violationsProvide structured feedback to improve model alignment and safety performanceCollaborate with AI researchers and safety teams on ongoing evaluation initiatives

Good Candidature

Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related fieldExcellent written English, critical thinking, and analytical reasoning skillsAbility to consistently evaluate nuanced and policy-sensitive scenarios

Nice to Have

Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluationFamiliarity with safety policies, content moderation, or evaluation rubric developmentExperience reviewing complex, high-risk, or ambiguous content