AI Safety Practitioner
AI Safety Practitioner [$60-$70/hr]
AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics and assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback
Role Responsibilities
Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall qualityReview content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domainsApply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarkingIdentify unsafe outputs, hallucinations, reasoning failures, and policy violationsProvide structured feedback to improve model alignment and safety performanceCollaborate with AI researchers and safety teams on ongoing evaluation initiatives
Good Candidature
Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related fieldExcellent written English, critical thinking, and analytical reasoning skillsAbility to consistently evaluate nuanced and policy-sensitive scenarios
Nice to Have
Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluationFamiliarity with safety policies, content moderation, or evaluation rubric developmentExperience reviewing complex, high-risk, or ambiguous content
