This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Policy Evaluator, Violence & Threats based in United States.
This role sits at the intersection of AI safety, content policy, and expert human judgment, helping improve how advanced AI models handle violent and threatening content.
You’ll evaluate user requests, model responses, and conversation context to distinguish legitimate fictional, educational, historical, or defensive content from material that could enable real-world harm.
The work focuses heavily on nuanced edge cases where intent, context, and a single detail can materially change the appropriate policy decision.
You’ll contribute not only to evaluations, but also to adversarial testing, policy refinement, calibration, and the identification of emerging safety gaps.
This is a non-engineering role for someone with deep experience in areas such as violent fiction, military or emergency response, crisis intervention, threat assessment, trust and safety, or related fields.
You’ll work in a structured, feedback-rich environment where clear reasoning, consistency, and the ability to separate personal beliefs from policy standards are essential.
Because the role involves regular exposure to difficult material, candidates must be prepared to engage with sensitive content carefully, professionally, and sustainably.
Accountabilities- Evaluate user requests, AI responses, and conversation histories involving violence, weapons, threats, self-harm, dark fiction, and related sensitive topics.
- Distinguish fictional, educational, historical, journalistic, defensive, or expressive content from requests that meaningfully facilitate real-world harm.
- Assess whether an AI response provides actionable real-world capability, regardless of how the original request is framed.
- Differentiate ordinary anger, frustration, venting, or dark humor from credible threats and potential crisis indicators.
- Apply relevant customer policies consistently while considering intent, context, precedent, and established team guidance rather than relying on rote annotation.
- Select defensible classifications for ambiguous cases and produce concise, well-supported rationales referencing policy language and relevant conversation details.
- Develop and refine adversarial or borderline prompts that test how AI systems handle difficult policy boundaries.
- Identify policy gaps, contradictions, recurring ambiguities, and emerging edge cases, escalating findings to project leads and policy teams.
- Participate actively in calibration and adjudication discussions, respectfully challenging interpretations and updating judgments when stronger reasoning emerges.
- Maintain high accuracy, consistency, and attention to detail across repetitive, feedback-heavy evaluation workflows.
- Contribute to improving AI safety standards by turning nuanced human judgment into clear, auditable evaluation guidance.
Requirements
- Demonstrated depth of experience in at least one relevant domain, such as violent fiction, game design or game mastering, film and television, military or law enforcement, emergency medicine, crisis counseling, threat assessment, trust and safety, content moderation, journalism, or law.
- Strong ability to distinguish fictional or contextualized depictions of violence from content that facilitates or signals real-world harm.
- Demonstrated understanding of how intent, context, language, and subtle changes in a request can materially affect a safety assessment.
- Ability to make nuanced judgment calls while separating personal beliefs from the policy standard being applied.
- Strong written communication skills, with the ability to explain complex decisions clearly enough for another evaluator to audit the reasoning.
- Ability to remain open-minded, challenge assumptions constructively, and revise conclusions when stronger evidence or reasoning emerges.
- Strong attention to detail and consistency when working through repeated evaluations involving difficult or sensitive material.
- Familiarity with AI tools and a strong interest in understanding where language models may over-refuse, under-refuse, or misunderstand user intent.
- Prior experience with AI evaluation, red teaming, data annotation, RLHF, trust and safety, content moderation, or language-model assessment is helpful but not required.
- Experience with calibration sessions, inter-rater agreement, adjudication workflows, or structured policy evaluation is a plus.
- Additional valuable experience may include published or produced work involving violence, military or security experience, crisis intervention, threat assessment, forensic or clinical psychology, violence prevention, weapons disciplines, or professional experience evaluating ChatGPT, Claude, Gemini, or similar AI systems.
- A degree, security clearance, or technical/software engineering background is not required.
- Ability to work remotely in the United States on a Monday-Friday schedule from 8:00 AM to 5:00 PM PT.
Benefits
- Compensation: $45–$55 per hour.
- Employment classification: W-2.
- Work arrangement: Fully remote within the United States.
- Schedule: Monday through Friday, 8:00 AM–5:00 PM Pacific Time.
- Assignment: Ongoing opportunity with a planned start date of September 21, 2026.
- Benefits eligibility: Eligible for available employee benefits.
- Structured evaluation frameworks, professional guidelines, content rotation, and exposure limits designed to support sustainable work with sensitive material.
- Access to mental health support given the nature of the content reviewed.
- Opportunity to contribute directly to the safety, reliability, and responsible development of advanced AI systems.
- Exposure to complex policy questions and emerging challenges at the intersection of AI, violence, threats, and content safety.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1