Kairos
Back to jobs

Safeguards Enforcement Analyst, Conventional Weapons

Remote
AnthropicSan Francisco, CA, US / Washington, DC, US2 weeks agoWebsite
May be filledHiring slowdown reported
Safeguards (Trust & Safety)

Compensation

$245,000-$330,000/yr
Apply
Share

Description

About the role

As a Safeguards Enforcement Analyst focused on Conventional Weapons, your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, specifically utilizing conventional weapons and dangerous technology. You will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas. 


Important context for this role: In this position you may be exposed to and engage with explicit content spanning a range of topics, including those of a violent, graphic, hateful, or psychologically disturbing nature.

Key responsibilities

  • Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy

  • Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements

  • Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations

  • Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging tactics

  • Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios

  • Develop and maintain enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content

  • Keep up to date with emerging weapons trends and applications, regulatory changes, and AI policy enforcement best practices, and apply these to inform our workflows and evals

  • Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent activity

Minimum qualifications

  • Have deep, applied expertise in weapons systems and can translate complex technical evidence to make enforcement decisions

  • Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, or physical harm facilitation

  • Experience standing up and scaling policy enforcement or content review workflows

  • Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health

  • Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams

  • Experience working with generative AI products, including writing effective prompts for content review and enforcement

  • Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space

Preferred qualifications

  • Subject matter expertise in conventional weapons and dangerous technology, autonomous systems, or critical infrastructure protection

  • Familiarity with relevant legal and regulatory frameworks governing dangerous technology, conventional weapons, and critical infrastructure

  • Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases

  • Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)

  • Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments

  • Experience with large language models and an understanding of how AI technology could provide meaningful uplift toward serious harm

  • Proficiency in Python for data analysis and workflow automation

  • Background in law enforcement, national security, defense, counterterrorism, or a relevant regulatory environment

  • Experience assessing the technical plausibility and real-world harm potential of content, including the ability to distinguish between general educational content and genuine operational uplift, and between protected speech and genuine incitement/mobilization

  • Familiarity with cross-platform threat analysis and OSINT techniques

Stack

LLMsGenerative AIPythonData ScienceSQL
Posted
Sep 2, 2026
Last seen
Sep 2, 2026
First seen
Sep 2, 2026

Similar roles

Browse more AI jobs