
Compensation
$245,000-$330,000/yrDescription
About the role
As a Safeguards Enforcement Analyst focused on Conventional Weapons, your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, specifically utilizing conventional weapons and dangerous technology. You will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas.
Important context for this role: In this position you may be exposed to and engage with explicit content spanning a range of topics, including those of a violent, graphic, hateful, or psychologically disturbing nature.
Key responsibilities
-
Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy
-
Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements
-
Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
-
Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging tactics
-
Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios
-
Develop and maintain enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content
-
Keep up to date with emerging weapons trends and applications, regulatory changes, and AI policy enforcement best practices, and apply these to inform our workflows and evals
-
Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent activity
Minimum qualifications
-
Have deep, applied expertise in weapons systems and can translate complex technical evidence to make enforcement decisions
-
Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, or physical harm facilitation
-
Experience standing up and scaling policy enforcement or content review workflows
-
Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
-
Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams
-
Experience working with generative AI products, including writing effective prompts for content review and enforcement
-
Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space
Preferred qualifications
-
Subject matter expertise in conventional weapons and dangerous technology, autonomous systems, or critical infrastructure protection
-
Familiarity with relevant legal and regulatory frameworks governing dangerous technology, conventional weapons, and critical infrastructure
-
Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases
-
Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
-
Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments
-
Experience with large language models and an understanding of how AI technology could provide meaningful uplift toward serious harm
-
Proficiency in Python for data analysis and workflow automation
-
Background in law enforcement, national security, defense, counterterrorism, or a relevant regulatory environment
-
Experience assessing the technical plausibility and real-world harm potential of content, including the ability to distinguish between general educational content and genuine operational uplift, and between protected speech and genuine incitement/mobilization
-
Familiarity with cross-platform threat analysis and OSINT techniques
Stack
- Posted
- Sep 2, 2026
- Last seen
- Sep 2, 2026
- First seen
- Sep 2, 2026