Anthropic
Red Team Engineer, Safeguards
Remote role where the employee must remain based in a particular country.
United States only
Employer listed it 5 weeks ago · Added 5 days ago
Been open since 5 weeks ago, still being checked, but it has been live a while.
Salary
$320k to $405k per year
Location
United States only
Timezone
Not stated
Contract
Full-time
Experience
Mid
Category
Software
Stated by the employer in the job description
Remote flexibility
Work from home
This is a remote role, but the employee must be based in United States. It is work from home rather than work from anywhere.
What the employer says
- Source listing states candidate location: "Remote-Friendly (Travel Required) | San Francisco, CA, Remote-Friendly US (Travel Required)"
What Nomaders makes of it
- Residency required in United States
- Payroll and tax are likely handled in that country only
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the role
Anthropic's Safeguards team is seeking a Red Team Engineer to help ensure the safety of our deployed AI systems and products. In this role, you'll take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors. Your work will span from technical infrastructure vulnerabilities on our products to emergent risks from advanced AI capabilities.
While you'll bring best practices from traditional security approaches, the focus is on broader safety implications and novel abuse unique to advanced AI systems and associated products. You'll investigate the full spectrum of potential abuse — from coordinated account manipulation and payment fraud to novel exploitation of product features — and simulate sophisticated threat actors who chain multiple attack vectors to achieve their objectives.
Key responsibilities
Conduct comprehensive adversarial testing across Anthropic's product surfaces, developing creative attack scenarios that combine multiple exploitation techniques
Research and implement novel testing approaches for emerging capabilities, including agent systems, tool use, and new interaction paradigms
Design and execute "full kill chain" attacks that emulate real-world threat actors attempting to achieve specific malicious objectives
Build and maintain systematic testing methodologies that evaluate every aspect of our systems
Develop automated testing frameworks to enable continuous assessment at scale
Collaborate with Product, Engineering, and Policy teams to translate findings into concrete improvements
Help establish metrics for measuring detection effectiveness of novel abuse
Minimum qualifications
Experience in penetration testing, red teaming, or application security
Experience in model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors
Strong technical skills in web application security, including hands-on expertise with security testing tools (e.g., Burp Suite, Metasploit, custom scripting frameworks)
Experience building custom automation, including LLM-specific testing frameworks
A track record of discovering novel attack vectors and chaining vulnerabilities in creative ways
A public body of work such as CVEs, blog posts, or disclosed bug bounty reports
Strong written and verbal communication skills, with the ability to explain technical concepts to varied audiences
Preferred qualifications
Experience with AI/ML security or adversarial machine learning
Understanding of AI safety considerations beyond traditional security, including modern guardrails against jailbreaks
Requirements
- ·Experience in penetration testing, red teaming, or application security
- ·Experience in model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors
- ·Strong technical skills in web application security, including hands-on expertise with security testing tools (e.g., Burp Suite, Metasploit, custom scripting frameworks)
- ·Experience building custom automation, including LLM-specific testing frameworks
- ·A track record of discovering novel attack vectors and chaining vulnerabilities in creative ways
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, work from home, matches where you plan to live and work.
- 2Tailor your CV to the role at Anthropic, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 5d ago. Last checked 23 Sept. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
$320k to $405k per year · You'll be taken to the employer's careers page.