Anthropic logo

Anthropic

Cyber Evaluations Engineer

Work from homeNew this week

Remote role where the employee must remain based in a particular country.

United States only

Employer listed it yesterday · Added yesterday

First listed yesterday.

Salary

$300k to $405k per year

Location

United States only

Timezone

US East

Contract

Full-time

Experience

Mid

Category

Software

Stated by the employer in the job description

Remote flexibility

Work from home

This is a remote role, but the employee must be based in United States. It is work from home rather than work from anywhere.

What the employer says

  • Source listing states candidate location: "Remote-Friendly, United States; San Francisco, CA | Washington, DC, San Francisco, CA"

What Nomaders makes of it

  • Residency required in United States
  • Payroll and tax are likely handled in that country only

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

We're hiring Cyber Evaluations Engineers to build and run the evaluations that measure cyber-relevant capabilities and safeguard robustness in our models. You'll design new evals, run per-release robustness testing, and dig into data on jailbreaks and prompt bypasses to understand where our safeguards hold up and where they don't. You'll also design many of the probes that detect cyber abuse in production and help shape the overall detection architecture alongside the policy team.

Key responsibilities

Design and run capability, uplift, and safety evaluations to assess cyber-relevant risk in new models

Execute per-release safeguard-robustness testing ahead of major model launches

Analyze evaluation results and communicate findings clearly to the team and to stakeholders

Design, prototype, and tune detection probes for cyber misuse

Work with the cyber policy team to turn policy lines into a layered, robust abuse-detection architecture, and measure its precision and coverage over time

Build and maintain internal tooling used to run and score evaluations

Collaborate with policy and engineering partners to translate eval findings into safeguard improvements

Minimum qualifications

Experience building or running evaluations, benchmarks, or test suites for software or ML systems, including delivering results on short, fixed timelines

Hands-on cybersecurity experience (e.g., CTF participation, vulnerability research, exploit development, or security research)

Proficiency in Python

Strong ability to communicate evaluation results with multiple cross-functional stakeholders or potential policy stakeholders

Preferred qualifications

Deep offensive-security or security-research experience, including experience building AI security benchmarks

Experience analyzing adversarial or abuse data (e.g., jailbreaks, prompt bypasses, intrusion or fraud telemetry)

Experience working onsite with government partners on testing or evaluation engagements

Experience with AI/ML evaluation frameworks

Familiarity with coordinated vulnerability disclosure practices

Experience testing pre-release or pre-deployment software or models under confidentiality constraints

Requirements

  • ·Experience building or running evaluations, benchmarks, or test suites for software or ML systems, including delivering results on short, fixed timelines
  • ·Hands-on cybersecurity experience (e.g., CTF participation, vulnerability research, exploit development, or security research)
  • ·Proficiency in Python

Benefits

No benefits package published with this listing. Ask about it at first interview.

How to apply

  1. 1Check the flexibility label above, work from home, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Anthropic, mentioning your remote working experience and working hours (US East).
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 1d ago. Last checked 23 Sept. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open software roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

$300k to $405k per year · You'll be taken to the employer's careers page.