Sentry logo

Sentry

Senior Software Engineer, AI Evals

Hybrid

Part remote, part office, you need to live within commuting distance of a named location.

Hybrid · San Francisco, California

Employer listed it 8 months ago · Added 4 days ago

Been open since 8 months ago. Long-running listings are sometimes left up after the role is filled.

Salary

$155,000–$400,000

Location

Hybrid · San Francisco, California

Timezone

Not stated

Contract

Full-time

Experience

Senior

Category

Software

Published by the employer

Remote flexibility

Hybrid

This role is only partly remote, the employer expects time in the office around San Francisco, California, Hybrid, so you need to live within commuting distance.

What the employer says

  • Source listing states candidate location: "San Francisco, California, Hybrid"
  • Listing mentions "Hybrid"

What Nomaders makes of it

  • Not suitable if you plan to move between countries

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

About Sentry

Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building.

Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future.

About the role

As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence.

In this role you will

Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems

Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data

Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows

Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria

Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring

You’ll love this job if you

Care deeply about correctness, rigor, and measurement in AI systems

Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics

Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team

Thrive in cross-functional environments and enjoy influencing model design through better evaluation

Qualifications

Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learning, or a related field

Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred)

Comfort writing production-quality code (we use Python and TypeScript)

Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines

Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts)

Bonus: experience evaluating LLMs, agentic systems, or AI-assisted developer tools

The base salary range (or hourly wage range, if applicable) that Sentry reasonably expects to pay for this position is $155,000 to $400,000 USD . A successful candidate’s actual base salary (or hourly wage) amount will be determined by a variety of relevant factors including, without limitation, the candidate’s work location, education, work and other relevant experience, skills, and job-related knowledge. A successful candidate will be eligible to participate in Sentry’s employee benefit plans/programs applicable to the candidate’s position (including incentive compensation, equity grants, paid time off, and group health insurance coverage). See Sentry Benefits for more details about the Company’s benefit plans/programs.

Requirements

  • ·Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learning, or a related field
  • ·Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred)
  • ·Comfort writing production-quality code (we use Python and TypeScript)
  • ·Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines
  • ·Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts)

Benefits

No benefits package published with this listing. Ask about it at first interview.

How to apply

  1. 1Check the flexibility label above, hybrid, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Sentry, mentioning your remote working experience.
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 5d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open software roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

$155,000–$400,000 · You'll be taken to the employer's careers page.