Sentry
Senior Software Engineer, AI Evals
Part remote, part office, you need to live within commuting distance of a named location.
Hybrid · San Francisco, California
Employer listed it 8 months ago · Added 4 days ago
Been open since 8 months ago. Long-running listings are sometimes left up after the role is filled.
Salary
$155,000–$400,000
Location
Hybrid · San Francisco, California
Timezone
Not stated
Contract
Full-time
Experience
Senior
Category
Software
Published by the employer
Remote flexibility
Hybrid
This role is only partly remote, the employer expects time in the office around San Francisco, California, Hybrid, so you need to live within commuting distance.
What the employer says
- Source listing states candidate location: "San Francisco, California, Hybrid"
- Listing mentions "Hybrid"
What Nomaders makes of it
- Not suitable if you plan to move between countries
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About Sentry
Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building.
Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future.
About the role
As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence.
In this role you will
Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems
Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data
Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows
Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria
Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring
You’ll love this job if you
Care deeply about correctness, rigor, and measurement in AI systems
Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics
Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team
Thrive in cross-functional environments and enjoy influencing model design through better evaluation
Qualifications
Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learning, or a related field
Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred)
Comfort writing production-quality code (we use Python and TypeScript)
Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines
Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts)
Bonus: experience evaluating LLMs, agentic systems, or AI-assisted developer tools
The base salary range (or hourly wage range, if applicable) that Sentry reasonably expects to pay for this position is $155,000 to $400,000 USD . A successful candidate’s actual base salary (or hourly wage) amount will be determined by a variety of relevant factors including, without limitation, the candidate’s work location, education, work and other relevant experience, skills, and job-related knowledge. A successful candidate will be eligible to participate in Sentry’s employee benefit plans/programs applicable to the candidate’s position (including incentive compensation, equity grants, paid time off, and group health insurance coverage). See Sentry Benefits for more details about the Company’s benefit plans/programs.
Requirements
- ·Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learning, or a related field
- ·Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred)
- ·Comfort writing production-quality code (we use Python and TypeScript)
- ·Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines
- ·Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts)
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, hybrid, matches where you plan to live and work.
- 2Tailor your CV to the role at Sentry, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 5d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
$155,000–$400,000 · You'll be taken to the employer's careers page.