Salla
More roles at SallaSenior AI Backend Engineer - Agent Evaluation & Quality
Remote role where the employee must remain based in a particular country.
Saudi Arabia only
Employer listed it 2 months ago · Added yesterday
Been open since 2 months ago. Long-running listings are sometimes left up after the role is filled.
Salary
Not stated
Location
Saudi Arabia only
Timezone
Not stated
Contract
Full-time
Experience
Senior
Category
Software
This employer didn't state pay. Jobs like this usually pay around $170k–$225k a year, a typical range taken from 599 senior-level software roles on Nomaders that do state pay. It's a guide, not an offer.
Remote flexibility
Work from home
This is a remote role, but the employee must be based in Saudi Arabia. It is work from home rather than work from anywhere.
What the employer says
- Source listing states candidate location: "Makkah, Saudi Arabia, Makkah, Makkah Province, Saudi Arabia, Makkah, Makkah Province, Saudi Arabia, Remote"
What Nomaders makes of it
- Residency required in Saudi Arabia
- Payroll and tax are likely handled in that country only
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About the role
We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that.
You'll build the evaluation systems behind our agents - the judges, test harnesses, and simulators that tell us whether an agent is working and where it's failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us.
Evaluation is the focus, but it won't be the boundary. Because you'll understand the agents' failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them.
Responsibilities
Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode.
Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes.
Build user simulators to generate test coverage and adversarial cases before real users hit them.
Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time.
Partner with product to turn "what good looks like" into concrete, measurable criteria.
Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.
Requirements
Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering.
Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break.
A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.
Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability.
5+ years software engineering, with recent hands-on LLM/agent work.
Nice to have
Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing.
Observability tooling (Arize, LangSmith, or similar).
Arabic language / NLP experience.
E-commerce or merchant-facing product experience.
Requirements
- ·Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering.
- ·Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break.
- ·A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.
- ·Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability.
- ·5+ years software engineering, with recent hands-on LLM/agent work.
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, work from home, matches where you plan to live and work.
- 2Tailor your CV to the role at Salla, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 1d ago. Last checked 25 Sept. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
Typically $170k to $225k per year · You'll be taken to the employer's careers page.