Salla logo

Salla

More roles at Salla

Senior AI Backend Engineer - Agent Evaluation & Quality

Work from home

Remote role where the employee must remain based in a particular country.

Saudi Arabia only

Employer listed it 2 months ago · Added yesterday

Been open since 2 months ago. Long-running listings are sometimes left up after the role is filled.

Salary

Not stated

Location

Saudi Arabia only

Timezone

Not stated

Contract

Full-time

Experience

Senior

Category

Software

This employer didn't state pay. Jobs like this usually pay around $170k–$225k a year, a typical range taken from 599 senior-level software roles on Nomaders that do state pay. It's a guide, not an offer.

Remote flexibility

Work from home

This is a remote role, but the employee must be based in Saudi Arabia. It is work from home rather than work from anywhere.

What the employer says

  • Source listing states candidate location: "Makkah, Saudi Arabia, Makkah, Makkah Province, Saudi Arabia, Makkah, Makkah Province, Saudi Arabia, Remote"

What Nomaders makes of it

  • Residency required in Saudi Arabia
  • Payroll and tax are likely handled in that country only

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

About the role

We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that.

You'll build the evaluation systems behind our agents - the judges, test harnesses, and simulators that tell us whether an agent is working and where it's failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us.

Evaluation is the focus, but it won't be the boundary. Because you'll understand the agents' failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them.

Responsibilities

Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode.

Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes.

Build user simulators to generate test coverage and adversarial cases before real users hit them.

Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time.

Partner with product to turn "what good looks like" into concrete, measurable criteria.

Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.

Requirements

Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering.

Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break.

A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.

Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability.

5+ years software engineering, with recent hands-on LLM/agent work.

Nice to have

Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing.

Observability tooling (Arize, LangSmith, or similar).

Arabic language / NLP experience.

E-commerce or merchant-facing product experience.

Requirements

  • ·Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering.
  • ·Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break.
  • ·A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.
  • ·Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability.
  • ·5+ years software engineering, with recent hands-on LLM/agent work.

Benefits

No benefits package published with this listing. Ask about it at first interview.

How to apply

  1. 1Check the flexibility label above, work from home, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Salla, mentioning your remote working experience.
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 1d ago. Last checked 25 Sept. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open software roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

Typically $170k to $225k per year · You'll be taken to the employer's careers page.