LiveKit
Research Engineer (Reinforcement Learning)
Remote work allowed only within certain countries or regions.
North America, Europe
Employer listed it 5 weeks ago · Added yesterday
Been open since 5 weeks ago, still checked daily, but it has been live a while.
Salary
$135,000–$300,000
Location
North America, Europe
Timezone
US East
Contract
Full-time
Experience
Mid
Category
Software
Published by the employer
Remote flexibility
Region Restricted
Remote work is allowed, but only for candidates based in North America, Europe.
What the employer says
- Source listing states candidate location: "North America, APJ, EMEA, Remote"
What Nomaders makes of it
- Applications outside the listed area are usually rejected
- Timezone overlap with the listed area is often expected
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About LiveKit
LiveKit is building the infrastructure layer for the voice-driven era of computing. Our platform gives developers everything they need to build, test, deploy, scale, and observe agents in production. Founded in 2021, LiveKit powers voice AI applications for OpenAI, xAI, Salesforce, Coursera, Spotify, and thousands of others, collectively facilitating billions of calls each year.
About This Role
We are looking for an exceptional engineer to build post-training at LiveKit. Our agents run over voice and increasingly over text channels like SMS and chat, and the interesting problems show up over long horizons: staying useful across many sessions, working with context that accumulates over time, and using tools reliably in the middle of a live conversation.
What You'll Do
Build the environments and verifiers our models train against
Own the synthetic data pipeline, from generation through the quality gates
Run training experiments end to end, and explain what moved the model
Build the evaluations a release has to clear
Choose and adapt open-weight base models for our tasks
Make trained behavior hold up for voice and text agents alike
Ship models into production and keep improving them on real usage
Who You Are
A strong Python engineer
Have carried a model from raw data through to production
Treat data as the product: coverage, diversity, leakage
Assume a model will exploit a weak reward, and design against it
Comfortable with GPUs and honest about their limits
Know when to train, and when not to
Comfortable working collaboratively in a remote environment
Nice to Have
Experience with post-training: fine-tuning, reward design, or reinforcement learning such as GRPO
RL and fine-tuning frameworks such as TRL, verl, or OpenRLHF, or a training loop you wrote yourself
Fast rollouts with vLLM or SGLang, multi-GPU training with FSDP
Requirements
The employer hasn't listed requirements separately, they're described in the role summary above and on the original listing.
Benefits
- ·Flexible vacation policy
How to apply
- 1Check the flexibility label above, region restricted, matches where you plan to live and work.
- 2Tailor your CV to the role at LiveKit, mentioning your remote working experience and working hours (US East).
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 1d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
$135,000–$300,000 · You'll be taken to the employer's careers page.