Babbel logo

Babbel

Senior Machine Learning Engineer (all genders) - Babbel Labs

Hybrid

Part remote, part office, you need to live within commuting distance of a named location.

Hybrid · Berlin

Employer listed it 13 days ago · Added yesterday

First listed 13 days ago and still open.

Salary

Not stated

Location

Hybrid · Berlin

Timezone

Not stated

Contract

Full-time

Experience

Senior

Category

Data

This employer didn't state pay. Jobs like this usually pay around $145k–$210k a year, a typical range taken from 255 senior-level data roles on Nomaders that do state pay. It's a guide, not an offer.

Remote flexibility

Hybrid

This role is only partly remote, the employer expects time in the office around Berlin, Hybrid, so you need to live within commuting distance.

What the employer says

  • Source listing states candidate location: "Berlin, Hybrid"
  • Listing mentions "Hybrid"

What Nomaders makes of it

  • Not suitable if you plan to move between countries

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

Senior Machine Learning Engineer (all genders) - Babbel Labs

About Babbel Labs

Babbel Labs is building the future of language learning. We are an AI-first, independent company within the Babbel group, based in Berlin. Our teams bring AI, research and product together to ship experiences that set a new standard for how people learn.

The role

Babbel's learner-personalisation engine tracks what a learner has and hasn't mastered, and decides what they should practice next. We are actively pushing personalisation further; moving beyond describing a user and prescribing targeted practice, to adapting the learning journey ahead of them.

This is a hands-on senior individual-contributor role on that team. You will own real subsystems end to end and make decisions about how to improve the personalisation engine. That includes everything from research and benchmarking through production, monitoring and incidents that follow six months later.

How you'll make an impact

Work directly with the Principal Scientist to take designs from spec into production, then keep improving what you’ve built on your own judgement rather than waiting to be told what’s next.

Help shape new features from the beginning - not just implementation of a spec handed to you.

Take real ownership of core personalisation and mastery-tracking subsystems, operate independently, and make decisions confidently and competently.

Design the evaluation that tells you whether a change is real: a rigorous offline benchmark against a real baseline, and the online experiment that confirms or kills it.

Run what you build. Instrument it, notice when it's silently wrong rather than only when it errors, and fix it before it becomes an incident.

Deliver with coding agents as a matter of course, and verify what they produce before you rely on it.

Your skills and qualifications

Strong, hands-on ML engineering that has shipped real models to production — recommendation, ranking, scoring, or trust-and-safety systems under real user load are the closest match. Research or competition experience is a plus.

Experience with probabilistic modeling, latent-variable modeling and Bayesian inference, or the equivalent rigor from an adjacent domain.

ML system evaluation: monitoring metrics you define, debugging output that doesn’t look right, and rolling out a change to a live scoring or ranking system without breaking it.

Rigorous experimentation practice: benchmarking against a real baseline, running or reading A/B tests correctly, and the judgement to know when an offline improvement won't survive contact with production.

TypeScript/Python as your primary languages, with enough command of our surrounding stack (AWS, Terraform, CI/CD) to ship and own your own service's delivery. This is not an infrastructure role, so depth there is not the bar.

Coding agents are part of your daily workflow, and you check their output before you rely on it. You are neither dismissive of them nor careless with them.

Nice to have

Experience with psychometric models, such as Item Response Theory.

Graph ML experience — embeddings, graph neural networks, or relational modeling — at real scale.

Public technical work: open-source contributions, writing, or competitive ML.

Requirements

  • ·Experience with probabilistic modeling, latent-variable modeling and Bayesian inference, or the equivalent rigor from an adjacent domain.
  • ·ML system evaluation: monitoring metrics you define, debugging output that doesn’t look right, and rolling out a change to a live scoring or ranking system without breaking it.
  • ·Rigorous experimentation practice: benchmarking against a real baseline, running or reading A/B tests correctly, and the judgement to know when an offline improvement won't survive contact with production.
  • ·Coding agents are part of your daily workflow, and you check their output before you rely on it. You are neither dismissive of them nor careless with them.
  • ·Experience with psychometric models, such as Item Response Theory.

Benefits

No benefits package published with this listing. Ask about it at first interview.

How to apply

  1. 1Check the flexibility label above, hybrid, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Babbel, mentioning your remote working experience.
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 1d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open data roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

Typically $145k–$210k · You'll be taken to the employer's careers page.