Speak logo

Speak

Assessment Design Lead

Work from home

Remote role where the employee must remain based in a particular country.

United States only

Employer listed it 6 weeks ago · Added yesterday

Been open since 6 weeks ago. Long-running listings are sometimes left up after the role is filled.

Salary

$100,000 to $140,000

Location

United States only

Work style

Async

Contract

Full-time

Experience

Lead

Category

Education

Published by the employer

Remote flexibility

Work from home

This is a remote role, but the employee must be based in United States. It is work from home rather than work from anywhere.

What the employer says

  • Source listing states candidate location: "US - Remote, Remote"

What Nomaders makes of it

  • Residency required in United States
  • Payroll and tax are likely handled in that country only

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

About us

Our mission is to reinvent the way people learn, starting with language.

Learning a language can change a life by opening doors to new cultures, careers, and communities. Two billion people around the world are actively trying to learn a language, but the best way to learn (one-on-one tutoring) is hard to access at scale and hasn’t been meaningfully improved in decades. Speak is building a human-level, AI-powered tutor in your pocket: a conversation-first experience that lets learners actually speak, get instant feedback, and progress through carefully designed lessons. The result is a complete path from beginner to confident speaker across multiple languages.

Speak first launched in South Korea in 2019, where Speak has now become the number one language learning app, and we now serve learners across many markets and 15+ languages. Speak is one of the world’s leading AI companies, with over $150m raised in venture investment from OpenAI, Accel, Founders Fund, Khosla Ventures, and more, with a distributed team across San Francisco, Seoul, Tokyo, Taipei, and Ljubljana.

About this role

Speak cares deeply about learners actually learning and improving with Speak. We have a dedicated Proficiency team to own how we measure learning efficacy and speaking proficiency, from unit-level mastery checks to standalone proficiency tests to onboarding placement, all in service of understanding users’ proficiency levels and learning gains in an accurate, transparent and actionable manner. Because everything happens remotely and asynchronously in the app, keeping scores fair, stable over time, and resistant to gaming is both hard and genuinely interesting.

We're looking for an Assessment Design Lead: the person who defines what we measure, why, how, and signs off on whether our assessments actually measure the right thing. You will be staffed on the Proficiency team and report to the Head of Learning Design and Curriculum. If solving reliable, at-scale speaking assessment excites you and you want to directly shape Speak's efficacy story, we'd love to hear from you.

What you'll be doing

Own Assessment Design — Define what Speak measures, why, and how often across three distinct assessment types (Curriculum Mastery Assessment, Proficiency Test, Placement Test) — with the Proficiency Test as the immediate focus, expanding to the other two as the pod's priorities evolve. Keep constructs (fluency, pronunciation, grammar, task achievement) clearly separated and each aligned to CEFR or a comparable speaking proficiency standard, so no single assessment conflates domains it wasn't designed to measure.

Define Constructs & Build the Rubric/Blueprint Layer — Translate fuzzy goals like "measure fluency" or "measure pronunciation" into concrete, scoreable constructs, item blueprints, and rubrics that an item writer can generate items against and an ML Engineer can build a grading model against.

Own Validity & the Quality Bar — Sign off on content validity for every assessment that ships. Decide what "mastery" or a passing score operationally means, catch cases where an assessment is measuring the wrong thing before it ships, own the rubric/rater guidelines behind the human-labeled data our ML scoring models are evaluated against, and audit items/rubrics for bias across learner subgroups. Making sure scores stay comparable as the assessment evolves and stay meaningful against attempts to game an unproctored test.

Design and run the validity evidence plan — so validity is built into the process rather than checked only after launch. This includes concurrent/criterion studies benchmarking Speak’s assessments against external proficiency measures (CEFR-anchored exams, expert human ratings), so we can say what a Speak Score means in terms the outside world already trusts.

Partner Tightly with Product and ML — Work closely with the Product Manager and ML Engineers on automated scoring, calibration, and feedback generation. You own the construct and quality bar, they own the model. Neither works without the other, and the loop between you is the product.

What we're looking for

Must-Haves

Assessment/Psychometric Design: 4+ years designing rubrics, blueprints, and item specs for a real, shipped language assessment product (or equivalent depth in closely related psychometric/measurement work) — not just academic theory. Can explain reliability and validity in plain language and knows how to catch a test that's measuring the wrong construct.

Language Proficiency Domain Expertise: Deep familiarity with frameworks like CEFR (or ACTFL, IELTS/TOEFL band descriptors) and what separates "did you learn what we taught you" from "how good is your speaking overall."

Fairness Across Learner Populations: Can identify whether an item or rubric unfairly penalizes specific L1 backgrounds or accents (differential item functioning) — essential for a speech-based test serving learners across dozens of native languages.

Translates Qualitative → Technical: Can turn a construct like "pronunciation quality" into something concrete enough for an ML engineer to build a scoring pipeline against, without either oversimplifying or getting lost in academic nuance.

Quantitative Rigor: Comfortable running or interpreting the statistics behind a rubric or rater system — inter-rater reliability (e.g., Cohen's/Fleiss' kappa), classical test theory, and basic IRT concepts — enough to know whether a scoring system is actually reliable, not just plausible.

Ownership of Quality Bar: Comfortable being the sign-off authority on content validity — makes the call clearly and follows through on it, rather than deferring to data alone or product pressure to ship.

AI Fluency & Judgment: Uses AI tools directly in their own workflow (e.g., drafting item variants, testing rubric language, exploring construct definitions) and has real judgment about when AI-generated output is precise enough to ship vs. needs a human rewrite — distinct from spec'ing work for the ML Engineer to build.

Comfort with Ambiguity: Comfortable operating in a 0-to-1 environment. Can wear multiple hats, take a fuzzy goal and turn it into a concrete plan, communicate tradeoffs clearly, and keep momentum without waiting for perfect clarity or team setup

Nice-to-Haves

Requirements

The employer hasn't listed requirements separately, they're described in the role summary above and on the original listing.

Benefits

No benefits package published with this listing. Ask about it at first interview.

How to apply

  1. 1Check the flexibility label above, work from home, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Speak, mentioning your remote working experience and working hours (Async).
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 1d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open education roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

$100,000 to $140,000 · You'll be taken to the employer's careers page.