Descript
Applied Research Scientist, AI Research
Remote role where the employee must remain based in a particular country.
United States only
Employer listed it 4 weeks ago · Added 4 days ago
Been open since 4 weeks ago, still being checked, but it has been live a while.
Salary
$197k to $263k per year
Location
United States only
Timezone
Not stated
Contract
Full-time
Experience
Mid
Category
Data
Stated by the employer in the job description
Remote flexibility
Work from home
This is a remote role, but the employee must be based in United States. It is work from home rather than work from anywhere.
What the employer says
- Source listing states candidate location: "San Francisco, CA or Remote, US, Remote, San Francisco"
- Job description states: "located in the Mission District of San Franci"
What Nomaders makes of it
- Payroll and tax are likely handled in that country only
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
Descript's Research team builds the models behind the product's most distinctive features: Video Regenerate and lipsync, video translation, zero-shot voice and roomtone cloning, and Studio Sound. We don't build general-purpose generative models. We pick specific problems in the editing workflow and build specialized models for them. This isn't research for its own sake. Everything we build is meant to ship, and most of it has, going from prototype to a production feature used by millions of creators within months.
This role is focused on multimodal understanding: training models to perceive edited media the way a human video editor does. Underlord, our AI editing agent, reasons about a project largely through a textual representation of it. Giving it direct perception of the media it's working on is what will let it judge its own output and reason about the creative choices in an edit, not just the structure of a project. It's also an open research problem, since there's no settled way to represent or evaluate editorial craft, whether a cut lands or whether the pacing works. We have a unique dataset to work with.
Some recent work from the team:
Audio editing by latent inpainting : regenerating a masked span of speech
Video Regenerate : regenerating a speaker's lower face to match new or translated audio
Jumpcut Smoothing : generating a bridge across a cut so the join plays like a continuous take
Anchored Tree Sampling : tree-based imputation that bounds drift in long video generation
PoDAR : disentangling power from semantics in audio latents to make them easier to model
More at descript.com/research .
What you'll do
Multimodal understanding: build vision-language systems that let Descript's agentic editing features reason over the visual and audio content of a project.
Evaluation: design the benchmarks and evals that make editorial quality measurable, and that balance quality against cost and latency.
Data: build the datasets your work depends on, including synthetic data generation where real examples don't exist at scale.
Training: train specialized models from scratch or fine-tune existing foundation models, whichever gets the capability we need.
Shipping: take models from prototype to production with the agent and engineering teams.
Direction-setting: identify the next research direction that should become a Descript feature, not just a paper. More senior candidates should expect to own this directly; more junior candidates will grow into it.
Publishing: take your work to academic venues if you'd like. We support it, but it isn't a requirement of the role.
What you bring
Required
Proven ability to design and implement deep learning algorithms, demonstrated by publications, open-source work, or models you've shipped.
Strong programming skills and deep fluency in PyTorch.
A track record of generating new ideas in machine learning. You produce more ideas than you can implement, and once an experiment setup is established, you can run and evaluate many of them quickly rather than being bottlenecked on infrastructure.
Strong experimental judgment. You test ideas fast, and you're honest with yourself and the team about which ones don't pan out.
Clear written and verbal communication, including when a direction isn't working, so the team doesn't waste time following a lead that's already dead.
Requirements
- ·Proven ability to design and implement deep learning algorithms, demonstrated by publications, open-source work, or models you've shipped.
- ·Strong programming skills and deep fluency in PyTorch.
- ·Strong experimental judgment. You test ideas fast, and you're honest with yourself and the team about which ones don't pan out.
- ·Clear written and verbal communication, including when a direction isn't working, so the team doesn't waste time following a lead that's already dead.
- ·A PhD or Master's in deep learning or a related field, or equivalent experience. We care about the track record more than the credential.
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, work from home, matches where you plan to live and work.
- 2Tailor your CV to the role at Descript, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 5d ago. Last checked 23 Sept. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open data roles with comparable remote rules.
Free to apply, no account needed.
$197k to $263k per year · You'll be taken to the employer's careers page.