Luma AI
Senior Site Reliability Engineer
Remote work allowed only within certain countries or regions.
Employer listed it 2 months ago · Added today
Been open since 2 months ago. Long-running listings are sometimes left up after the role is filled.
This listing wasn't found on its job board during the last check. It may have just closed, check the original posting before spending time on an application.
Salary
$168,000–$252,000
Location
Timezone
Not stated
Contract
Full-time
Experience
Senior
Category
Software
Published by the employer
Remote flexibility
Region Restricted
Remote work is allowed, but the listing limits candidates to Redwood City, CA.
What the employer says
- Source listing states candidate location: "Redwood City, CA"
What Nomaders makes of it
- Nomaders could not match this to a wider region
- Check the original posting
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
Team: Infra Reliability · SF Bay Area / Remote (US)
You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference clusters reliable and fast, and you help redesign them for the next level of scale.
This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it.
What You'll Own
Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
Tune Linux performance deeply, at the OS and kernel level.
Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.
First 90 Days
One way the first 90 could unfold.
Days 1–30 — Immerse & Diagnose: Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
Days 30–60 — Ship & Validate: Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
Days 60–90 — Scale & Systemize: Contribute to the next-gen re-architecture and harden security and compliance practices.
What You Bring
5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
Working experience with Terraform, Airflow, and Ray.
Strong experience with AWS or OCI.
Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
Comfort in a less-structured, fast-paced environment.
Nice to Have
Requirements
- ·5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
- ·Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
- ·Working experience with Terraform, Airflow, and Ray.
- ·Strong experience with AWS or OCI.
- ·Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, region restricted, matches where you plan to live and work.
- 2Tailor your CV to the role at Luma AI, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 1d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
$168,000–$252,000 · Free accounts get one application link on us.