Lambda
Staff Software Engineer - Compute
Part remote, part office, you need to live within commuting distance of a named location.
Hybrid · Bellevue Office
Employer listed it 6 weeks ago · Added 5 days ago
Been open since 6 weeks ago, still being checked, but it has been live a while.
Salary
$314,000 to $465,000
Location
Hybrid · Bellevue Office
Timezone
Not stated
Contract
Full-time
Experience
Lead
Category
Software
Published by the employer
Remote flexibility
Hybrid
This role is only partly remote, the employer expects time in the office around Bellevue Office, San Jose Office (First St), San Francisco Office (Fremont St), Hybrid, so you need to live within commuting distance.
What the employer says
- Source listing states candidate location: "Bellevue Office, San Jose Office (First St), San Francisco Office (Fremont St), Hybrid"
- Listing mentions "Hybrid" and 4 days per week in the office
What Nomaders makes of it
- Not suitable if you plan to move between countries
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
About the Role
As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration. The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.
What You'll Do
We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:
Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.
Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
Guide design of compute platform multi-tenant security model
Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.
Work with customers on translating vague customer technical requirements into concrete engineering deliverables.
Set engineering standards and lead design reviews for mission-critical cloud software at scale.
Who You are
10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.
Nice to Have
Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
Requirements
The employer hasn't listed requirements separately, they're described in the role summary above and on the original listing.
Benefits
- ·Health, dental, and vision coverage for you and your dependents
- ·Wellness and commuter stipends for select roles
- ·401k Plan with 2% company match (USA employees)
- ·Flexible paid time off plan that we all actually use
- ·Equal Opportunity Employer
How to apply
- 1Check the flexibility label above, hybrid, matches where you plan to live and work.
- 2Tailor your CV to the role at Lambda, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 5d ago. Last checked 23 Sept. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
$314,000 to $465,000 · You'll be taken to the employer's careers page.