SF Compute
HPC/ GPU Cluster Architect
Part remote, part office, you need to live within commuting distance of a named location.
Hybrid · San Francisco, CA
Employer listed it 9 days ago · Added yesterday
First listed 9 days ago and still open.
Salary
$220,000–$300,000
Location
Hybrid · San Francisco, CA
Timezone
Not stated
Contract
Full-time
Experience
Mid
Category
Software
Published by the employer
Remote flexibility
Hybrid
This role is only partly remote, the employer expects time in the office around San Francisco, CA, Remote, Hybrid, so you need to live within commuting distance.
What the employer says
- Source listing states candidate location: "San Francisco, CA, Remote, Hybrid"
- Listing mentions "Hybrid"
What Nomaders makes of it
- Not suitable if you plan to move between countries
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About San Francisco Compute
San Francisco Compute exists to give ambitious teams large scale supercomputers without taking on a catastrophic balance sheet risk. Most GPU clouds force you to sign long term contracts with no way out. If you underbuy, you miss the window. If you overbuy, you’ll blow up with no way out. It’s a desperate gamble that San Francisco Compute does not force you to play.
Originally, SFC was an AI lab that bought too big of a GPU cluster and was forced to sublease it or go under. Today, we build data centers, supercomputers, and then lease those machines on contracts that let our customers sublease. More than anyone else in the industry, we deeply felt the pain and fear of being forced to bet the farm with your neocloud vendor. To make it possible to sublease, we invented the first compute market and became the first one to offer physical settlement. That means we operate GPU clusters like a neocloud and derisk our customers like a market.
Customers buying GPU clusters don’t want to work through a network of brokers. They want to transact with the folks that have the keys to the cluster and can own the SLA.
To keep the industry safe from unnecessary risk, SFC needs to scale fast. To do that, we serve financially motivated parties as a technical partner to build, operate, & lease supercomputers on their behalf. Basically, we help folks who want an economic return build GPU clusters on their balance sheet, but operated in full by us. That allows us to retain the technical control needed to operate a cloud, but scale fast like a market. SFC’s ability to offer subleasing can significantly increase the levered returns for cluster owners. Our hybrid model out-performs brokered contracts or other “compute markets”, who have to go through a network of counterparties to solve problems. It also lets us design custom solutions for our customers, like clusters deployed in regions next to your dataset or physical eval set or a large CPU fleet deployed colocated with your GPU cluster.
Our team includes senior & technical leadership from places like Lambda, Crusoe, Digital Ocean, AWS, and Hut8. Our CTO is the cofounder and former CEO of Voltage Park. In prior roles, our team has deployed 8GW of datacenter capacity & hundreds of thousands of GPUs. We’ve been described as having “the highest talent density in the space.” If you are an honorable, gritty person with eyes wide open & good epistemics, who wants to see AI go right, we’d love to work with you. There has never been a more critical moment in history and it is up to us to shape it on behalf of those who come after us.
About the Role
GPU clusters are some of the most performant computers on the planet. Even smaller clusters by today’s standards would have ranked in the TOP500 five years ago. Our infrastructure team is responsible for architecting and deploying new clusters around the world and keeping them running smoothly. You’ll participate in on-call rotation, deploy new environments, fix issues when they arise, and lean into automation to enable deployments at scale. We’re a small but ambitious team so you’ll be an early contributor helping to shape culture, mentor junior engineers, and learn from our customers.
About You
You will have 5+ years of experience with hands-on designing, architecting and scaling at least one HPC or GPU compute cluster in production (ideally >1,000 GPUs, but not required)
You deeply understand server hardware fundamentals, including GPUs, NICs, PCIe, memory, thermals, and power
You’re comfortable debugging performance and reliability issues across hardware, OS, drivers, and networking layers; full-stack.
The idea of automating fleet operations (provisioning, monitoring, remediation) excites you — you embrace infrastructure-as-code
You appreciate, value, and generate strong operational documentation and runbooks
You have the ability and willingness to mentor junior engineers and contribute to team culture
You’re open to domestic travel when required
Some Nice to Haves
Familiarity with data center operations including power, cooling, and colo/vendor engagements
Strong Linux systems administration experience, including kernel drivers, RDMA stack tuning, and performance analysis
Experience with schedulers and orchestration systems such as Slurm and Kubernetes
Exposure to virtualization technologies (KVM, QEMU, libvirt)
Experience utilizing telemetry pipelines for predictive hardware failure detection
Experience troubleshooting high-speed fabrics such as InfiniBand and/or RoCEv2 Ethernet
Benefits
Requirements
- ·You will have 5+ years of experience with hands-on designing, architecting and scaling at least one HPC or GPU compute cluster in production (ideally >1,000 GPUs, but not required)
- ·You deeply understand server hardware fundamentals, including GPUs, NICs, PCIe, memory, thermals, and power
- ·You’re comfortable debugging performance and reliability issues across hardware, OS, drivers, and networking layers; full-stack.
- ·The idea of automating fleet operations (provisioning, monitoring, remediation) excites you — you embrace infrastructure-as-code
- ·You appreciate, value, and generate strong operational documentation and runbooks
Benefits
- ·Generous equity grant
- ·Team members are offered a competitive salary along with equity in the company
- ·Visa Sponsorships
- ·Yes, we sponsor visas and work permits
- ·Retirement matching
How to apply
- 1Check the flexibility label above, hybrid, matches where you plan to live and work.
- 2Tailor your CV to the role at SF Compute, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 1d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
$220,000–$300,000 · You'll be taken to the employer's careers page.