OpenAI
CPU/Storage/PoP-WAN Program Manager
Part remote, part office, you need to live within commuting distance of a named location.
Hybrid · San Francisco
Employer listed it 5 months ago · Added 4 days ago
Been open since 5 months ago. Long-running listings are sometimes left up after the role is filled.
Salary
$226,000–$285,000
Location
Hybrid · San Francisco
Timezone
Not stated
Contract
Full-time
Experience
Senior
Category
Operations
Published by the employer
Remote flexibility
Hybrid
This role is only partly remote, the employer expects time in the office around San Francisco, Seattle, Hybrid, so you need to live within commuting distance.
What the employer says
- Source listing states candidate location: "San Francisco, Seattle, Hybrid"
- Listing mentions "Hybrid"
What Nomaders makes of it
- Not suitable if you plan to move between countries
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About the Team
OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical.
The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly.
About the Role
We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity.
In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity.
This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions.
This role is based in San Francisco, CA, with travel as needed.
Key Responsibilities
Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint
Drive readiness to convert contracted compute capacity into schedulable production clusters
Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives
Build integrated schedules spanning procurement, logistics, installation, storage readiness, network turn-up, testing, and production handoff
Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones
Partner with engineering teams to align compute, storage, and networking dependencies before cluster activation
Manage deployment of storage systems supporting training and inference workloads, including readiness, validation, performance checks, and scaling plans
Coordinate backbone capacity expansion, cross-connects, inter-region pathing, and cloud interconnect readiness with Azure and third-party providers
Lead physical deployment execution including rack-and-stack, hardware bring-up, L1 validation, and site acceptance criteria
Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale
Identify risks early across supply chain, site readiness, technical constraints, and vendor execution, then drive mitigation plans
Communicate milestones, escalations, and capacity forecasts to senior leadership
Qualifications
8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations
Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems
Requirements
- ·8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations
- ·Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems
- ·Working knowledge of servers, clusters, storage arrays, routers, switches, optics, and structured cabling
- ·Experience owning cross-functional programs across engineering, operations, supply chain, and external vendors
- ·Strong understanding of deployment lifecycles from planning and procurement through production handoff
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, hybrid, matches where you plan to live and work.
- 2Tailor your CV to the role at OpenAI, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 5d ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open operations roles with comparable remote rules.
Free to apply, no account needed.
$226,000–$285,000 · You'll be taken to the employer's careers page.