Fireworks AI
Member of Technical Staff, Cloud Infrastructure
Part remote, part office, you need to live within commuting distance of a named location.
Hybrid · San Mateo
Employer listed it 15 months ago · Added today
Been open since 15 months ago. Long-running listings are sometimes left up after the role is filled.
Salary
$175,000–$220,000
Location
Hybrid · San Mateo
Timezone
Not stated
Contract
Full-time
Experience
Lead
Category
Software
Published by the employer
Remote flexibility
Hybrid
This role is only partly remote, the employer expects time in the office around San Mateo, New York, Hybrid, so you need to live within commuting distance.
What the employer says
- Source listing states candidate location: "San Mateo, New York, Hybrid"
- Listing mentions "Hybrid"
What Nomaders makes of it
- Not suitable if you plan to move between countries
The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.
About the role
About Us:
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
The Role:
As a Software Engineer on our Cloud Infrastructure team, you'll be at the forefront, architecting and building the foundational systems that power Fireworks AI's revolutionary generative AI platform. You'll spearhead the creation of one of the world's first virtual clouds, seamlessly serving AI workloads across the globe and every cloud provider. Your mission: to deliver unparalleled reliability, efficiency, and scalability, fueling the world's most innovative AI products.This is a highly technical role requiring deep expertise in distributed systems, cloud-native infrastructure, and machine learning platforms. You’ll partner closely with engineering partners, product teams, and infrastructure stakeholders to design solutions that balance performance, cost-efficiency, and operational simplicity across compute, storage, and networking layers.
Key Responsibilities:
Architect and build scalable, resilient, and high-performance backend infrastructure to support distributed training, inference, and data processing pipelines.
Lead technical design discussions, mentor other engineers, and establish best practices for building and operating large-scale ML infrastructure.
Design and implement core backend services (e.g., job schedulers, resource managers, autoscalers, model serving layers) with a focus on efficiency and low latency.
Drive infrastructure optimization initiatives, including compute cost reduction, storage lifecycle management, and network performance tuning.
Collaborate cross-functionally with ML, DevOps, and product teams to translate research and product needs into robust infrastructure solutions.
Continuously evaluate and integrate cloud-native and open-source technologies (e.g., Kubernetes, Kubeflow, MLFlow) to enhance our platform’s capabilities and reliability.
Own end-to-end systems from design to deployment and observability, with a strong emphasis on reliability, fault tolerance, and operational excellence.
Ensuring System Reliability: Ensure systems are designed and implemented with high availability, scalability, and performance. Focus on fault tolerance, disaster recovery, identifying and removing scaling bottlenecks, and performance optimization across our multi-cloud infrastructure.
Observability & Monitoring: Develop, implement, and maintain comprehensive monitoring, alerting, logging, and tracing solutions to provide deep insights into system health and performance
Minimum qualifications:
Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
5+ years of experience designing and building backend infrastructure in cloud environments (e.g., AWS, GCP, Azure).
Proven experience in ML infrastructure and tooling (e.g., PyTorch, TensorFlow, Vertex AI, SageMaker, Kubernetes, etc.).
Strong software development skills in languages like Python, or C++.
Deep understanding of distributed systems fundamentals: scheduling, orchestration, storage, networking, and compute optimization.
Preferred qualifications:
Master’s or PhD in Computer Science or related field.
Experience leading infrastructure projects supporting large-scale ML/AI workloads or high-throughput systems.
Familiarity with infrastructure-as-code and CI/CD tooling (e.g., Terraform, ArgoCD, GitOps).
Requirements
- ·Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
- ·5+ years of experience designing and building backend infrastructure in cloud environments (e.g., AWS, GCP, Azure).
- ·Proven experience in ML infrastructure and tooling (e.g., PyTorch, TensorFlow, Vertex AI, SageMaker, Kubernetes, etc.).
- ·Strong software development skills in languages like Python, or C++.
- ·Deep understanding of distributed systems fundamentals: scheduling, orchestration, storage, networking, and compute optimization.
Benefits
No benefits package published with this listing. Ask about it at first interview.
How to apply
- 1Check the flexibility label above, hybrid, matches where you plan to live and work.
- 2Tailor your CV to the role at Fireworks AI, mentioning your remote working experience.
- 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.
Found 23h ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.
Listing sourced from Company boards.
Similar roles
Other open software roles with comparable remote rules.
Free to apply, no account needed.
$175,000–$220,000 · You'll be taken to the employer's careers page.