Moonlite logo

Moonlite

Senior Software Engineer, Platform Infrastructure

Work from home

Remote role where the employee must remain based in a particular country.

United States only

Employer listed it 4 months ago · Found 8h ago

Been open since 4 months ago. Long-running listings are sometimes left up after the role is filled.

Salary

$165,000 to $225,000

Location

United States only

Timezone

Not stated

Contract

Full-time

Experience

Senior

Category

Software

Stated by the employer in the job description

Remote flexibility

Work from home

This is a remote role, but the employee must be based in United States. It is work from home rather than work from anywhere.

What the employer says

  • Source listing states candidate location: "Chicago, IL or Remote, United States (Remote)"
  • Job description states: "located in yours"

What Nomaders makes of it

  • Payroll and tax are likely handled in that country only

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads.We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists combines bare-metal performance with cloud-native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise-grade reliability and compliance.

Your Role:

You will be foundational to building the comprehensive infrastructure platform that bridges our physical infrastructure – bare-metal servers, GPU clusters, high-performance storage, and networking fabric – with the systems our customers depend on for large-scale computation, inference, simulations, and training. Working closely with product, your platform team members, and infrastructure specialists, you’ll design and implement the orchestration layer, APIs, and automation framework that make thousands of servers, petabytes of storage, and high-speed networks feel like a unified, programmable platform.

Job Responsibilities

Infrastructure Abstraction Layer: Design and build systems that bridge physical infrastructure (bare-metal servers, storage clusters, network fabric) with customer-facing services, enabling programmatic management of compute, networking, and storage at scale.

Research Cluster Provisioning: Design and implement systems for provisioning and managing research computing environments including Kubernetes and SLURM clusters, enabling automated deployment, resource scheduling, and workload orchestration for distributed AI training and HPC workloads.

Platform Orchestration: Implement comprehensive orchestration systems that coordinate across compute, storage and networking to deliver unified experience for complex research workloads.

Network Automation & Placement: Design and build network provisioning automation including intelligent VM placement decisions for optimal network topology, automated VLAN and subnet configuration, and software-designed networking orchestration for high-performance interconnects.

Enterprise APIs & SDKs: Develop robust APIs and SDKs that enable researchers and engineering teams to programmatically provision and manage infrastructure resources across all platform domains.

Observability & Telemetry: Implement comprehensive observability, telemetry, and logging systems that provide visibility into infrastructure health, performance, and utilization across the infrastructure footprint.

Performance Engineering: Build and optimize platform services that deliver consistent high-throughput low-latency networking for demand research applications and data-intensive workloads.

Cross-Team Collaboration: Work closely with engineering, infrastructure, and product to define requirements, drive infrastructure-product-rollouts, and improve resource lifecycle management.

Compliance & Security: Implement platform-wide compliance and security features supporting SOC 2, ISO 27001, and enterprise regulatory requirements including comprehensive audit logging, access controls, and data residency management.

Requirements

Experience: 5+ years in software engineering with a proven track record of infrastructure platforms, distributed systems, or cloud platforms for production environments.

Kubernetes & Container Orchestration: Strong familiarity with Kubernetes architecture, container orchestration concepts, and experience deploying workloads in Kubernetes environments. Understanding of pods, deployments, services, and basic Kubernetes operations.

Infrastructure Systems: Strong understanding of infrastructure fundamentals including compute orchestration, storage systems, networking technologies, and how they integrate together to deliver complete platform experiences.

Programming Skills: Experience with systems programming languages (Go, C/C++, Rust, Python) for performance-critical components is a strong plus.

Linux Production Experience: Strong experience with linux in production environments, including systems administration, performance tuning, and troubleshooting.

Bare-Metal & Virtualization: Deep knowledge of bare-metal infrastructure, provisioning systems, out-of-band management, and virtualization technologies (KVM, Kubernetes, etc).

API & Platform Design: Proven experience designing and building APIs, SDKs, and automation frameworks that enable programmatic infrastructure management.

Cloud Platform Knowledge: Strong familiarity with cloud environments (AWS, GCP, Azure) and understanding of how to translate cloud-native patterns to bare-metal infrastructure.

Infrastructure Automation: Experience with Infrastructure-as-code tools (Terraform, Ansible) and building automated deployment pipelines.

Problem Solving & Autonomy: Self-starter who can navigate ambiguity, balance pragmatic shipping with good long-term architecture, and independently drive complex technical initiatives.

Requirements

  • ·Experience: 5+ years in software engineering with a proven track record of infrastructure platforms, distributed systems, or cloud platforms for production environments.
  • ·Programming Skills: Experience with systems programming languages (Go, C/C++, Rust, Python) for performance-critical components is a strong plus.
  • ·Linux Production Experience: Strong experience with linux in production environments, including systems administration, performance tuning, and troubleshooting.
  • ·Bare-Metal & Virtualization: Deep knowledge of bare-metal infrastructure, provisioning systems, out-of-band management, and virtualization technologies (KVM, Kubernetes, etc).
  • ·API & Platform Design: Proven experience designing and building APIs, SDKs, and automation frameworks that enable programmatic infrastructure management.

Benefits

No benefits package published with this listing. Ask about it at first interview.

How to apply

  1. 1Check the flexibility label above, work from home, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Moonlite, mentioning your remote working experience.
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 9h ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open software roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

$165,000 to $225,000 · You'll be taken to the employer's careers page.