Mirantis logo

Mirantis

Senior Site Reliability Engineer

Work from home

Remote role where the employee must remain based in a particular country.

India only

Employer listed it 4 days ago · Added today

First listed 4 days ago and still open.

Salary

Not stated

Location

India only

Timezone

APAC

Contract

Full-time

Experience

Senior

Category

Software

This employer didn't state pay. Jobs like this usually pay around $170k–$225k a year, a typical range taken from 597 senior-level software roles on Nomaders that do state pay. It's a guide, not an offer.

Remote flexibility

Work from home

This is a remote role, but the employee must be based in India. It is work from home rather than work from anywhere.

What the employer says

  • Source listing states candidate location: "Hyderabad, , India, Hyderabad, in, Remote"

What Nomaders makes of it

  • Residency required in India
  • Payroll and tax are likely handled in that country only

The quotes above are the employer's own words; the reading is ours. Always check the original listing and employment terms before working from another country.

About the role

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

We are looking for a senior Kubernetes-focused DevOps/SRE engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments, pipelines, and infrastructure tooling that our engineering teams across the US, Europe, and APAC  depend on to ship daily.

This role spans both sides of the line. You will make our development, test, and pre-production clusters fast and reproducible, harden the Helm and CI/CD path from commit to release, and carry operational ownership — including on-call — for one of our smaller customer-facing production regions, under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated — the person other teams consult, including the teams running our larger regions. Working within an agile framework, you will directly shape how quickly and safely changes reach production, and be accountable for how they behave once there.

Main Responsibilities:

Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing production region — local kind clusters, shared dev and QA environments, and multi-cluster/multi-region topologies.

Operate your production region against defined SLOs: capacity and upgrade planning, patching, backup and restore, disaster recovery drills, and participation in an on-call rotation.

Lead incident response for your region — detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems that feed fixes back into the platform.

Build and maintain Helm charts and umbrella releases for the platform's services and dependencies, including versioning, values hygiene, and upgrade paths.

Own the CI/CD pipelines end to end — build, test, image publishing, chart packaging, release cutting, and hotfix/backport flows.

Automate environment bootstrap and seeding so any engineer can bring up a full stack — control plane, identity, gateway, database, workflow engine — with one command.

Operate and troubleshoot the supporting stack across test and production: PostgreSQL, Temporal, Keycloak, API gateway, message broker, and observability components.

Build observability and diagnostics — metrics, dashboards, alerting, log and audit access — that serve both engineering environments and production operations.

Consult with product teams and with the teams operating our larger regions on deployment topology, GPU and resource scheduling, RBAC, networking, and failure modes; validate upgrade and migration procedures and hand over runbooks.

Enforce security and tenant isolation in production: least-privilege access, secret handling, certificate and TLS lifecycle, image and dependency scanning, and audit evidence for compliance reviews.

Drive infrastructure as code and repeatability — no snowflake environments, no undocumented manual steps.

Mentor engineers on Kubernetes and operational practice, and raise the team's bar through review and documentation.

Required Skills/Abilities: 

10+ years in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.

Expert-level Kubernetes: workloads, networking, storage, RBAC, resource management, CRDs and operators, and cluster upgrades — able to debug from kubectl and cluster internals rather than dashboards alone.

Proven incident response under SLA pressure — on-call rotations, escalation paths, postmortems, and follow-through on corrective action.

Strong CI/CD engineering — pipelines as code, reproducible builds, artifact and release management (GitHub Actions or equivalent).

Solid scripting and automation ability, and enough Go familiarity to read service code, trace a failure into it, and file a precise bug.

Experience running the stateful supporting stack — relational databases, identity providers, gateways, and message brokers — in Kubernetes, including backup, restore, and upgrade.

Track record as a technical consultant to other engineering teams: clear runbooks, design feedback, and incident write-ups across global time zones (strong written English).

Requirements

  • ·10+ years in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.
  • ·Expert-level Kubernetes: workloads, networking, storage, RBAC, resource management, CRDs and operators, and cluster upgrades — able to debug from kubectl and cluster internals rather than dashboards alone.
  • ·Proven incident response under SLA pressure — on-call rotations, escalation paths, postmortems, and follow-through on corrective action.
  • ·Strong CI/CD engineering — pipelines as code, reproducible builds, artifact and release management (GitHub Actions or equivalent).
  • ·Solid scripting and automation ability, and enough Go familiarity to read service code, trace a failure into it, and file a precise bug.

Benefits

  • ·We are a  Leader for Container Management  in G2 (#2 after AWS)!

How to apply

  1. 1Check the flexibility label above, work from home, matches where you plan to live and work.
  2. 2Tailor your CV to the role at Mirantis, mentioning your remote working experience and working hours (APAC).
  3. 3Apply directly on the employer's careers page using the button below. Nomaders never handles your application.

Found 15h ago. Last checked today. Always confirm the details on the original posting, salary and location can change after publication.

Listing sourced from Company boards.

Similar roles

Other open software roles with comparable remote rules.

Browse all open roles

Free to apply, no account needed.

Typically $170k to $225k per year · You'll be taken to the employer's careers page.