ASAI job platform logo
  1. Home
  2. /Jobs
  3. /Cloud Engineer Jobs in Bengaluru
  4. /Infrastructure SRE - HPC
sarvam

Infrastructure SRE - HPC

Bengaluru · Senior

Older listing - lower visibility likelyVerified listingPosted 101d ago

Applicants who checked fit first are 3.1× more likely to hear back

Your match scoreCalculated · locked
86Overall
64Skills
97Experience

Your score for this role already exists

ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.

No credit card · 1 tap with Google

What we know about this role

Hiring pulse

LOW

sarvam is showing limited hiring activity lately. Expect slight delayed response.

Apply window

First 72 hours

Window passed - posted 101d ago

Early applicants get seen before the pile builds.

Not a repost

The first time we've seen this listing - it hasn't been closed and reopened.

Skills required

7 listed
Graphics Processing Unit (GPU)KubernetesPython (Programming Language)Slurm (Batch Scheduling Software)InfiniBandMetal Inert Gas (MIG) WeldingFabric.io
Graphics Processing Unit (GPU)KubernetesPython (Programming Language)Slurm (Batch Scheduling Software)InfiniBand

You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.

Job description

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

Sarvam runs a large, multi-vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs and must run uninterrupted for weeks, and inference services that must hold a flat p99 under production load. Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve.

This is not a Kubernetes administration role. We assume Kubernetes fluency as a baseline. The difficulty lies above and below it - in parallel filesystems under heavy checkpoint load, in RDMA fabrics that degrade quietly, in NCCL hangs whose root cause may be the network or the kernel, in driver and firmware drift across heterogeneous hardware, and in distributed training failures that masquerade as infrastructure faults.

We are hiring a team of specialists rather than a set of identical generalists. This posting covers five areas of focus. We expect candidates to bring genuine depth in one and working fluency across the others, because on a shared fleet a storage problem often first appears as a training hang, and the engineer on call must route an incident correctly before anyone can resolve it.

When you apply, please indicate the area of focus that best matches your experience. Strong generalists are welcome; we will place you where your depth is most useful.

What You’ll Do

  • Operate the GPU fleet end to end across training and serving - provisioning, observability, capacity, and fleet health.

  • Hold a meaningful on-call rotation, write runbooks that hold up under pressure, and drive postmortems that produce durable fixes.

  • Build the internal tooling the team relies on, rather than operating off-the-shelf systems alone.

  • Partner with ML and platform teams to keep large runs alive and serving latency predictable.

What We're Looking For

  • 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.*

  • Demonstrated on-call ownership of infrastructure that mattered, with a track record of postmortems that led to real change.

  • Proficiency in Python or Go, used to build and maintain internal tooling.

  • Working fluency across all five areas of focus below - enough to recognize, triage, and route a problem outside your specialty, even if the fix belongs to a teammate.

* For the Storage and Fabric areas of focus, we will weigh deep domain expertise against the GPU-cluster requirement; exceptional specialists with less direct GPU-fleet time are encouraged to apply.

Bring depth in one of the five areas below; expect to be conversational across the rest.

  1. Distributed high-performance storage — operate a parallel filesystem (Lustre, GPFS, WEKA, or BeeGFS) at scale and keep it from stalling under checkpoint-write storms.

  2. Fabric & RDMA networking — InfiniBand or RoCE health, NVLink/NVSwitch topology, and RDMA debugging that catches degradation before the workload feels it.

  3. GPU systems reliability — NCCL debugging, driver and firmware lifecycle across a mixed fleet, and DCGM-based node health at scale.

  4. Kubernetes platform reliability — the GPU operator stack, scheduling (Slurm-on-k8s or pure k8s), multi-tenant isolation, and cost/SLO primitives.

  5. Training & inference workload reliability — hang and straggler detection, checkpoint/restart, and protecting serving p99 with HA and DR.

Bonus Points

  • Slurm and Kubernetes hybrid environments.

  • On-premise GPU deployment, including coordination with datacenter operations on power, cooling, and InfiniBand cabling.

  • Experience with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud etc.

  • Multi-tenant GPU isolation (MIG, MPS, time-slicing) in production.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.

Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

High ownership and high impact, from day one

Everything we do is AI-first, from the way we build and ship to the way we think about problems

You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

 
 
 
 

Free · no signup

Get tomorrow's jobs before you have to search

Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.

Join WhatsAppJoin Telegram

No spam. Just jobs and resources.

Why people use ASAI

Someone shared one job with you. ASAI keeps finding the rest.

  • Scored, not searched. Every role ranked against your actual profile.

  • Alerts as often as hourly. Reach new roles while the pile is still small.

  • Skill gaps, spelled out. See exactly which requirements you don't meet yet.

  • Verified jobs, only. Say no to ghost jobs. Your time deserves respect.

More Cloud Engineer roles in Bengaluru

See all

Staff Engineer - Software Development (Cloud Networking & Network Security)

Aviatrix · Bengaluru

Assistant Manager - Cloud & AI Architect

KPMG India · Bengaluru

Assistant Manager - Cloud & AI Architect

KPMG India · Bengaluru

Sr Analyst I Infrastructure Architecture

Dxctechnology · Bengaluru

Keep browsing

All open roles at sarvamAll Cloud Engineer jobs in Bengaluru

Two ways in

Applicants who checked fit first are 3.1× more likely to hear back

Your match scoreCalculated · locked
86Overall
64Skills
97Experience

Your score for this role already exists

ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.

No credit card · 1 tap with Google

Free · no signup

Get tomorrow's jobs before you have to search

Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.

Join WhatsAppJoin Telegram

No spam. Just jobs and resources.

Why people use ASAI

Someone shared one job with you. ASAI keeps finding the rest.

  • Scored, not searched. Every role ranked against your actual profile.

  • Alerts as often as hourly. Reach new roles while the pile is still small.

  • Skill gaps, spelled out. See exactly which requirements you don't meet yet.

  • Verified jobs, only. Say no to ghost jobs. Your time deserves respect.

ASAI job platform logo

A job platform finally, balanced in your favour.

Jobs by CityJobs by CompanyGuidesHow We VerifyAboutPrivacy PolicyTerms of Service

Built in India 🇮🇳

© 2026 ASAI. All rights reserved.