Jobs›Ops Engineer, Noida

ML Ops Engineer

Cloudkeeper · Noida
Pay₹30L–50La year, as listed
WhereNoidaUttar Pradesh
TypeFull time7-12 years
Posted5 Oct6 days ago, via the company site
Kaam checked
No fee, deposit or pay-to-apply signs
Good spoken and written English expected
Day work
No vehicle or licence needed
Skills they list3 named
Graphics Processing Unit (GPU)MLOpsLarge Language Models (LLM)
About this job
Designation
Lead MLOps Engineer (GPU Optimization)

About CloudKeeper

CloudKeeper is a cloud cost optimization partner that combines the power of group buying & commitments management, expert cloud consulting & support, and an enhanced visibility & analytics platform to reduce cloud cost & help businesses maximize the value from AWS, Microsoft Azure, & Google Cloud.

A certified AWS Premier Partner, Azure Technology Consulting Partner, Google Cloud Partner, and FinOps Foundation Premier Member, CloudKeeper has helped 350+ global companies save an average of 20% on their cloud bills, modernize their cloud set-up and maximize value — all while maintaining flexibility and avoiding any long-term commitments or cost.

CloudKeeper hived off from TO THE NEW, a digital technology service company with 2500+ employees and an 8-time GPTW winner.

To know more, please visit - https://www.cloudkeeper.com/

Responsibilities

  • Drive R&D and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform — building the Tuner AI / Commit AI capability on GPU and ML workloads
  • Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement
  • Extend the optimization stack to LLM-era workloads — caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures
  • Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers
  • Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion → optimization engine → customer-facing recommendation)
  • Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales
  • (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands

Must Have

  • B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience
  • Production experience with GPU workloads — has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work
  • Strong performance engineering background — must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction)
  • Strong Python + Linux + systems fundamentals
  • Solid understanding of the ML model lifecycle — training, serving, inference — able to reason about what is running on the GPU and why
  • MLOps fluency — model deployment, monitoring, observability, GPU cluster operations
  • Hands-on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar)
  • Familiarity with at least one modern LLM inference framework — vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML
  • Strong communication skills — able to translate deep technical optimization into customer / business outcomes
  • (Lead level) Experience managing or technically leading a team of 3+ engineers

Good to Have

  • Deep LLM-era optimization expertise — KV caching, semantic caching, model routing, dynamic batching, quantization (FP16 → INT8 → INT4), model distillation, structured outputs
  • Familiarity with LLM workload patterns — RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant)
  • CUDA, NCCL, mixed-precision training and inference
  • Experience with managed ML training platforms — SageMaker, Azure ML, Vertex AI, Databricks Mosaic
  • Exposure to GPU-native clouds — CoreWeave, Lambda Labs, RunPod, Crusoe
  • Open source contributions to ML infrastructure projects — vLLM, llama.cpp, TGI, Ray, Triton, KubeRay
  • Adjacent experience in cloud cost optimization / FinOps — Spot.io, ScaleOps, Granulate, CAST AI
  • Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability)

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on the company site
Opens cutshort.io in a new tab