Jobs›Platform

Senior Platform & Site Reliability Engineer

HyrHub · Work from home
SeniorWork from home
Pay₹10L–30La year, as listed
WhereWork from home
TypeFull time8-12 years
Posted21 Sep16 days ago, via the company site
Skills they list4 named
TerraformgrafanaprometheusAmazon Web Services (AWS)
About this job

What You Will Own

Platform Architecture

Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure

Define, build, and enforce platform standards across portfolio products

Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed

Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines

Design and maintain batch processing infrastructure alongside live event flows

Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products

Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures

Deployment standards
release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

Own the full observability stack: Grafana, Prometheus, and Loki across all products

SLOs and error budgets defined per product; reliability tracked consistently

Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen

Incident response
on-call design, escalation playbooks, post-mortem facilitation

Automated remediation scoped to a defined set of safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns. Novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture

Migration plan and execution for each portfolio company joining the platform.

Target
full platform integration within a defined window per acquisition

What We're Looking For

Experience & Background

8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time

Strong Terraform depth across multi-environment, multi-account setups

CI/CD ownership across a multi-product environment with GitHub Actions

Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management

Hands-on Grafana, Prometheus, and Loki in production

AWS operational depth
ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
SRE fundamentals
SLOs, error budgets, on-call design, post-mortem culture

Acquisition or greenfield platform integration experience strongly preferred

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on the company site
Opens cutshort.io in a new tab