
Senior Platform & Site Reliability Engineer
What You Will Own
Platform Architecture
Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
Define, build, and enforce platform standards across portfolio products
Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
Self-service developer platform so product teams ship without waiting on platform
Event Streaming & Pipeline Infrastructure
Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
Design and maintain batch processing infrastructure alongside live event flows
Ensure pipeline reliability, throughput, and cost are actively managed at scale
CI/CD & Deployment
Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards
- release management, rollback mechanisms, canary and blue-green patterns where justified
Observability & Reliability
Own the full observability stack: Grafana, Prometheus, and Loki across all products
SLOs and error budgets defined per product; reliability tracked consistently
Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
- Incident response
- on-call design, escalation playbooks, post-mortem facilitation
Automated remediation scoped to a defined set of safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns. Novel or ambiguous failures escalate to a human with full context attached
Acquisition Onboarding
Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
Migration plan and execution for each portfolio company joining the platform.
- Target
- full platform integration within a defined window per acquisition
What We're Looking For
Experience & Background
8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
Strong Terraform depth across multi-environment, multi-account setups
CI/CD ownership across a multi-product environment with GitHub Actions
Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth
- ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals
- SLOs, error budgets, on-call design, post-mortem culture
Acquisition or greenfield platform integration experience strongly preferred
Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →