Jobs›Staff Engineer

Staff Engineer - Distributed Systems

HighLevel · Work from home
Work from home
PayPay not listed
WhereWork from home
TypeFull timeSenior
Posted2 Oct3 days ago, via Himalayas
Kaam checked
No fee, deposit or pay-to-apply signs
Day work
No vehicle or licence needed
Skills they list7 named
Staff EngineerSoftware EngineerPrincipal EngineerSenior Staff EngineerDistributed Systems EngineerRemote Senior Distributed Systems EngineerStaff Systems Engineer
About this job

About HighLevel:HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. Today, HighLevel supports SMBs across 150+ countries, fueling community-driven growth rooted in real customer outcomes.

To date, businesses operating on HighLevel have generated over $7 billion in ecosystem value, demonstrating the impact of shared infrastructure at scale. By centralizing conversations, automation and intelligence into one system, we help businesses move faster, reduce complexity and execute efficiently.

Behind the platform, HighLevel powers more than 4 billion API hits and 2.5 billion message events daily. With 250 terabytes of distributed data, 250+ microservices and over 1 million domain names supported, our architecture is built for performance, resilience and long-term scalability.

Our people

With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership. We value initiative, clarity and execution, creating space for ambitious people to build systems that support millions of businesses worldwide. Here, innovation thrives, ideas are celebrated and people come first, no matter where they call home.

Our impact

Every month, HighLevel enables more than 1.5 billion messages, 200 million leads and 20 million conversations for the more than 1 million businesses we support. Behind those numbers are real people building independence, expanding opportunity and creating measurable impact. We’re proud to be a part of that.

Learn more about us on our YouTube Channel or Blog PostsWhat are we hiring for?

Thousands of pods. 50+ kinds of deployments. Four databases. Three teams shipping fast. And until now, nobody whose whole job is the system itself.

Our teams are excellent at building products. Each service has owners, each feature has engineers, each database has experts. What we don't have is the person who holds the entire distributed system in their head, who sees that the retry policy in one service and the queue configuration two hops away are, together, a cascading failure waiting for Black Friday.

That's the seat. Your job is to find the failure before it finds production.

You start with Workflows, HighLevel's automation engine, and one of the largest systems in the company: 3.1 billion enrollments and 21.5 billion action executions every month, traffic peaking at 28,000+ requests per second, running across thousands of pods on GCP with Pub/Sub, Cloud Tasks, Redis, and multiple database engines underneath. From there, your blast radius grows, into Conversations (2.6 billion messages a month) and a brand-new Ticketing system being built right now, where you get to make sure it's born right instead of fixed later.

This is not an architect role where you draw boxes and hand them to someone else. And it's not a feature role where you own a backlog. It's the role in between that most companies never create, and most staff engineers spend their careers wishing existed.

Responsibilities

Own the architecture health of a billion-scale distributed system, its failure modes, capacity limits, consistency guarantees, and the interactions between 50+ deployments that no single team can see

Approve critical-path designs. Changes that touch the system's core go through you, not as bureaucracy, but as the person accountable for the whole staying sound. When there's a disagreement, you make your case on merit

Hunt gaps proactively, single points of failure, unbounded queues, missing idempotency, thundering herds, quiet data-loss windows, and drive the fixes before they become incidents

Build the parts nobody else can. No sprint tickets. You prototype the risky architectural bets yourself, ship the remediation after serious incidents, and pair into the gnarliest cross-team bugs, roughly a quarter to a third of your time in code, all of it on the hardest problems

Make resilience a property of the system, not a heroic act, degradation strategies, backpressure, isolation boundaries, capacity models that survive 10%+ month-over-month growth

Raise the teams around you. Design reviews that teach, post-mortems that change architecture (not just add alerts), and patterns that 80+ engineers build on

Set the standard for how AI-assisted engineering works safely on systems this critical, where a bad merge doesn't cost a demo, it costs real businesses their revenue

The terrain

Runtime
Node.js (TypeScript), Go, thousands of pods on GKE
Messaging & async
GCP Pub/Sub, Cloud Tasks, Redis
Storage
MongoDB, Firestore, ClickHouse, ElasticSearch
Scale
21.5B automation actions/month, 2.6B messages/month, 28.5K req/s peaks, ~226B async events across the org

Requirements

10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems, comparable scale strongly preferred: hundreds of services or thousands of instances, billions of daily events

You've carried sole accountability for a production system through real failures, not adjacent to it, not advising on it. Owned it

Deep command of queueing and async architectures, delivery semantics, ordering, backpressure, idempotency, exactly-once myths and at-least-once realities

Strong with multiple storage engines (SQL and NoSQL), you reason about consistency models, indexing at scale, and when each engine is the wrong choice

Expert-level depth in Redis or comparable in-memory systems, including their failure modes under memory pressure and network partition

Production experience on Kubernetes at scale, resource limits, autoscaling behavior, what actually happens when a node pool dies

Exceptional design communication, docs, diagrams, and RCAs that drive decisions across multiple teams

Fluent in Node.js and/or Go, enough to prototype your own proposals and ship fixes on the critical path

What Success Looks Like

The critical paths of Workflows have named owners, capa

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on Himalayas
Opens himalayas.app in a new tab