Remote (India) · Staff/Principal · Remote
Applicants who checked fit first are 3.1× more likely to hear back
Your score for this role already exists
ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.
No credit card · 1 tap with Google
HIGH
HighLevel is actively reviewing profiles and moving candidates through the pipeline right now.
First 72 hours
Still inside it - posted 3h agoEarly applicants get seen before the pile builds.
Not a repost
The first time we've seen this listing - it hasn't been closed and reopened.
You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.
Thousands of pods. 50+ kinds of deployments. Four databases. Three teams shipping fast. And until now, nobody whose whole job is the system itself.
Our teams are excellent at building products. Each service has owners, each feature has engineers, each database has experts. What we don't have is the person who holds the entire distributed system in their head, who sees that the retry policy in one service and the queue configuration two hops away are, together, a cascading failure waiting for Black Friday.
That's the seat. Your job is to find the failure before it finds production.
You start with Workflows, HighLevel's automation engine, and one of the largest systems in the company: 3.1 billion enrollments and 21.5 billion action executions every month, traffic peaking at 28,000+ requests per second, running across thousands of pods on GCP with Pub/Sub, Cloud Tasks, Redis, and multiple database engines underneath. From there, your blast radius grows, into Conversations (2.6 billion messages a month) and a brand-new Ticketing system being built right now, where you get to make sure it's born right instead of fixed later.
This is not an architect role where you draw boxes and hand them to someone else. And it's not a feature role where you own a backlog. It's the role in between that most companies never create, and most staff engineers spend their careers wishing existed.
Own the architecture health of a billion-scale distributed system, its failure modes, capacity limits, consistency guarantees, and the interactions between 50+ deployments that no single team can see
Approve critical-path designs. Changes that touch the system's core go through you, not as bureaucracy, but as the person accountable for the whole staying sound. When there's a disagreement, you make your case on merit
Hunt gaps proactively, single points of failure, unbounded queues, missing idempotency, thundering herds, quiet data-loss windows, and drive the fixes before they become incidents
Build the parts nobody else can. No sprint tickets. You prototype the risky architectural bets yourself, ship the remediation after serious incidents, and pair into the gnarliest cross-team bugs, roughly a quarter to a third of your time in code, all of it on the hardest problems
Make resilience a property of the system, not a heroic act, degradation strategies, backpressure, isolation boundaries, capacity models that survive 10%+ month-over-month growth
Raise the teams around you. Design reviews that teach, post-mortems that change architecture (not just add alerts), and patterns that 80+ engineers build on
Set the standard for how AI-assisted engineering works safely on systems this critical, where a bad merge doesn't cost a demo, it costs real businesses their revenue
The terrain:
Runtime: Node.js (TypeScript), Go, thousands of pods on GKE
Messaging & async: GCP Pub/Sub, Cloud Tasks, Redis
Storage: MongoDB, Firestore, ClickHouse, ElasticSearch
Scale: 21.5B automation actions/month, 2.6B messages/month, 28.5K req/s peaks, ~226B async events across the org
10+ years of engineering experience with deep, hands-on ownership of large-scale distributed systems, comparable scale strongly preferred: hundreds of services or thousands of instances, billions of daily events
You've carried sole accountability for a production system through real failures, not adjacent to it, not advising on it. Owned it
Deep command of queueing and async architectures, delivery semantics, ordering, backpressure, idempotency, exactly-once myths and at-least-once realities
Strong with multiple storage engines (SQL and NoSQL), you reason about consistency models, indexing at scale, and when each engine is the wrong choice
Expert-level depth in Redis or comparable in-memory systems, including their failure modes under memory pressure and network partition
Production experience on Kubernetes at scale, resource limits, autoscaling behavior, what actually happens when a node pool dies
Exceptional design communication, docs, diagrams, and RCAs that drive decisions across multiple teams
Fluent in Node.js and/or Go, enough to prototype your own proposals and ship fixes on the critical path
The critical paths of Workflows have named owners, capacity models, and tested failure modes, because you made it so
Incident count trends down while traffic grows double-digit percent month over month
Your design reviews are the ones engineers want their proposals to survive
Your scope has expanded on results, more systems, more surface, more trust
You've been injured in production. A lot. You've owned distributed systems at serious scale, through the outages, the migrations, the 3 AM discoveries, and every scar changed how you design
You've operated at staff scope, whatever your title said, the engineer everyone routed the hardest systems questions to
You think in failure modes by default: when you see a design, you instinctively ask what happens at the tail, under partition, at 10x load
You write design docs and RCAs that people reference years later, clear trade-offs, honest risks, real recommendations
You can disagree with a team and still make them better, influence through rigor and respect, not title
You'd rather prevent ten incidents quietly than be the hero of one loudly
You've made AI agents genuinely productive on complex systems, and know how to keep AI-generated code from becoming AI-generated incidents
GCP-native experience: Pub/Sub, Cloud Tasks, GKE, Firestore
You've done this job before under another name, "the systems person," principal engineer, architect-who-still-codes
Experience taking a 0→1 system to production alongside hardening mature ones
EEO Statement:
The company is an Equal Opportunity Employer. As an employer subject to affirmative action regulations, we invite you to voluntarily provide the following demographic information. This information is used solely for compliance with government recordkeeping, reporting, and other legal requirements. Providing this information is voluntary and refusal to do so will not affect your application status. This data will be kept separate from your application and will not be used in the hiring decision.
We encourage you to review our Privacy Policy before submitting your application
Free · no signup
Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.
No spam. Just jobs and resources.
Why people use ASAI
Scored, not searched. Every role ranked against your actual profile.
Alerts as often as hourly. Reach new roles while the pile is still small.
Skill gaps, spelled out. See exactly which requirements you don't meet yet.
Verified jobs, only. Say no to ghost jobs. Your time deserves respect.
Keep browsing