Remote (India) · Staff/Principal · Remote
Applicants who checked fit first are 3.1× more likely to hear back
Your score for this role already exists
ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.
No credit card · 1 tap with Google
MEDIUM
Jobgether is reviewing applications at a steady pace. Expect a standard response time as they evaluate the current pool.
First 72 hours
Still inside it - posted 1h agoEarly applicants get seen before the pile builds.
Not a repost
The first time we've seen this listing - it hasn't been closed and reopened.
You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer, Production Engineering based in India.
This is a senior production engineering opportunity focused on keeping large-scale cloud and infrastructure services highly available, reliable, and secure. You’ll support production Kubernetes services within a global 24/7 reliability operations environment, with a strong emphasis on automation and reducing manual operational work. The role combines deep systems expertise with incident management, observability, cluster administration, and reliability engineering. You’ll work closely with Site Reliability Engineering, security, DevOps, development teams, and service owners to prevent and resolve complex production issues. Your work will directly contribute to service availability, customer experience, and operational resilience at significant scale. The position is well suited to an engineer who enjoys solving complex infrastructure problems and working with advanced cluster technologies.
Support production Kubernetes services as part of a global 24/7 production engineering operation, including flexibility to work split-weekend shifts.
Administer and maintain large-scale Kubernetes clusters, systems, and infrastructure while protecting service availability, integrity, reliability, and SLAs.
Automate operational processes and continuously identify opportunities to reduce manual tasks and improve engineering efficiency.
Use monitoring, observability, alerts, and alarms to proactively detect, prevent, investigate, and respond to production incidents.
Analyze logs, metrics, system behavior, and infrastructure signals to troubleshoot complex issues and determine root causes.
Lead incident management calls, coordinating timely detection, escalation, investigation, and resolution of critical production issues.
Engage subject matter experts, service owners, and cross-functional engineering teams to resolve complex incidents efficiently.
Develop and improve monitoring, alerting, and reliability mechanisms in collaboration with development teams.
Perform systems administration and security monitoring across large-scale infrastructure environments.
Apply deep knowledge of Linux, networking, Kubernetes, and cluster infrastructure to maintain reliable production services.
Contribute to the architecture, deployment, and ongoing improvement of Kubernetes environments operating at significant scale.
Continuously evaluate emerging infrastructure and high-performance computing technologies and identify opportunities for innovation.
7+ years of demonstrated experience administering large-scale production Kubernetes environments within high-availability Internet, cloud, or data-center environments, with strong on-premises experience preferred.
Bachelor's degree in Computer Science, Engineering, Mathematics, or a related discipline, or equivalent professional experience.
Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management.
Familiarity with GPU/DPU hardware and high-performance computing cluster environments.
Strong Linux systems administration experience, including DNS, DHCP, IP tables, routing, firewalls, and core Linux networking.
Proven ability to troubleshoot and maintain services across large-scale bare-metal infrastructure.
Experience with CI/CD technologies and tools such as Jenkins and ArgoCD.
Scripting or programming experience in Python, Golang, or Rust is preferred but not mandatory.
Strong understanding of observability, incident management, reliability engineering, and production operations.
Excellent analytical and troubleshooting skills, with the ability to work effectively under pressure during complex incidents.
Strong communication and interpersonal skills, including the ability to clearly present technical information and influence cross-functional stakeholders.
Ability to learn new technologies quickly and adapt to evolving infrastructure environments.
Experience architecting, building, and deploying Kubernetes environments at large scale is highly valuable.
Passion for innovation and advanced high-performance cluster technologies is an advantage.
Full-time opportunity with a remote working option in India.
Opportunity to work on large-scale production Kubernetes and infrastructure environments.
Exposure to advanced cloud, bare-metal, GPU/DPU, and high-performance computing technologies.
Opportunity to work alongside SRE, DevOps, security, development, and other specialized engineering teams.
Significant technical ownership across reliability, automation, observability, incident response, and infrastructure operations.
Opportunity to solve complex engineering challenges at global scale.
Continuous exposure to emerging technologies and opportunities to develop advanced infrastructure expertise.
24/7 production engineering environment offering substantial experience in incident management and high-availability operations.
Compensation, healthcare, leave, and other employment benefits are provided according to the applicable employment package and location-specific terms.
Free · no signup
Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.
No spam. Just jobs and resources.
Why people use ASAI
Scored, not searched. Every role ranked against your actual profile.
Alerts as often as hourly. Reach new roles while the pile is still small.
Skill gaps, spelled out. See exactly which requirements you don't meet yet.
Verified jobs, only. Say no to ghost jobs. Your time deserves respect.
Keep browsing