Job Title: Production Site Reliability Engineer (SRE) – Digital Payments
Role Overview
We are seeking a highly technical and driven Production SRE Engineer to manage and monitor mission-critical payment platforms including UPI, IMPS, and Payment Hub systems & various Payment Applications.
The role focuses on ensuring high availability, low latency, and seamless transaction experience for customers. The incumbent will collaborate with cross-functional teams (Engineering, Business, Compliance) and external regulators (RBI, NPCI) to maintain resilient and scalable payment infrastructure.
Key Responsibilities
Production Support & Incident Management
- Provide L2/L3 production support for UPI, IMPS, and Payment Hub platforms & various Payment Applications.
- Diagnose, triage, and resolve transaction failures, timeouts, and API disruptions.
- Lead and participate in Major Incident Management (MIM) calls and ensure timely stakeholder communication.
- Manage incidents, service requests, and problem tickets via Jira, ServiceNow.
- Provide regular updates to internal stakeholders and regulatory bodies (NPCI/RBI) during critical issues.
Reliability Engineering & RCA
- Perform deep-dive Root Cause Analysis (RCA) for recurring payment and system issues.
- Implement preventive and corrective measures to improve system stability.
- Drive SRE best practices including error budgets, SLIs/SLOs, and system resilience.
- Experience in managing DR Drills & Documentations.
- Reviewing the SOPs & its relative documentations.
Monitoring, Observability & System Engineering
- Monitor key performance indicators:
- Transaction success rates
- Latency and response times
- Failure trends and retries
- Build and maintain dashboards using:
- ELK Stack, Grafana, Kibana, Splunk, Datadog, Prometheus
- Establish proactive alerting and anomaly detection mechanisms.
- Work closely with engineering teams to design and optimize:
- High-throughput payment switches
- Routing logic
- Settlement and reconciliation systems
- Understand and support UPI architecture, IMPS rails, and payment orchestration layers & various Payment Applications.
- Trace end-to-end transaction lifecycle across distributed systems.
External Partner & Regulatory Coordination
- Coordinate with NPCI, partner banks, and TPAPs during outages, reconciliation issues, or network disruptions.
- Lead integrations and ensure seamless onboarding of ecosystem participants.
- Ensure compliance with:
- RBI guidelines and data localization mandates
- NPCI operational and technical standards
Technical Skills & Expertise
Payments Domain Knowledge
- Strong expertise in:
- UPI architecture and flows
- IMPS rails
- Payment gateway / switch systems
- Payment Hub orchestration & various Payment Applications.
Core Technical Skills
- Advanced SQL proficiency (joins, aggregations, stored procedures)
- Strong hands-on experience in:
- Linux/UNIX systems administration
- Shell scripting
- Ability to:
- Read , Write and interpret All types documentation (SOPs, workflows, etc.)
- Understand database schemas
- Analyse system architecture and latency
Monitoring & Observability Tools
- Hands-on expertise with:
- ELK Stack (Elasticsearch, Logstash, Kibana)
- Grafana, Prometheus
- Splunk, Datadog
DevOps & Cloud
- Experience with:
- CI/CD pipelines, Containerization (Docker, Kubernetes)
- Cloud platforms:
Key Competencies
- Strong problem-solving and analytical skills
- High ownership in production environments
- Ability to work under pressure in real-time systems
- Strong stakeholder communication and coordination
- Focus on reliability, scalability, and performance
- Must have can do, takes initiative, Drives end to end deliverables.
- Proactive, solution-oriented mindset with ownership to resolve production issues under pressure.
- Ability to clearly articulate incidents, updates, and RCA to stakeholders, leadership, and regulators.
- Works effectively with cross-functional teams (engineering, product, partners, regulators).
- Structured thinking to diagnose complex system failures and drive long-term fixes.
- Ability to stay calm and effective during high-severity incidents and critical outages.
- Knowledge of PCI-DSS compliance, Financial data governance & security best practices
- Quickly adapts to changing technologies, incidents, and regulatory requirements in a fast-evolving payments ecosystem.
- Precision in analysing logs, transactions, and system behaviour to avoid critical errors in production.
- Effectively manage multiple incidents, tasks, and escalations in a high-pressure environment.
- Ability to handle expectations and coordinate with internal teams, partners, and regulators efficiently.
- Takes ownership to make quick, informed decisions during outages or critical production incidents & communications to various Stake holders including Regulatory.