Fynd is a frontier technology company. We started at the intersection of technology and retail because that is where technology was the least available. Over the years, we became one of India’s largest retail technology platforms. But retail was the entry point, not the boundary.
Today, Fynd builds intelligent software that runs business operations. Not tools that help people work faster, but systems that absorb entire functions: manufacturing, marketing, logistics, commerce, quality control. We sit inside our customers’ businesses, harvest deep domain context, and build AI systems that operate autonomously. We are expanding from retail into manufacturing, generative media, physical AI, and healthcare.
About the role
Retail OS powers live store operations across fulfilment, delivery, complaints, and operational incidents. We are hiring an SDE - 2 or 3 who has built and operated products at scale and can bring that engineering discipline to AI agents, root-cause analysis, attribution, and autonomous workflows.
You will own services and features from design through production, combining performance, availability, and recovery engineering with applied AI. You will partner with senior engineers, operations teams, and upstream platform owners.
What you will do
- Build scalable backend services and event-processing pipelines with measurable throughput, latency, and availability objectives.
- Improve performance through profiling, efficient data access, concurrency management, capacity planning, and load testing.
- Engineer resilience through replication, failover, backpressure, graceful degradation, and recovery procedures; validate backup restoration and disaster recovery.
- Build incident correlation, root-cause analysis, and attribution systems that reconstruct operational timelines and link conclusions to evidence.
- Develop agents that investigate incidents using operational data, metrics, and runbooks, then execute permitted actions with deduplication, audit trails, escalation, and verified closure.
- Own production health through observability, on-call participation, incident response, postmortems, and preventive improvements.
What you must bring
- Experience operating products at scale: Direct ownership of production services with significant traffic, event volumes, or concurrency. You can explain their scale, bottlenecks, performance targets, and availability outcomes.
- Strong backend and distributed-systems fundamentals: Python, APIs, asynchronous processing, data modelling, consistency, idempotency, duplicate and late events, checkpoints, and replay.
- Hands-on availability and recovery experience: Replication, failover, backup and restore validation, and disaster-recovery exercises, including recovery time and recovery point objectives (RTO/RPO).
- Production debugging and performance depth: Experience diagnosing application, database, and infrastructure failures using logs, metrics, traces, profiling, and query analysis.
- Applied AI engineering: Experience shipping LLM applications or agents with tool calling, structured outputs, retrieval, and evaluations of correctness, latency, cost, and failure behaviour.
- AI-native development practices: Effective use of coding agents while independently reviewing, testing, and taking ownership of the resulting software.
- Sound operational judgment: Clear reasoning about evidence, uncertainty, permissions, rollback, and when human intervention is required.
Useful additional experience
Commerce, fulfilment, logistics, payments, observability, or workflow automation; Kubernetes/GCP; MongoDB; React; and operational attribution systems.
Our environment
Python, FastAPI, MongoDB, GKE, Databricks, React, Prometheus, and Grafana.
What success looks like
Your services meet agreed performance and availability objectives, recover predictably during failures, and have tested recovery procedures. Your AI workflows reduce investigation effort, improve attribution quality, and complete permitted actions with observable, verifiable outcomes.
What do we offer?
Growth
At Fynd, growth is limitless. We nurture a culture that encourages innovation, embraces challenges, and supports continuous learning. As we expand into new product lines and global markets, we’re seeking talented individuals eager to grow with us.
We believe in empowering our people to take ownership, lead with confidence, and shape their careers.
- Learning Wallet: Enrol in external courses or certifications to upskill—we’ll reimburse the costs to support your development.
Culture
We believe in building strong teams and lasting connections.
- Regular community engagement and team-building activities
- Biannual events to celebrate achievements, foster collaboration, and strengthen our workplace culture
Wellness
Your well-being is our priority. Comprehensive Mediclaim policy for you, your spouse, children, and parents
Work Environment
We thrive on collaboration and creativity. Our teams work from the office five days a week to encourage open communication, teamwork, and innovation.
Join us to be part of a dynamic environment where your ideas make an impact!