Job
Description
Senior
Devops Engineer
Location:
Pune | Hybrid
Experience:
6-8 years
Primary
Skill: Linux programming , AWS,
Terraform/OpenTofu, Cloud Migration, DevOps/SRE
AI-Driven Infrastructure Modernization
About
the Role
We are
hiring a Senior Cloud Platform / DevOps / SRE Engineer to help migrate,
modernize, operate, and improve infrastructure across a portfolio of
businesses. This role will focus first on AWS, with future expansion
into Azure and GCP.
The ideal
candidate has experience moving legacy infrastructure from colo, hosted,
data center, or existing cloud environments into AWS. This includes both colo-to-cloud
and cloud-to-cloud migrations. Early work may include lift-and-shift
migrations; over time, this role will help standardize, automate, secure,
observe, and modernize those environments.
This is
also a production reliability role. We expect this team to support
production uptime, participate in on-call or escalation processes, and help
ensure our infrastructure and deployment practices support reliable
customer-facing systems.
This is an
AI-first infrastructure role. We expect the right person to use AI tools
to accelerate infrastructure analysis, Terraform/OpenTofu development,
migration planning, CI/CD improvement, troubleshooting, documentation,
automation, and operational remediation.
Key
Responsibilities
- Lead and support migrations
from colo, hosted, legacy, and existing cloud environments into AWS.
- Assess current-state
infrastructure, document dependencies, identify risks, and create
practical migration, cutover, rollback, and validation plans.
- Map source environments —
including networking, IAM, storage, databases, DNS, certificates, security
controls, observability, and deployment workflows — into appropriate AWS
target architectures.
- Build, maintain, and
standardize cloud infrastructure using Terraform or OpenTofu, with
awareness of AWS CDK/CloudFormation, Azure Bicep, Pulumi, and other
cloud-native IaC tools.
- Operate and improve
multi-account AWS environments using AWS Organizations and related
governance patterns.
- Help establish a centralized
cloud operating model across AWS, with a path toward Azure and GCP.
- Design and support cloud
networking, IAM, security, logging, monitoring, backups, disaster
recovery, and high availability patterns.
- Build and improve CI/CD
pipelines for application and infrastructure deployments.
- Define opinionated best
practices for infrastructure automation, CI/CD, observability,
reliability, security, and cloud operations.
- Implement monitoring, logging,
tracing, alerting, dashboards, and service health indicators so issues are
detected quickly and resolved effectively.
- Build automation and
self-healing workflows so common production failures can be corrected
automatically where possible.
- Participate in production
support, incident response, escalation, and on-call processes as needed.
- Partner with engineering teams
to containerize legacy applications using Docker and deploy to ECS, EKS,
Kubernetes, or similar platforms.
- Use AI tools to accelerate
infrastructure analysis, IaC creation, pipeline development,
troubleshooting, documentation, incident response, and operational
automation.
Required
Qualifications
- 7+ years of experience in
DevOps, SRE, cloud infrastructure, platform engineering, systems
engineering, or similar roles.
- Strong hands-on experience
operating production workloads in AWS.
- Experience migrating
infrastructure from colo, data center, hosted, legacy, or existing cloud
environments into AWS.
- Experience with cloud-to-cloud
migrations, including service mapping, data migration, networking,
identity/access, DNS, cutover, rollback, and validation.
- Strong production experience
with Terraform or OpenTofu, including modules, remote state,
environments, plan/apply workflows, secrets handling, policy checks, and
CI/CD integration.
- Experience with AWS
networking, including VPCs, subnets, routing, VPNs, load balancers, DNS,
certificates, NAT gateways, and security groups.
- Experience with AWS
Organizations, IAM, centralized logging, cloud governance, and
multi-account patterns.
- Experience building and
maintaining CI/CD pipelines for application and infrastructure delivery.
- Experience with observability,
including monitoring, logging, alerting, dashboards, metrics, traces, and
service health indicators.
- Experience owning or
supporting production uptime, including incident response, on-call or
escalation workflows, runbooks, alerting, and post-incident improvement.
- Strong troubleshooting skills
across infrastructure, networking, application, and cloud layers.
- Demonstrated use of AI tools
in infrastructure, DevOps, SRE, or software delivery workflows.
- Strong communication skills
and ability to work across multiple engineering teams and business units.
Preferred
Qualifications
- Experience with Azure
and/or GCP in addition to AWS.
- Direct experience migrating
workloads from Azure to AWS, GCP to AWS, or other cloud
providers into AWS.
- Experience with Azure Bicep,
AWS CDK/CloudFormation, Pulumi, or other infrastructure-as-code
approaches.
- Experience with AWS Control
Tower, IAM Identity Center, CloudTrail, Config, GuardDuty, Security Hub,
or similar governance and security services.
- Experience with Azure
management groups, subscriptions, policies, identity, networking, and
governance patterns.
- Experience with GCP
organizations, folders, projects, IAM, networking, and organization
policies.
- Experience with Docker, ECS,
EKS, Kubernetes, Helm, or other container platforms.
- Experience containerizing
legacy applications and moving them toward automated deployment models.
- Experience with GitHub
Actions, GitLab CI, Azure DevOps, Jenkins, Argo CD, CircleCI, or similar
CI/CD platforms.
- Experience with Datadog, New
Relic, Grafana, Prometheus, CloudWatch, OpenTelemetry, ELK/OpenSearch,
Splunk, or similar observability platforms.
- Experience with automated
remediation, self-healing infrastructure, auto-scaling, event-driven
operations, and runbook automation.
- Experience with
policy-as-code, security-as-code, backup, disaster recovery, high
availability, and compliance automation.
- Enough application
architecture knowledge to assess whether an application should be
lifted-and-shifted, containerized, re-platformed, or more deeply
modernized.