We are looking for a Cloud
Solution Architect whose primary responsibility is to design, build, manage,
and optimize the enterprise cloud infrastructure that runs our AI
ecosystem.
Job Description:
Role: Cloud Solution
Architect – Enterprise AI Infrastructure & FinOps
Experience: 12–18 years total, with at least 6 years in cloud
architecture and at least 3 years architecting production AI/ML or GenAI
platforms
Domain: Enterprise AI Platform (Pharma / Life Sciences, GxP-aware)
Education: B.E./B.Tech in Computer Science, Information Technology,
Information Security, Electronics & Communication, or related Engineering
disciplines.
Location: Ahmedabad/Mumbai
Role Summary
The individual hired as Cloud
Solution Architect will be responsible for cloud foundation end to end: landing
zones, networking, compute, storage, identity, security, Kubernetes platforms,
GPU capacity, infrastructure-as-code, monitoring, resilience, and day-to-day
infrastructure operations.
The role will ensure that AI and
GenAI workloads operate securely, reliably, and at scale while driving cloud
governance, operational excellence, and FinOps practices to maintain
transparency, control, and optimization of cloud and AI spending.
This is a hands-on infrastructure
leadership role. Application and AI engineering teams build agents and
workflows; you make sure the platform they run on is secure, stable, scalable,
compliant, and cost-efficient.
Key Responsibilities
1. Cloud Infrastructure
Architecture
- Own the enterprise cloud infrastructure
architecture on Azure, AWS, and/or GCP, including multi-cloud and
hybrid designs.
- Design and maintain enterprise landing zones:
management group and account/subscription structure, policies, guardrails,
naming and tagging standards, and environment separation (dev, test,
validation, production).
- Architect network infrastructure:
hub-and-spoke or Virtual WAN topologies, VNets/VPCs, subnets, private
endpoints, DNS, load balancers, application gateways, WAF,
ExpressRoute/Direct Connect, site-to-site VPN, and hybrid connectivity to
data centers and manufacturing sites.
- Design compute platforms across VMs, VM scale sets,
containers, Kubernetes (AKS, EKS, GKE), serverless (Functions,
Lambda, Cloud Run), and managed PaaS services.
- Architect storage and database infrastructure:
object storage, file shares, managed disks, backup vaults, and managed
databases (SQL, PostgreSQL, Cosmos DB, DynamoDB, Redis).
- Design for high availability, multi-region
resilience, disaster recovery, and business continuity with defined
RTO and RPO targets.
- Maintain architecture documentation, reference
designs, and architecture decision records (ADRs).
2. Cloud Infrastructure
Management and Operations
- Own the day-to-day health, stability, and
performance of cloud infrastructure supporting AI and enterprise
workloads.
- Lead infrastructure provisioning, configuration,
patching, upgrades, and lifecycle management for VMs, Kubernetes clusters,
networking, and platform services.
- Manage Kubernetes platform operations:
cluster upgrades, node pools (CPU and GPU), ingress, service mesh,
autoscaling, namespaces, and multi-tenancy for AI teams.
- Define and track infrastructure SLOs, SLAs, and
capacity plans, and lead major incident response, root cause analysis,
and problem management.
- Implement backup, restore, and DR testing on
a regular schedule and maintain runbooks.
- Establish ITIL-aligned processes for incident,
change, problem, and service request management in collaboration with IT
operations and managed service providers.
- Manage cloud vendors and managed service partners,
including SLA and performance governance.
- Drive operational automation to reduce manual
effort and human error.
3. Infrastructure-as-Code,
Automation and Platform Engineering
- Define and enforce infrastructure-as-code
standards using Terraform, Bicep/ARM, Pulumi, or CloudFormation.
- Build reusable IaC modules, templates, and
"golden paths" so teams can self-serve compliant infrastructure
quickly.
- Implement CI/CD and GitOps for
infrastructure (GitHub Actions, Azure DevOps, GitLab, Argo CD, Flux).
- Apply policy-as-code (Azure Policy, AWS
SCPs/Config, OPA/Gatekeeper) to prevent misconfiguration and enforce
security and cost guardrails.
- Automate environment provisioning, scaling,
patching, compliance checks, and cleanup.
- Build an internal developer platform experience for
AI, data, and application teams.
4. AI Infrastructure Enablement
- Design and operate the infrastructure foundation
for AI workloads, covering agent runtimes, workflow engines, data
pipelines, model serving, and vector and search services.
- Provision and manage GPU and accelerator
capacity (NVIDIA A100/H100/L40S, Inferentia/Trainium, TPUs) for
inference, fine-tuning, and batch processing, including quota management
and capacity reservations.
- Host and secure LLM access infrastructure:
Azure OpenAI, AWS Bedrock, Vertex AI, and self-hosted open-source models
(vLLM, TGI, Triton, KServe, Ray).
- Deploy and operate a centralized AI/LLM gateway
for secure model access, routing, rate limiting, caching, logging, and
cost attribution.
- Provide infrastructure for agent frameworks and
orchestration tools (e.g., LangGraph, Semantic Kernel, Azure AI Foundry,
Bedrock Agents, Temporal, Airflow, Logic Apps, Step Functions), including
MCP servers and secure tool connectors.
- Support data platforms used by pipelines, such as
Databricks, Snowflake, Microsoft Fabric, Synapse, Kafka/Event Hub, and
vector stores (Azure AI Search, OpenSearch, pgvector, Pinecone, Qdrant).
- Ensure network isolation, private connectivity, and
secure integration between AI services and enterprise systems (SAP, LIMS,
MES, QMS, document repositories, and on-premises data).
- Partner with AI engineering teams on LLMOps/MLOps
infrastructure: model registries, experiment tracking, CI/CD for AI
assets, and environment promotion.
- Architect
the agent runtime platform: hosting, lifecycle management, state
and memory, tool execution, session handling, and scaling.
- Define
standards for agent frameworks and patterns (e.g., LangGraph, Semantic
Kernel, AutoGen, CrewAI, OpenAI Agents SDK, AWS Bedrock Agents, Azure AI
Foundry Agent Service).
- Design multi-agent orchestration patterns:
supervisor/worker, hierarchical, planner-executor, and agent-to-agent
(A2A) collaboration.
- Design
the orchestration layer for AI-enabled business workflows combining
agents, deterministic logic, APIs, and human tasks.
- Standardize
on durable orchestration and workflow engines (e.g., Temporal, Azure
Durable Functions, Logic Apps, AWS Step Functions, Airflow, Prefect, n8n,
Power Automate) based on use case.
- Define event-driven architectures using Kafka,
Event Hub, EventBridge, or Pub/Sub for real-time, asynchronous triggers.
5. Cloud FinOps and AI Cost
Management
- Establish and run the cloud FinOps practice
based on FinOps Foundation principles (Inform, Optimize, Operate) and the
FOCUS billing standard.
- Implement mandatory tagging and cost allocation
so every resource and AI workload maps to an owner, cost center,
application, agent, workflow, or pipeline.
- Deliver showback and chargeback reporting to
business units and product teams.
- Track cloud spend across compute, storage,
networking, data egress, PaaS services, GPUs, and LLM token consumption.
- Define unit economics such as cost per agent
interaction, cost per workflow run, cost per pipeline execution, and cost
per environment.
- Drive infrastructure cost optimization through:
- Rightsizing VMs, clusters, and databases
- Reserved instances, savings plans, and committed
use discounts
- Spot and preemptible capacity for non-critical and
batch workloads
- Autoscaling, scale-to-zero, and scheduled shutdown
of non-production environments
- Storage tiering, lifecycle policies, and orphaned
resource cleanup
- GPU utilization improvement and workload
consolidation
- LLM cost controls such as model routing, caching,
batch APIs, and provisioned throughput decisions
- Set budgets, quotas, and anomaly alerts per
subscription, team, and AI workload.
- Produce forecasts and executive dashboards linking
cloud and AI spend to business value.
- Support Finance and Procurement on enterprise
agreements, commitments, and vendor negotiations.
6. Cloud Security and
Compliance
- Design and enforce cloud security architecture
based on zero-trust principles.
- Own identity and access architecture (Entra ID, AWS
IAM, GCP IAM), RBAC, privileged access management, managed identities, and
workload identity for agents and services.
- Implement secrets and key management (Key Vault,
KMS, HashiCorp Vault), encryption at rest and in transit, and certificate
lifecycle management.
7. Monitoring, Observability
and Reliability
- Architect enterprise monitoring and observability
for infrastructure and AI workloads using Azure Monitor, CloudWatch,
Google Cloud Operations, Datadog, Dynatrace, Prometheus/Grafana, and
OpenTelemetry.
- Monitor infrastructure health, performance,
capacity, GPU utilization, network latency, and cost in unified
dashboards.
- Integrate AI-level telemetry (token usage, latency,
failures) from tools such as Langfuse or LangSmith with infrastructure
observability.
- Apply SRE practices: error budgets,
proactive alerting, chaos testing, and post-incident reviews.
8. Technical Leadership and
Governance
- Act as the design authority for cloud
infrastructure, reviewing and approving infrastructure designs from
project and product teams.
- Define cloud standards, policies, and best
practices, and drive consistent adoption across teams.
- Lead and mentor cloud engineers, DevOps/platform
engineers, and operations staff.
Required Skills and Experience
- 12+ years in IT infrastructure, with at least 7
years of hands-on cloud infrastructure architecture and operations
on Azure, AWS, or GCP (multi-cloud preferred).
- Deep expertise in cloud networking, identity,
compute, storage, and security services.
- Strong hands-on experience with Kubernetes
(AKS/EKS/GKE) and containers in production, including GPU node pools.
- Advanced infrastructure-as-code skills
(Terraform preferred) and CI/CD/GitOps practices.
- Proven experience managing production cloud
environments at enterprise scale, including incidents, DR, patching, and
change management.
- Demonstrated experience implementing cloud
FinOps: tagging, allocation, optimization, and reporting.
- Practical experience supporting AI/ML or GenAI
workloads: LLM services, model serving, GPU infrastructure, vector
databases, and data platforms.
- Strong understanding of cloud security, zero trust,
and compliance frameworks.
- Scripting skills in Python, PowerShell, or Bash.
- Experience in regulated industries; pharma or life
sciences GxP experience is a strong advantage.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science,
Information Technology, Engineering, or a related field.
- Certifications such as Azure Solutions Architect
Expert, AWS Solutions Architect Professional, Google Professional Cloud
Architect, CKA/CKAD, HashiCorp Terraform Associate, FinOps
Certified Practitioner/Engineer, Azure AI Engineer, AWS ML Specialty,
or ITIL.
- Experience with hybrid infrastructure connecting
cloud to on-premises data centers and manufacturing/OT environments.
- Exposure to AI governance and Responsible AI
frameworks.