ASAI job platform logo
  1. Home
  2. /Jobs
  3. /Cloud Solution Architect – Enterprise AI Infrastructure & FinOps
Zydus Group

Cloud Solution Architect – Enterprise AI Infrastructure & FinOps

Ahmedabad · Executive

Good timing - competition is buildingVerified listingPosted 10d ago

Applicants who checked fit first are 3.1× more likely to hear back

Your match scoreCalculated · locked
86Overall
64Skills
97Experience

Your score for this role already exists

ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.

No credit card · 1 tap with Google

What we know about this role

Hiring pulse

HIGH

Zydus Group is actively reviewing profiles and moving candidates through the pipeline right now.

Apply window

First 72 hours

Window passed - posted 10d ago

Early applicants get seen before the pile builds.

Not a repost

The first time we've seen this listing - it hasn't been closed and reopened.

Skills required

70 listed
Service Request ManagementCI/CDCloud InfrastructureCloud GovernanceNetwork InfrastructureDirect ConnectObject StorageDisaster Recovery+62 more
Service Request ManagementCI/CDCloud InfrastructureCloud GovernanceNetwork Infrastructure

You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.

Job description

We are looking for a Cloud Solution Architect whose primary responsibility is to design, build, manage, and optimize the enterprise cloud infrastructure that runs our AI ecosystem.

                                                     

Job Description:

Role: Cloud Solution Architect – Enterprise AI Infrastructure & FinOps
Experience: 12–18 years total, with at least 6 years in cloud architecture and at least 3 years architecting production AI/ML or GenAI platforms
Domain: Enterprise AI Platform (Pharma / Life Sciences, GxP-aware)
Education: B.E./B.Tech in Computer Science, Information Technology, Information Security, Electronics & Communication, or related Engineering disciplines.

Location: Ahmedabad/Mumbai

 

Role Summary

The individual hired as Cloud Solution Architect will be responsible for cloud foundation end to end: landing zones, networking, compute, storage, identity, security, Kubernetes platforms, GPU capacity, infrastructure-as-code, monitoring, resilience, and day-to-day infrastructure operations.

 

The role will ensure that AI and GenAI workloads operate securely, reliably, and at scale while driving cloud governance, operational excellence, and FinOps practices to maintain transparency, control, and optimization of cloud and AI spending.

 

This is a hands-on infrastructure leadership role. Application and AI engineering teams build agents and workflows; you make sure the platform they run on is secure, stable, scalable, compliant, and cost-efficient.

 

Key Responsibilities

1. Cloud Infrastructure Architecture

  • Own the enterprise cloud infrastructure architecture on Azure, AWS, and/or GCP, including multi-cloud and hybrid designs.

  • Design and maintain enterprise landing zones: management group and account/subscription structure, policies, guardrails, naming and tagging standards, and environment separation (dev, test, validation, production).

  • Architect network infrastructure: hub-and-spoke or Virtual WAN topologies, VNets/VPCs, subnets, private endpoints, DNS, load balancers, application gateways, WAF, ExpressRoute/Direct Connect, site-to-site VPN, and hybrid connectivity to data centers and manufacturing sites.

  • Design compute platforms across VMs, VM scale sets, containers, Kubernetes (AKS, EKS, GKE), serverless (Functions, Lambda, Cloud Run), and managed PaaS services.

  • Architect storage and database infrastructure: object storage, file shares, managed disks, backup vaults, and managed databases (SQL, PostgreSQL, Cosmos DB, DynamoDB, Redis).

  • Design for high availability, multi-region resilience, disaster recovery, and business continuity with defined RTO and RPO targets.

  • Maintain architecture documentation, reference designs, and architecture decision records (ADRs).

2. Cloud Infrastructure Management and Operations

  • Own the day-to-day health, stability, and performance of cloud infrastructure supporting AI and enterprise workloads.

  • Lead infrastructure provisioning, configuration, patching, upgrades, and lifecycle management for VMs, Kubernetes clusters, networking, and platform services.

  • Manage Kubernetes platform operations: cluster upgrades, node pools (CPU and GPU), ingress, service mesh, autoscaling, namespaces, and multi-tenancy for AI teams.

  • Define and track infrastructure SLOs, SLAs, and capacity plans, and lead major incident response, root cause analysis, and problem management.

  • Implement backup, restore, and DR testing on a regular schedule and maintain runbooks.

  • Establish ITIL-aligned processes for incident, change, problem, and service request management in collaboration with IT operations and managed service providers.

  • Manage cloud vendors and managed service partners, including SLA and performance governance.

  • Drive operational automation to reduce manual effort and human error.

3. Infrastructure-as-Code, Automation and Platform Engineering

  • Define and enforce infrastructure-as-code standards using Terraform, Bicep/ARM, Pulumi, or CloudFormation.

  • Build reusable IaC modules, templates, and "golden paths" so teams can self-serve compliant infrastructure quickly.

  • Implement CI/CD and GitOps for infrastructure (GitHub Actions, Azure DevOps, GitLab, Argo CD, Flux).

  • Apply policy-as-code (Azure Policy, AWS SCPs/Config, OPA/Gatekeeper) to prevent misconfiguration and enforce security and cost guardrails.

  • Automate environment provisioning, scaling, patching, compliance checks, and cleanup.

  • Build an internal developer platform experience for AI, data, and application teams.

4. AI Infrastructure Enablement

  • Design and operate the infrastructure foundation for AI workloads, covering agent runtimes, workflow engines, data pipelines, model serving, and vector and search services.

  • Provision and manage GPU and accelerator capacity (NVIDIA A100/H100/L40S, Inferentia/Trainium, TPUs) for inference, fine-tuning, and batch processing, including quota management and capacity reservations.

  • Host and secure LLM access infrastructure: Azure OpenAI, AWS Bedrock, Vertex AI, and self-hosted open-source models (vLLM, TGI, Triton, KServe, Ray).

  • Deploy and operate a centralized AI/LLM gateway for secure model access, routing, rate limiting, caching, logging, and cost attribution.

  • Provide infrastructure for agent frameworks and orchestration tools (e.g., LangGraph, Semantic Kernel, Azure AI Foundry, Bedrock Agents, Temporal, Airflow, Logic Apps, Step Functions), including MCP servers and secure tool connectors.

  • Support data platforms used by pipelines, such as Databricks, Snowflake, Microsoft Fabric, Synapse, Kafka/Event Hub, and vector stores (Azure AI Search, OpenSearch, pgvector, Pinecone, Qdrant).

  • Ensure network isolation, private connectivity, and secure integration between AI services and enterprise systems (SAP, LIMS, MES, QMS, document repositories, and on-premises data).

  • Partner with AI engineering teams on LLMOps/MLOps infrastructure: model registries, experiment tracking, CI/CD for AI assets, and environment promotion.

  • Architect the agent runtime platform: hosting, lifecycle management, state and memory, tool execution, session handling, and scaling.

  • Define standards for agent frameworks and patterns (e.g., LangGraph, Semantic Kernel, AutoGen, CrewAI, OpenAI Agents SDK, AWS Bedrock Agents, Azure AI Foundry Agent Service).

  • Design multi-agent orchestration patterns: supervisor/worker, hierarchical, planner-executor, and agent-to-agent (A2A) collaboration.

  • Design the orchestration layer for AI-enabled business workflows combining agents, deterministic logic, APIs, and human tasks.

  • Standardize on durable orchestration and workflow engines (e.g., Temporal, Azure Durable Functions, Logic Apps, AWS Step Functions, Airflow, Prefect, n8n, Power Automate) based on use case.

  • Define event-driven architectures using Kafka, Event Hub, EventBridge, or Pub/Sub for real-time, asynchronous triggers.

5. Cloud FinOps and AI Cost Management

  • Establish and run the cloud FinOps practice based on FinOps Foundation principles (Inform, Optimize, Operate) and the FOCUS billing standard.

  • Implement mandatory tagging and cost allocation so every resource and AI workload maps to an owner, cost center, application, agent, workflow, or pipeline.

  • Deliver showback and chargeback reporting to business units and product teams.

  • Track cloud spend across compute, storage, networking, data egress, PaaS services, GPUs, and LLM token consumption.

  • Define unit economics such as cost per agent interaction, cost per workflow run, cost per pipeline execution, and cost per environment.

  • Drive infrastructure cost optimization through:

    • Rightsizing VMs, clusters, and databases

    • Reserved instances, savings plans, and committed use discounts

    • Spot and preemptible capacity for non-critical and batch workloads

    • Autoscaling, scale-to-zero, and scheduled shutdown of non-production environments

    • Storage tiering, lifecycle policies, and orphaned resource cleanup

    • GPU utilization improvement and workload consolidation

    • LLM cost controls such as model routing, caching, batch APIs, and provisioned throughput decisions

  • Set budgets, quotas, and anomaly alerts per subscription, team, and AI workload.

  • Produce forecasts and executive dashboards linking cloud and AI spend to business value.

  • Support Finance and Procurement on enterprise agreements, commitments, and vendor negotiations.

6. Cloud Security and Compliance

  • Design and enforce cloud security architecture based on zero-trust principles.

  • Own identity and access architecture (Entra ID, AWS IAM, GCP IAM), RBAC, privileged access management, managed identities, and workload identity for agents and services.

  • Implement secrets and key management (Key Vault, KMS, HashiCorp Vault), encryption at rest and in transit, and certificate lifecycle management.

7. Monitoring, Observability and Reliability

  • Architect enterprise monitoring and observability for infrastructure and AI workloads using Azure Monitor, CloudWatch, Google Cloud Operations, Datadog, Dynatrace, Prometheus/Grafana, and OpenTelemetry.

  • Monitor infrastructure health, performance, capacity, GPU utilization, network latency, and cost in unified dashboards.

  • Integrate AI-level telemetry (token usage, latency, failures) from tools such as Langfuse or LangSmith with infrastructure observability.

  • Apply SRE practices: error budgets, proactive alerting, chaos testing, and post-incident reviews.

8. Technical Leadership and Governance

  • Act as the design authority for cloud infrastructure, reviewing and approving infrastructure designs from project and product teams.

  • Define cloud standards, policies, and best practices, and drive consistent adoption across teams.

  • Lead and mentor cloud engineers, DevOps/platform engineers, and operations staff.

Required Skills and Experience

  • 12+ years in IT infrastructure, with at least 7 years of hands-on cloud infrastructure architecture and operations on Azure, AWS, or GCP (multi-cloud preferred).

  • Deep expertise in cloud networking, identity, compute, storage, and security services.

  • Strong hands-on experience with Kubernetes (AKS/EKS/GKE) and containers in production, including GPU node pools.

  • Advanced infrastructure-as-code skills (Terraform preferred) and CI/CD/GitOps practices.

  • Proven experience managing production cloud environments at enterprise scale, including incidents, DR, patching, and change management.

  • Demonstrated experience implementing cloud FinOps: tagging, allocation, optimization, and reporting.

  • Practical experience supporting AI/ML or GenAI workloads: LLM services, model serving, GPU infrastructure, vector databases, and data platforms.

  • Strong understanding of cloud security, zero trust, and compliance frameworks.

  • Scripting skills in Python, PowerShell, or Bash.

  • Experience in regulated industries; pharma or life sciences GxP experience is a strong advantage.

Preferred Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.

  • Certifications such as Azure Solutions Architect Expert, AWS Solutions Architect Professional, Google Professional Cloud Architect, CKA/CKAD, HashiCorp Terraform Associate, FinOps Certified Practitioner/Engineer, Azure AI Engineer, AWS ML Specialty, or ITIL.

  • Experience with hybrid infrastructure connecting cloud to on-premises data centers and manufacturing/OT environments.

  • Exposure to AI governance and Responsible AI frameworks.

Free · no signup

Get tomorrow's jobs before you have to search

Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.

Join WhatsAppJoin Telegram

No spam. Just jobs and resources.

Why people use ASAI

Someone shared one job with you. ASAI keeps finding the rest.

  • Scored, not searched. Every role ranked against your actual profile.

  • Alerts as often as hourly. Reach new roles while the pile is still small.

  • Skill gaps, spelled out. See exactly which requirements you don't meet yet.

  • Verified jobs, only. Say no to ghost jobs. Your time deserves respect.

Keep browsing

All open roles at Zydus Group

Two ways in

Applicants who checked fit first are 3.1× more likely to hear back

Your match scoreCalculated · locked
86Overall
64Skills
97Experience

Your score for this role already exists

ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.

No credit card · 1 tap with Google

Free · no signup

Get tomorrow's jobs before you have to search

Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.

Join WhatsAppJoin Telegram

No spam. Just jobs and resources.

Why people use ASAI

Someone shared one job with you. ASAI keeps finding the rest.

  • Scored, not searched. Every role ranked against your actual profile.

  • Alerts as often as hourly. Reach new roles while the pile is still small.

  • Skill gaps, spelled out. See exactly which requirements you don't meet yet.

  • Verified jobs, only. Say no to ghost jobs. Your time deserves respect.

ASAI job platform logo

A job platform finally, balanced in your favour.

Jobs by CityJobs by CompanyGuidesHow We VerifyAboutPrivacy PolicyTerms of Service

Built in India 🇮🇳

© 2026 ASAI. All rights reserved.