AI Infrastructure Architect (AWS AI Services | AWS Cloud Architecture)
Role description
Role Overview
We are seeking a highly experienced AI Infrastructure Architect to design, build, and govern scalable, secure, resilient, and cost-efficient AI platforms across AWS, Microsoft Azure, and Google Cloud Platform. The role will lead end-to-end architecture for AI/ML workloads, including data platforms, model training, fine-tuning, inference, MLOps, GPU infrastructure, observability, security, and governance for enterprise and production-grade use cases.
The ideal candidate combines deep cloud infrastructure expertise, hands-on AI/ML platform knowledge, and strong enterprise architecture leadership across multi-cloud environments.
Key Responsibilities AI/ML Infrastructure Architecture
Lead the design of end-to-end AI infrastructure for model experimentation, training, fine-tuning, inference, deployment, monitoring, and ongoing operations.
Architect scalable platforms for batch and real-time ML workloads, LLM-based solutions, Generative AI pipelines, and enterprise AI applications.
Define standards for model experimentation, versioning, registry, promotion, lifecycle management, and retirement.
Create reusable reference architectures, blueprints, guardrails, and design patterns for AI workloads.
Ensure platforms meet scalability, availability, performance, disaster recovery, and operational requirements.
- Multi-Cloud Platform Design
- AWS, Azure & GCP
Architect cloud-native and cloud-agnostic AI platforms across AWS SageMaker, EKS, EC2 GPU, S3 and IAM; Azure Machine Learning, AKS, Azure OpenAI and GPU VM series; and GCP Vertex AI, GKE and TPU/GPU infrastructure.
Define workload placement principles based on capability, security, latency, resilience, portability, cost, and strategic vendor alignment.
Enable workload portability and standardized operating practices across cloud environments.
Define hybrid and multi-cloud AI operating models, including connectivity, identity, observability, governance, and disaster recovery.
MLOps, DevOps & Platform Engineering
Establish MLOps frameworks for CI/CD and continuous training of models, pipelines, features, and AI applications.
Design automation for model lifecycle management, experiment tracking, model registry, validation, deployment, rollback, and retraining.
Implement monitoring for service health, model performance, drift, data quality, latency, throughput, reliability, and cost.
Integrate AI delivery pipelines with enterprise DevOps, platform engineering, security, and change-management standards.
Data & Compute Architecture
Design scalable data ingestion, feature stores, training datasets, data lakes, and data access patterns for AI/ML workloads.
Architect accelerator strategies using NVIDIA GPUs, TPUs, and fit-for-purpose inference compute.
Optimize utilization, performance, scheduling, capacity, and cost for training and inference workloads.
Define storage, networking, caching, and distributed-compute patterns for large-scale AI platforms.
Security, Governance & Compliance
Define AI security architecture covering identity, privileged access, data access, network isolation, secrets and key management, encryption, supply-chain security, and tenant/workload isolation.
Implement governance controls for model usage, data privacy, lineage, approvals, responsible AI, risk management, and compliance.
Align AI platforms with enterprise security architecture, regulatory obligations, audit requirements, and internal governance frameworks.
Embed security-by-design, policy-as-code, traceability, and evidence collection into platform workflows.
Leadership & Advisory
Act as the technical authority for AI infrastructure and platform architecture decisions.
Guide cloud architects, platform engineers, data engineers, ML engineers, security teams, and application teams.
Support AI platform roadmaps, cloud strategy, capability assessments, investment decisions, and architecture reviews.
Communicate architectural choices, trade-offs, risks, and recommendations to business, engineering, and leadership stakeholders.
Mentor teams and promote reusable engineering practices and architecture standards.
Core Technical Skills Cloud Platforms & Architecture
Advanced architecture expertise across AWS, Microsoft Azure, and GCP.
Strong experience in cloud networking, IAM, security architecture, landing zones, resilience, and multi-cloud governance.
AI/ML Platforms
Hands-on experience designing and deploying enterprise AI/ML infrastructure.
Expertise with Azure Machine Learning, AWS SageMaker, and GCP Vertex AI.
Experience with Generative AI and LLM platforms supporting training, fine-tuning, evaluation, inference, and monitoring.
Infrastructure & Platform Engineering
Kubernetes expertise across EKS, AKS, and GKE.
GPU/accelerator infrastructure architecture, scheduling, performance tuning, capacity management, and cost optimization.
Infrastructure as Code using Terraform, ARM/Bicep and/or CloudFormation.
Containerization, platform automation, service mesh, networking, and observability.
MLOps & Automation
CI/CD and continuous training for ML pipelines and AI applications.
Model registry, experiment tracking, feature/pipeline versioning, deployment automation, and inference scaling.
Monitoring, logging, ing, drift detection, reliability engineering, and performance tuning.
Data Systems
Large-scale data platforms for AI/ML workloads, including batch and streaming architectures.
Feature stores, data ingestion, data quality, lineage, governance, and secure data-access patterns.
Strong understanding of distributed systems and high-performance computing concepts.
Preferred Qualifications
Experience designing and governing enterprise AI platforms at scale.
Exposure to Responsible AI frameworks, model risk management, and AI governance operating models.
Strong background in cost optimization and FinOps for GPU-intensive AI workloads.
Consulting, client
Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →