Skip to content
All projects

AI Infrastructure on Amazon EKS

Production-style AI infrastructure on Amazon EKS for provisioning, sharing and observing NVIDIA GPUs — Karpenter, GPU Operator, CUDA workloads and DCGM observability.

Source code
  • GPU Operator
  • Karpenter
  • CUDA
  • Time Slicing
  • Observability

The problem

GPU capacity is expensive, scarce and easy to strand: nodes sit idle between jobs, a single workload can hold a whole accelerator, and standard Kubernetes gives you almost no visibility into what the GPU is actually doing.

Constraints

Architecture

Observability

Grafana
Prometheus
DCGM Exporter

Driver & device lifecycle

GPU Operator
Device Plugin
Node Feature Discovery
Container Toolkit
Time Slicing

Capacity

GPU Node — provisioned by Karpenter
GPU Workloads (CUDA)

Key decisions

What was chosen, what it was chosen over, and why.