Karthik Orugonda
Senior Platform Engineer & SRE
Kubernetes · Multi-cloud · IDPs · Observability
10+ years building cloud-native platforms across AWS, Azure and GCP. I build Kubernetes-based internal developer platforms, Terraform-driven self-service infrastructure, GitOps delivery and the observability that keeps them reliable.
Tech Skills
The stack I build platforms with
Containers & Orchestration
KubernetesCKA + CKAD
HelmLibrary charts, OCI distribution
Docker
Kubernetes Operators
KarpenterNode autoscaling, GPU NodePools
Istio
IaC & GitOps
Terraform
Terragrunt
Argo CD
JenkinsShared libraries for 150+ teams
Ansible
GitHub Actions
GitLab CI
Observability & Reliability
OpenTelemetryCollector · Gateway · vendor-neutral
Dynatrace
Prometheus
Grafana
Loki / Tempo
- Incident ResponseOn-call · blameless postmortems · MTTR reduction
Datadog
Security & Governance
OPA Gatekeeper
Kyverno
- External SecretsAWS Secrets Manager · Azure Key Vault
- Kubernetes RBAC & Network Policies
- IaC / container scanningSAST · DAST · image scanning
Cloud Platforms
AWSEKS, VPC, IAM/IRSA
AzureAKS, Entra ID, Key Vault
GCPGKE, Cloud IAM
Linux & Networking
- LinuxAdministration · production troubleshooting
- NetworkingDNS · TCP/IP · TLS · load balancing · ingress
- VPC & subnet designCross-account connectivity
Software Engineering
PythonPlatform APIs, operators, automation
Bash
GoCLIs — actively deepening
Java / GroovyJenkins pipeline libraries
AI & GPU Infrastructure
NVIDIA GPU OperatorDevice plugin · time slicing · GPU metrics
- Agentic coding workflowsClaude Code · Copilot · Cursor
- LLM servingOllama · llama.cpp · FastAPI gateway
Projects
What I built, and the decisions behind it
5 documented decisions
OpenTelemetry & LGTM Platform
Telemetry ends up coupled to whichever vendor was chosen first — instrumentation is rewritten every time the backend changes, and metrics, logs and traces stay in three unconnected tools.
OpenTelemetry
LGTM Stack
Prometheus
Grafana
Loki
Tempo
8 documented decisions
IDP & GitOps Reference Architecture
Onboarding a new service means a ticket, a wait, and a platform engineer hand-assembling the same manifests again — the platform team becomes the bottleneck for every team it serves.
IDP
GitOps
Argo CD
Kubernetes
Terraform
Go
6 documented decisions
Enterprise AWS Infrastructure
Multi-environment Terraform drifts: dev and prod diverge through copy-paste, and nothing stops a change that is syntactically valid but violates security or blows the budget.
Terragrunt
Terraform
AWS
OPA/Conftest
GitHub Actions
More projects
4 documented decisions
AI Infrastructure on Amazon EKSProduction-style AI infrastructure on Amazon EKS for provisioning, sharing and observing NVIDIA GPUs — Karpenter, GPU Operator, CUDA workloads and DCGM observability.
7 documented decisions
FinOps Kubernetes OperatorKubernetes operator that scales non-production workloads to zero during declared sleep windows, with per-workload exclusions.
Experience
15 years building infrastructure and the teams that run it
Dec 2022 – Present
Senior Platform Engineer & SRE
Aldi Süd
Platform and reliability engineering for a Kubernetes-based internal developer platform
- Set the deployment standard used across multiple engineering teams — reusable Terraform modules, Helm library charts and GitOps golden paths that replaced per-team bespoke pipelines on the department's Kubernetes-based internal developer platform.
- Owned the observability practice end to end: OpenTelemetry and Dynatrace with alerting-as-code and SLO frameworks, cutting MTTR and false-positive alerts by ~30%.
- Brought agentic coding tools (GitHub Copilot, Claude Code) into daily platform work, accelerating delivery of IaC modules and GitOps workflows.
- Platform Engineering
- SRE
- Kubernetes
- GitOps
- Observability
- Golden Paths
May 2018 – Nov 2022
Technical Lead — DevOps, Cloud & Platform
Rakuten
Multi-tenant platform and delivery tooling for 400+ engineers
- Owned the CI contract for 150+ teams — refactored the shared Jenkins libraries every team built on, and moved delivery onto GitOps pipelines with automated canary and blue-green rollouts.
- Ran multi-tenant Kubernetes platforms and CI/CD systems serving 400+ engineers across multiple business domains.
- Led the migration off legacy PaaS (Mesos/Marathon) to Kubernetes, and then to private cloud, across multiple business units — directing a team of 5 engineers.
- Platform Engineering
- Kubernetes
- Helm
- Azure
- GCP
- Private Cloud
- Security Automation
Sep 2015 – Apr 2018
IT Operations Lead & DevOps Engineer
Hewlett Packard Enterprise
Production operations for high-traffic enterprise e-commerce
- Led a 25-engineer production operations team for high-traffic enterprise e-commerce platforms.
- Replaced manual release runbooks with automated pipelines and infrastructure provisioning workflows.
- Production Operations
- Infrastructure Automation
- Incident Management
- Release Engineering
Dec 2010 – Aug 2015
Senior Software Engineer
Tech Mahindra
Backend systems for Vodafone UK
- Built automated reporting tooling for Vodafone UK backend systems, cutting manual operational effort by ~30%.
- Backend Systems
- High Availability
- Linux
Certifications & Education
Qualifications
Certifications

CKA
Certified Kubernetes Administrator

CKAD
Certified Kubernetes Application Developer
Education
Bachelor of Engineering in Information Technology
SVEC, affiliated to JNT University, India
Graduated 2010
Languages
- EnglishFluent (C2)
- GermanBeginner (A1)