Skip to content
Open to opportunities · Berlin or remote

Karthik Orugonda

Senior Platform Engineer & SRE

15+ years in IT, including 10+ years building and operating cloud-native platforms — primarily AWS, with Azure and GCP experience — where observability and reliability are the foundation, not an afterthought: SLOs, incident response, and golden-signal dashboards cutting MTTR by ~30%. I lead platform initiatives across engineering teams, mentoring engineers on reliability practices and turning shared infrastructure — CI/CD, GitOps, self-service golden paths — into the developer experience that lets teams ship without waiting on a ticket.

10+
Years cloud native
400+
Engineers served
~30%
MTTR reduction

Tech Skills

The stack I build platforms with

Cloud Platforms

  • AWS
    IAM/IRSA·Multi-account
  • Azure
    AKS·Entra ID·Key Vault
  • GCP
    GKE·Cloud IAM

Containers & Orchestration

  • Kubernetes
    Multi-tenancy·Operators·RBAC
  • Docker
  • Helm
  • Kustomize
  • Istio

Infrastructure as Code

  • Terraform
    Reusable modules·Multi-env·TfLint
  • Terragrunt
    Multi-env·OPA gates
  • Ansible

GitOps

  • Argo CD
    ApplicationSets

CI/CD

  • Jenkins
  • GitHub Actions
  • GitLab CI

Observability & Reliability

  • Grafana
  • Prometheus
  • OpenTelemetry
  • Loki
  • Tempo
  • Datadog
  • Dynatrace

Security & Governance

  • OPA Gatekeeper
  • Kyverno
  • External Secrets
    AWS Secrets Manager·Azure Key Vault
  • Vulnerability Scanning & Remediation
    SAST·DAST·Trivy·Checkov

Data & Messaging

  • SQL
  • Redis
  • RabbitMQ

Software Engineering

  • Python
    Platform APIs·Operators·Automation
  • Bash
    Automation·Operational tooling
  • Go
    CLIs·actively deepening
  • Java / Groovy
    Pipeline libraries·Groovy DSL

Linux & Networking

  • Linux
    Administration·Troubleshooting
  • Networking
    DNS·TCP/IP·TLS·Load balancing
  • VPC & subnet design
    Cross-account·connectivity

AI & GPU Infrastructure

  • NVIDIA GPU Operator
    Device plugin·Time slicing·GPU metrics
  • Agentic coding workflows
    Claude Code·Copilot·Cursor·MCP
  • LLM serving
    Ollama·llama.cpp·FastAPI gateway

Projects

What I built, and the decisions behind it

Each page documents what was chosen and what it was chosen over.

More projects

Experience

15 years building infrastructure and the teams that run it

Platform, reliability and delivery ownership across retail, e-commerce and telecom.

  1. Dec 2022 – Present

    Current

    Senior Platform Engineer & SRE

    Aldi DX

    Reliability and observability across multiple engineering departments, on a Kubernetes internal developer platform

    • Build and own the internal developer platform on Kubernetes, including GitOps workflows and application CI/CD pipelines supporting diverse services, with reusable Helm charts adopted across multiple development teams for Aldi’s multi-country e-commerce.
    • Led the strategic migration of the org-wide observability stack from New Relic to OpenTelemetry, standardising metrics, logs and traces, trace-to-log correlation and sampling, while reducing annual vendor licensing costs by ~40% and eliminating vendor lock-in.
    • Delivered observability through GitOps, versioning dashboards, SLOs, alerts and error budgets, integrating AIOps event correlation for automated RCA, impact analysis, cutting false-positive alerts by ~30%.
    • Built reusable Terraform modules, CI/CD pipelines and governance frameworks for shared cloud infrastructure, integrating policy-as-code, security scanning, drift detection and automated remediation.
    • Mentored platform and application engineers through design reviews and documented golden paths, enabling self-service adoption across teams.
    • Brought agentic coding tools (GitHub Copilot, Claude Code) into daily platform work, accelerating delivery of IaC modules and GitOps workflows.

    Platform Engineering · SRE · Kubernetes · Terraform · GitOps · OpenTelemetry · Observability · Golden Paths

  2. May 2018 – Nov 2022

    Technical Lead — DevOps, Cloud & Platform

    Rakuten

    Multi-tenant platform and delivery tooling for 400+ engineers

    • Spearheaded multi-cloud and platform migrations from legacy platforms to Kubernetes-based IDP, Azure and GCP, leading a team of 5 engineers, defining architecture and driving platform modernization and cost optimisation.
    • Owned platform capacity planning and scaling activities for major sales events, preparing platform infrastructure for massive traffic surges and maintaining reliability under peak demand.
    • Operated and evolved multi-tenant Kubernetes platforms providing centralized compute, cluster capabilities and developer self-service for 400+ engineers across multiple business domains.
    • Standardized CI/CD across 60+ teams by refactoring Jenkins shared libraries into reusable components and establishing GitOps deployment workflows for canary and blue-green releases, with RBAC and multi-tenant Kubernetes patterns supporting secure team isolation.
    • Built centralized DevSecOps pipelines integrating security and code analysis tools across the organization.

    Platform Engineering · Kubernetes · Helm · Azure · GCP · Private Cloud · Security Automation

  3. Sep 2015 – Apr 2018

    IT Operations Lead & DevOps Engineer

    Hewlett Packard Enterprise

    Production operations for Vodafone UK's e-commerce platform

    • Designed and maintained build and release pipelines for legacy monolithic applications using Jenkins and Puppet, automating workflows across on-prem infrastructure.
    • Steered production operations for Vodafone UK’s e-commerce platform, managing a 25-member team, owning incident lifecycle (detection → RCA → resolution) and driving MTTR reduction through post-incident improvements.

    Production Operations · Infrastructure Automation · Incident Management · Release Engineering

  4. Dec 2010 – Aug 2015

    Senior Software Engineer

    Tech Mahindra

    Backend and payment systems for Vodafone UK

    • Supported backend and payment gateway systems for Vodafone UK using WebLogic and Linux infrastructure, including onsite operations at Vodafone UK HQ.
    • Developed automated reporting tools reducing manual effort by 40%, improving operational efficiency.

    Backend Systems · Payment Gateways · High Availability · Linux

Certifications & Education

Qualifications

Certifications

  • CKA

    Certified Kubernetes Administrator

  • CKAD

    Certified Kubernetes Application Developer

Education

Bachelor of Engineering in Information Technology

SVEC, affiliated to JNT University, India

Graduated 2010

Languages

  • EnglishFluent (C2)
  • GermanBeginner (A1)

Recommendations

What colleagues say

View on LinkedIn
“Karthik always been a great leader with vast expertise on different set of tools and technologies. His leadership skills is one of the best I have met in corporate. His dedication to his craft is nothing short of inspiring, and his ability to coach/guide is incredible. Always helpful and brings idea for the solution to the problem in a optimal way. He's a great asset in DevOps world.”

Rohan Das

Site Reliability Engineer

Karthik was senior to Rohan, not his direct manager

“Karthik is an excellent DevOps engineer and a wonderful human being. His expertise as a DevOps engineer is considerable, and it helped our team come up with more efficient solutions on different projects. He is creative, smart, has excellent technical skills. I love his positive attitude and he is my go-to person when I need advice. I would recommend and endorse Karthik.”

Sachin Rajput

Principal Engineer, DevOps

Worked with Karthik on the same team

“During my career I've rarely come across real professionals like Karthik. I've learned so much from working with Karthik, he helped me to become a better professional. Such a great human being with vast knowledge in areas of DevOps and cloud. Highest recommendations to Karthik, not just as a thoughtful leader but as a team player as well.”

Arun Babu

DevOps Engineer

Reported to Karthik directly