Senior Platform Engineer Amazon
Architected and governed multi-region, multi-account Kubernetes platform serving as Internal Developer Platform for 30+ teams at 99.99% availability. Built 4-account AWS landing zone (Organizations/Control Tower/SCPs) in Terraform, cutting provisioning from 5 days to 2 hours (95% reduction) and standardizing it enterprise-wide. Implemented multi-region EKS cluster/workload account separation, minimizing blast radius and simplifying IAM policy management across 30+ production, staging, and dev teams. Built GitOps platform with ArgoCD, Helm, and blue/green deployments, enabling self-service and increasing deployment frequency 3× (bi-weekly to daily) while cutting incidents 70%, mentored teams on adoption. Replaced static AWS credentials with IRSA for per-service-account least-privilege access, improving security posture and enabling full CloudTrail auditability for SOC 2 Type II. Enforced DevSecOps via policy-as-code (OPA + Checkov + tfsec) in CI pipelines, blocking 100% of non-compliant changes and reducing high/critical CVEs by 80% through Trivy scanning and KMS encryption. Engineered comprehensive SLI/SLO observability stack (Prometheus + Grafana + Loki + OpenTelemetry), decreasing MTTR 35% (85 to 55 minutes) with proactive anomaly detection. Drove FinOps program (Reserved Instances + Spot + rightsizing), delivering $1.2M annual savings (25–30% reduction) with zero impact to reliability or performance, including cost allocation and chargeback models. Established EKS upgrade governance (disruption budgets + staged rollouts + PDBs), achieving zero-downtime upgrades across all clusters in line with Well-Architected reliability practices. Standardized secure pod-to-AWS access patterns using IRSA and RAM subnet sharing, enabling least-privilege workload isolation in multi-account environments.