SRE Manager HCL Technologies
Architected and governed SLOs for more than 3,000 servers supporting enterprise data warehouse analytics workloads, achieving a 99.9% success rate, four-hour RTO, and 15-minute RPO. Operated Kafka clusters, monitored Spark workloads on Azure Databricks and Hadoop, and managed Airflow DAGs across hybrid infrastructure.
Reduced critical incident MTTR by 40% through Prometheus and Grafana alerting and automated remediation runbooks. Led ITIL v4 P1 incident response, RCA, recurrence prevention, Kubernetes monitoring, and hybrid cloud observability. Automated migration and failover processes with Python and Kubernetes tooling, reducing cutover time by 35%.
Hired and onboarded eight SRE engineers, led cross-functional teams of more than 55 people across six APAC countries, managed a $2 million annual infrastructure budget, and established technical development and AI-assisted engineering governance programs.