I am a Senior Site Reliability and Cloud Infrastructure Engineer with more than 14 years of experience building resilient, highly available enterprise platforms. I specialize in multi-cloud architecture, automated Day-2 operations, reliability engineering, and infrastructure automation across AWS, Google Cloud Platform, and Microsoft Azure.
I design declarative, maintainable infrastructure using Terraform, Ansible, CloudFormation, OpenTofu, and Pulumi. My work focuses on eliminating operational toil, improving deployment reliability, and creating scalable self-service platform capabilities for engineering teams.
I have deep experience with Kubernetes, including Amazon EKS and Google Kubernetes Engine, along with Docker, Helm, Kustomize, ArgoCD, and service-mesh technologies. I have supported platforms serving up to 50,000 concurrent users while maintaining availability targets as high as 99.99%.
I build observability and incident-response practices using Prometheus, Grafana, Datadog, OpenSearch, and PagerDuty. I am focused on reducing MTTR through actionable telemetry, automated rollback processes, disaster-recovery procedures, and reliable runbooks.
I have delivered secure CI/CD and internal developer platform solutions using GitLab CI/CD, Harness CD, GitHub Actions, Jenkins, Backstage, and cloud-native APIs. I also mentor cross-functional engineering teams on SRE principles, production readiness, and automated disaster recovery.
My technical background includes Python, Go, JavaScript/TypeScript, Bash, REST API design, cloud networking, IAM, zero-trust security, and VMware virtualization. I am motivated by building robust systems that make software delivery safer, faster, and easier for developers.