I am a Site Reliability Engineer and DevOps Engineer with more than six years of total IT experience, including over three years focused on architecting and owning cloud infrastructure end to end. I build reliable, scalable, and self-service platforms for fintech, healthcare SaaS, and enterprise GitOps products.
I have deep hands-on experience with AWS, GCP, Azure, Kubernetes, Terraform, Docker, Helm, ArgoCD, CI/CD, and multi-region production environments. I have led zero-downtime Kubernetes migrations and platform modernization programs supporting more than 60 microservices across four global regions.
I currently lead a team of four DevOps/SRE engineers and own reliability, escalation management, technical reviews, and customer-success delivery for enterprise platform customers. I have resolved critical P0/P1 incidents, established on-call processes, and supported customers from proof of concept through production onboarding.
I specialize in observability and incident management, having built monitoring, logging, tracing, alerting, and escalation systems using SigNoz, OpenTelemetry, Prometheus, VictoriaMetrics, Grafana, Datadog, PagerDuty, and Zenduty. My work has reduced false-positive alerts, deployment failures, and incident response times.
I also focus on cloud cost governance, security, and compliance. Through Karpenter optimization, compute right-sizing, network redesign, Kubernetes RBAC, Vault, IAM, and network policies, I have reduced cloud spending by 21–23% while supporting regulatory and audit requirements.
I enjoy turning ambiguous infrastructure challenges into resilient platforms, collaborating closely with engineering, product, customers, and leadership to improve reliability, developer experience, and business outcomes.