Vaishak T Nair
Vaishak T Nair

Site Reliability Engineer / DevOps Engineer

Open to offers · Member since 1 Oct 2026
Location
Chengannur, India
Desired salary
Confidential
Work preference
Remote Only / Full Time
Experience level
Mid

About

Professional summary

I am a Site Reliability Engineer and DevOps Engineer with more than six years of total IT experience, including over three years focused on architecting and owning cloud infrastructure end to end. I build reliable, scalable, and self-service platforms for fintech, healthcare SaaS, and enterprise GitOps products.

I have deep hands-on experience with AWS, GCP, Azure, Kubernetes, Terraform, Docker, Helm, ArgoCD, CI/CD, and multi-region production environments. I have led zero-downtime Kubernetes migrations and platform modernization programs supporting more than 60 microservices across four global regions.

I currently lead a team of four DevOps/SRE engineers and own reliability, escalation management, technical reviews, and customer-success delivery for enterprise platform customers. I have resolved critical P0/P1 incidents, established on-call processes, and supported customers from proof of concept through production onboarding.

I specialize in observability and incident management, having built monitoring, logging, tracing, alerting, and escalation systems using SigNoz, OpenTelemetry, Prometheus, VictoriaMetrics, Grafana, Datadog, PagerDuty, and Zenduty. My work has reduced false-positive alerts, deployment failures, and incident response times.

I also focus on cloud cost governance, security, and compliance. Through Karpenter optimization, compute right-sizing, network redesign, Kubernetes RBAC, Vault, IAM, and network policies, I have reduced cloud spending by 21–23% while supporting regulatory and audit requirements.

I enjoy turning ambiguous infrastructure challenges into resilient platforms, collaborating closely with engineering, product, customers, and leadership to improve reliability, developer experience, and business outcomes.

Notice period: 30 days

Skills

32 capabilities

Tech stack & tools

Working toolkit

Experience

Career history

Lead Site Reliability Engineer / DevOps Engineer One2N Consulting

Worked across multiple client engagements, progressing from an individual-contributor SRE to leading DevOps/SRE delivery. Owned cloud infrastructure, reliability, GitOps processes, observability, incident management, customer onboarding, security, and cost optimization for fintech, healthcare SaaS, and enterprise platforms.

Led technical direction for platform reliability and customer success, supporting multi-cloud Kubernetes environments and large-scale production workloads. Delivered zero-downtime migrations, compliance-aligned access controls, monitoring platforms, and infrastructure modernization programs.

SRE & Customer Success Engineering Lead One2N Consulting (Devtron project)

Lead and manage a team of four DevOps/SRE engineers delivering reliability and customer-success engineering for an enterprise CD/GitOps platform serving 24 enterprise customers. Own technical reviews, escalation management, and resolution of more than 20 P0/P1 incidents across over eight client accounts within SLA.

Own the customer-success lifecycle from POC through production onboarding, including workload migrations and readiness validation. Led onboarding of 56 microservices for an automotive marketplace, delivered EKS/GKE/AKS upgrades, established RBI-aligned RBAC, Vault, and network-policy controls, and operated VictoriaMetrics, VictoriaLogs, Fluent Bit, Prometheus, and Zenduty alerting.

DevOps-SRE One2N Consulting (Neurowyzr project)

Acted as the sole embedded SRE for a healthcare SaaS platform operating more than 60 microservices across four global AWS production regions. Rebuilt outdated infrastructure from scratch with Terraform, including networking and environment-segregated infrastructure repositories.

Architected and led an eight-month, zero-downtime migration to parallel EKS clusters using ArgoCD GitOps blue-green synchronization and Route 53 DNS cutover. Built SigNoz and OpenTelemetry observability, reduced false-positive alerts by approximately 30%, improved release readiness, reduced deployment failures by approximately 20%, and implemented PagerDuty escalation policies.

DevOps-SRE One2N Consulting (Kulu project)

Built the observability foundation for a fintech payments platform using Helm- and CI/CD-managed Datadog alerting and synthetic monitoring across AWS services. Implemented OpenTelemetry log and trace correlation to address production monitoring gaps for application and infrastructure alerts.

Coordinated production releases across more than 10 microservices, managing environment configuration, database migrations, and release readiness. Resolved over 200 infrastructure support requests, introduced SonarQube quality gates, and led Karpenter, compute right-sizing, and NAT Gateway optimization that reduced monthly cloud costs by 21–23% per environment.

DevOps Engineer Datavizz

Migrated CI workloads from Jenkins to GitHub Actions, automated Docker image builds, and reduced build times by 30%. Delivered a zero-downtime production rollout using blue-green deployment practices with SSO and role-based access control.

Optimized a separate AWS CodePipeline workflow, packaged deployment and configuration assets as Helm charts, deployed applications to AWS EKS, and managed development and test environments.

DevOps Engineer LSN Software Services Pvt Ltd

Led the design and implementation of an on-premises Kubernetes cluster for the Equitybrix and Intellasphere projects. Deployed HashiCorp Vault to provide secure secrets management for platform workloads.

Built Jenkins CI pipelines integrated with ArgoCD continuous delivery, deployed Prometheus and Grafana monitoring, and managed database infrastructure through Terraform.

Software Engineer Inapp Information Technologies

Worked as a backend developer on the VSSC (Vikram Sarabhai Space Centre) project using Java and Spring Boot. Contributed to application development and backend service implementation.

Automated API performance testing using JMeter and Katalon, improving test execution and performance validation processes.

Education

Learning history

Mar Baselios College of Engineering and Technology, Trivandrum

Bachelor of Technology, Computer Science

Army Public School, Trivandrum

Higher Secondary School

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

Jobs Talent AI Tools Salaries
Menu