![Jagveer Singh](https://jobicy.com/react/themes/app/images/avatar.jpg)
Jagveer Singh

# SRE Manager | Data Platform Infrastructure

Actively looking · Member since 22 Jul 2026 [Message](https://jobicy.com/enter4715.php.md)

ShareLocationSingapore, SingaporeDesired salaryUnspecifiedWork preferenceRemote Only / Full TimeExperience levelLead
* [Overview](#overview)
* [Portfolio 0](#portfolio)
* [Services 0](#services)

## About

Professional summaryI am an SRE and infrastructure leader with more than 15 years of experience managing enterprise-scale data platforms, distributed systems, cloud infrastructure, and production reliability operations.

I spent nine years supporting UBS programs and have led the migration of more than 5,000 servers across storage, application, and cloud environments. I have consistently delivered 99.9% or higher SLO performance, reduced critical incident MTTR by 40%, and maintained strong change success rates across complex production environments.

My technical expertise includes Kafka, Spark, Airflow, Kubernetes, Azure, AWS, GCP, Terraform, Python, and observability platforms such as Prometheus, Grafana, Splunk, CloudWatch, and Azure Monitor. I design SLOs, error budgets, disaster recovery strategies, and automated operational processes for real-time and batch analytics platforms.

I am an experienced incident commander with strong ITIL v4 practices, including RCA completion within 48 hours, recurrence prevention, production readiness reviews, and governance for high-risk changes. I have also supported GDPR and MAS TRM compliance, encryption controls, and hybrid cloud security hardening.

I have hired and developed SRE talent, led cross-functional teams of more than 55 people across six APAC countries, managed annual infrastructure budgets exceeding $2 million, and created structured career development programs that supported multiple promotions.

I am passionate about reliable, automated, and well-governed engineering practices. I have driven the adoption of AI-assisted development tools with human-in-the-loop controls, quality guardrails, and review standards designed to protect production reliability.

## Skills

39 capabilities[Airflow](https://jobicy.com/talent/airflow.md)[Ansible](https://jobicy.com/talent/ansible.md)[Application Security](https://jobicy.com/talent/application-security.md)[Automation](https://jobicy.com/talent/automation.md)[AWS](https://jobicy.com/talent/aws.md)[Azure](https://jobicy.com/talent/azure.md)[BASH](https://jobicy.com/talent/bash.md)[Budget Management](https://jobicy.com/talent/budget-management.md)[Capacity Planning](https://jobicy.com/talent/capacity-planning.md)[Change Management](https://jobicy.com/talent/change-management.md)[Cloud Infrastructure](https://jobicy.com/talent/cloud-infrastructure.md)[Compliance](https://jobicy.com/talent/compliance.md)[Data Integrity](https://jobicy.com/talent/data-integrity.md)[Disaster Recovery](https://jobicy.com/talent/disaster-recovery.md)[Distributed Systems](https://jobicy.com/talent/distributed-systems.md)[Error Budgets](https://jobicy.com/talent/error-budgets.md)[GCP](https://jobicy.com/talent/gcp.md)[GDPR](https://jobicy.com/talent/gdpr.md)[GitOps](https://jobicy.com/talent/gitops.md)[Grafana](https://jobicy.com/talent/grafana.md)[Incident Response](https://jobicy.com/talent/incident-response.md)[Infrastructure as Code](https://jobicy.com/talent/infrastructure-as-code.md)[Kafka](https://jobicy.com/talent/kafka.md)[Kubernetes](https://jobicy.com/talent/kubernetes.md)[Linux Administration](https://jobicy.com/talent/linux-administration.md)[Mentoring](https://jobicy.com/talent/mentoring.md)[Observability](https://jobicy.com/talent/observability.md)[People Management](https://jobicy.com/talent/people-management.md)[Perl](https://jobicy.com/talent/perl.md)[PowerShell](https://jobicy.com/talent/powershell.md)[Prometheus](https://jobicy.com/talent/prometheus.md)[Python](https://jobicy.com/talent/python.md)[Root Cause Analysis](https://jobicy.com/talent/root-cause-analysis.md)[Site Reliability Engineering](https://jobicy.com/talent/site-reliability-engineering.md)[Spark](https://jobicy.com/talent/spark.md)[SQL](https://jobicy.com/talent/sql.md)[Team Leadership](https://jobicy.com/talent/team-leadership.md)[Terraform](https://jobicy.com/talent/terraform.md)[Virtualization](https://jobicy.com/talent/virtualization.md)

## Tech stack & tools

Working toolkit

### Analytics

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/grafana/grafana-original.svg) Grafana](https://jobicy.com/talent?tool=grafana.md)

### Application Hosting

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/azure/azure-original.svg) Azure](https://jobicy.com/talent?tool=azure.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/centos/centos-original.svg) CentOS](https://jobicy.com/talent?tool=centos.md)

### Collaboration

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/jira/jira-original.svg) Jira](https://jobicy.com/talent?tool=jira.md)

### Data Stores

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/cassandra/cassandra-original.svg) Cassandra](https://jobicy.com/talent?tool=cassandra.md)

### Development

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/ansible/ansible-original.svg) Ansible](https://jobicy.com/talent?tool=ansible.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/bash/bash-original.svg) Bash](https://jobicy.com/talent?tool=bash.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/git/git-original.svg) Git](https://jobicy.com/talent?tool=git.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/helm/helm-original.svg) Helm](https://jobicy.com/talent?tool=helm.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/kubernetes/kubernetes-original.svg) Kubernetes](https://jobicy.com/talent?tool=kubernetes.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/powershell/powershell-original.svg) PowerShell](https://jobicy.com/talent?tool=powershell.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/terraform/terraform-original.svg) Terraform](https://jobicy.com/talent?tool=terraform.md)

### Languages & Frameworks

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/hadoop/hadoop-original.svg) Apache Hadoop](https://jobicy.com/talent?tool=hadoop.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/perl/perl-original.svg) Perl](https://jobicy.com/talent?tool=perl.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/python/python-original.svg) Python](https://jobicy.com/talent?tool=python.md)

### Monitoring

[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/prometheus/prometheus-original.svg) Prometheus](https://jobicy.com/talent?tool=prometheus.md)[![Image](https://cdn.jsdelivr.net/gh/devicons/devicon@master/icons/splunk/splunk-original.svg) Splunk](https://jobicy.com/talent?tool=splunk.md)

## Experience

Career history

### SRE Manager · HCL Technologies

May 2021 – Mar 2026Architected and governed SLOs for more than 3,000 servers supporting enterprise data warehouse analytics workloads, achieving a 99.9% success rate, four-hour RTO, and 15-minute RPO. Operated Kafka clusters, monitored Spark workloads on Azure Databricks and Hadoop, and managed Airflow DAGs across hybrid infrastructure.

Reduced critical incident MTTR by 40% through Prometheus and Grafana alerting and automated remediation runbooks. Led ITIL v4 P1 incident response, RCA, recurrence prevention, Kubernetes monitoring, and hybrid cloud observability. Automated migration and failover processes with Python and Kubernetes tooling, reducing cutover time by 35%.

Hired and onboarded eight SRE engineers, led cross-functional teams of more than 55 people across six APAC countries, managed a $2 million annual infrastructure budget, and established technical development and AI-assisted engineering governance programs.

### SRE – Application Migration · Hays

Aug 2018 – May 2021Led SLO definition and reliability engineering for more than 1,500 business-critical application migrations across three countries, achieving 99.95% availability and less than 30 minutes of downtime. Executed IaaS, PaaS, and SaaS migration strategies with SQL database transformations and CI/CD integration.

Validated post-migration performance, including sub-150ms 95th percentile latency and error rates below 0.1%. Troubleshot Kafka consumer lag and Spark streaming failures and configured Azure Application Insights for real-time monitoring.

Managed schema and stored procedure migrations for more than 200 OLTP databases, supported Airflow ETL cutover validation, and developed Python reconciliation scripts. Ensured GDPR and MAS TRM compliance through encryption, Azure Key Vault controls, TLS 1.3, and security hardening reviews.

### SRE – Storage & Unix Infrastructure · Cognizant

Sep 2017 – Jul 2018Orchestrated more than 500 TB of EMC-to-Hitachi storage migrations while maintaining 100% data integrity and zero corruption events. Established validation procedures for ZFS, VxVM, fstab, and multipath configurations and maintained service continuity during live storage cutovers.

Performed readiness assessments on more than 500 Linux and Solaris servers and developed Bash and Perl automation for storage provisioning, LUN mapping, and multipath validation. Implemented Nagios and SCOM monitoring for storage arrays and Unix hosts.

Created capacity forecasting models and maintained runbooks for storage failover procedures across six APAC locations, preventing potential storage exhaustion incidents.

### Senior Systems Engineer · Digital Minds Software Solutions

2016 – 2017Provided Tier 3 technical support and administration for Active Directory, Exchange, Office 365, Citrix, VMware, and enterprise systems. Troubleshot complex infrastructure issues and supported reliable operations for business users.

### Senior Systems Engineer · Cognizant Technology Solutions

2014 – 2016Provided Level 2 support for Active Directory, DHCP, DNS, TCP/IP, VMware, email systems, and virtual desktops in a 24/5 global operations environment. Supported infrastructure troubleshooting and service restoration.

### Technical Support Specialist · Dell International Services

2011 – 2014Provided technical support for Microsoft Office and Outlook issues and supported Windows 8 and Windows 8.1 deployment activities. Assisted users with troubleshooting and enterprise desktop operations.

## Education

Learning history

### JNTU-GNEC

2009Bachelor of Technology, Electronics & Communication Engineering

Bachelor of Technology degree in Electronics and Communication Engineering.

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.