Muhammad Hay
Muhammad Hay

Lead DevOps / SRE Engineer

Open to offers · Member since 14 Sep 2026
Location
Waterbury, United States
Desired salary
Unspecified
Work preference
Remote Only
Experience level
Lead

About

Professional summary

I am a Site Reliability and DevOps Engineer with more than eight years of hands-on experience building reliable, scalable, and automated cloud platforms. I specialize in AWS, Kubernetes, infrastructure as code, CI/CD, GitOps, observability, and production operations.

I design and operate highly available cloud infrastructure using AWS services including EKS, VPC, IAM, EC2, RDS, S3, load balancers, and Auto Scaling. I focus on creating secure, efficient platforms that enable teams to deploy and operate workloads confidently.

I build repeatable infrastructure and configuration workflows with Terraform, Ansible, Helm, and Kustomize. I have modernized delivery processes through GitHub Actions, GitLab CI/CD, Jenkins, Argo CD, and GitOps practices, improving release consistency and rollback capabilities.

I am experienced in reliability engineering, including SLI/SLO management, incident response, root cause analysis, capacity planning, disaster recovery, and MTTR/MTTD improvement. I implement comprehensive observability using Prometheus, Grafana, OpenTelemetry, Alertmanager, CloudWatch, and ELK.

I also apply DevSecOps principles throughout delivery and infrastructure workflows. My experience includes container security, vulnerability scanning with Trivy, code-quality controls with SonarQube, secrets management with HashiCorp Vault, cloud security controls, and AWS cost optimization.

I enjoy simplifying complex operational problems, improving engineering workflows, and building self-service platform capabilities. I bring strong Linux administration, scripting, troubleshooting, and cross-functional platform engineering expertise to production environments.

Skills

32 capabilities

Tech stack & tools

Working toolkit

Experience

Career history

Lead DevOps / SRE Engineer StopShopREI LLC

Led the design and evolution of AWS cloud infrastructure and Kubernetes platforms for highly available production workloads. Used Amazon EKS, VPC, IAM, EC2, RDS, S3, load balancers, and Auto Scaling; improved platform scalability with workload scheduling, HPA, Karpenter, RBAC, Ingress, and resource management.

Standardized multi-environment infrastructure using Terraform modules and Ansible. Modernized delivery with GitHub Actions, GitLab CI/CD, Argo CD, Helm, and Kustomize, while establishing observability through Prometheus, Grafana, OpenTelemetry, Alertmanager, and CloudWatch. Applied SRE and DevSecOps practices including SLI/SLOs, incident response, disaster recovery, Trivy, SonarQube, HashiCorp Vault, security controls, and FinOps optimization.

Senior DevOps / SRE Engineer Botco.ai

Designed and maintained Amazon EKS Kubernetes environments, using Docker and Helm to manage scalable workloads across development, staging, and production. Built CI/CD pipelines with Jenkins, GitLab CI/CD, and GitHub Actions to automate builds, testing, containerization, and deployment.

Provisioned AWS resources with Terraform and Ansible, including EC2, VPC, IAM, S3, and RDS. Implemented monitoring with Prometheus, Grafana, AWS CloudWatch, and ELK; established Argo CD GitOps workflows; and strengthened reliability and security through SLI/SLO tracking, incident response, capacity planning, SonarQube, Trivy, and HashiCorp Vault.

DevOps Engineer VibeLogics

Built and maintained Jenkins and Git-based CI/CD pipelines to automate application builds, testing, and deployments. Supported AWS-based application infrastructure, including Linux servers, networking components, Nginx configurations, and application runtime environments.

Containerized applications with Docker and integrated container workflows into Jenkins pipelines. Automated provisioning with Terraform and Ansible, and developed Bash and Python scripts for deployments, system checks, log processing, and Linux administration. Implemented foundational monitoring, logging, troubleshooting, and incident investigation practices to improve availability.

Education

Learning history
No education data available.

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

People also viewed

All talent ›
Jobs Talent AI Tools Salaries
Menu