Krishnaveni Bijjam
Krishnaveni Bijjam

Site Reliability Engineer

Actively looking · Member since 31 Aug 2026
Message
Location
Bengaluru, India
Desired salary
Unspecified
Work preference
Remote Only
Experience level
Mid

About

Professional summary

I am a Site Reliability Engineer with over four years of experience supporting production systems, cloud-native applications, and enterprise services.

I specialize in improving infrastructure reliability through proactive monitoring, incident management, release validation, and operational automation.

I have hands-on experience with Kubernetes, Docker, Jenkins, Terraform, AWS, Azure, Linux, Python, and Bash.

I build and manage observability solutions using Datadog, Splunk, Grafana, and CloudWatch, including dashboards, alerts, log analysis, and incident troubleshooting.

I work closely with development, QA, architecture, and operations teams to resolve application, API, and infrastructure issues while maintaining service-level commitments.

I am experienced in 24/7 on-call support, major incident management, root-cause analysis, deployment rollbacks, and post-deployment validation, with a focus on reducing MTTR and sustaining high availability.

Skills

17 capabilities

Tech stack & tools

Working toolkit

Analytics

Application Hosting

Collaboration

Languages & Frameworks

Monitoring

Experience

Career history

Site Reliability Engineer LTIMindtree

I provide Site Reliability Engineering and production support services for Allstate India Pvt Ltd. I support production deployments, release validation, rollback activities, health checks, smoke testing, and non-production application monitoring for Identity, Profile, and Payment services. I also provide 24/7 on-call support and major incident management, resolving incidents within SLA timelines and collaborating with cross-functional teams to troubleshoot application, API, and infrastructure issues.

I develop and manage observability solutions using Datadog and Splunk, including dashboards, alerts, log analysis, SPL queries, application health monitoring, incident troubleshooting, and root-cause analysis. I manage Jenkins-based deployments and scheduled redeployments, validate deployments through Gatekeeper, and partner with QA and architecture teams to improve release readiness. My monitoring and alerting improvements reduced downtime by 25%, supported faster incident response and MTTR reduction, and helped maintain 99.9% uptime for critical enterprise systems.

Education

Learning history
No education data available.

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

People also viewed

All talent ›
Jobs Talent AI Tools Salaries
Menu