[![Image]() Meet Jobicy Copilot — free AI autofill for job applications + remote job alerts ›](#)   [All remote jobs](https://jobicy.com/jobs.md)Open role[![Playson logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/6f266d9c-221-1.png)](https://jobicy.com/company/playson.md)Remote opportunity at[Playson](https://jobicy.com/company/playson.md)

# Senior Site Reliability Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/playson.md)Share14 Sep 2026Published63Listing views2Application actions14 Oct 2026Apply before  Opportunity details

## About this role.

AI SummaryThis Senior Site Reliability Engineer role owns reliability, performance, and operational stability for a high-traffic platform handling approximately 5–7k requests per second. The engineer will manage production incidents, participate in a 24/7 on-call rotation, conduct root-cause analysis, and implement durable corrective actions. Core technical work includes Kubernetes on AWS EKS, Terraform, Helm, GitOps with FluxCD or ArgoCD, CI/CD, and observability tooling. The position is highly hands-on and requires independent judgment during time-sensitive production events. Close collaboration with engineering teams is essential to deploy safely and reduce customer impact.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

4/5IndependentCollaborative

AI insightThe role requires senior-level infrastructure expertise across cloud, Kubernetes, automation, observability, and incident management in a continuously operating, high-load environment. A 24/7 on-call rotation and ownership of real-time production decisions make the operational pressure and technical complexity especially high.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate$165,000US market range$140k–$200k0$220k

AI insightNo salary range is disclosed; "Competitive Salary" is not a quantifiable compensation offer. Estimated US-market annual base salary for a Senior Site Reliability Engineer with Kubernetes, AWS, Terraform, GitOps, and high-scale on-call ownership is approximately $140,000–$200,000 USD, with a midpoint estimate of $165,000 USD. Actual compensation may differ based on location, employment arrangement, bonus structure, and local market.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Site Reliability Engineering](https://jobicy.com/jobs?search_keywords=Site%20Reliability%20Engineering.md)[Kubernetes](https://jobicy.com/jobs?search_keywords=Kubernetes.md)[AWS EKS](https://jobicy.com/jobs?search_keywords=AWS%20EKS.md)[Terraform](https://jobicy.com/jobs?search_keywords=Terraform.md)[GitOps](https://jobicy.com/jobs?search_keywords=GitOps.md)[FluxCD](https://jobicy.com/jobs?search_keywords=FluxCD.md)[ArgoCD](https://jobicy.com/jobs?search_keywords=ArgoCD.md)[CI/CD](https://jobicy.com/jobs?search_keywords=CICD.md)[Observability](https://jobicy.com/jobs?search_keywords=Observability.md)[Incident Response](https://jobicy.com/jobs?search_keywords=Incident%20Response.md)

Sample interview questionsHow would you investigate and mitigate a sudden rise in latency on a Kubernetes-based production service?I would first assess user impact and inspect key signals such as request latency, error rate, saturation, pod health, node capacity, and recent deployments. I would correlate metrics, logs, traces, and Kubernetes events to isolate whether the cause is application behavior, dependency latency, networking, resource contention, or infrastructure failure. I would mitigate safely through actions such as rollback, scaling, traffic controls, or dependency protection, then document the incident and drive a root-cause fix.

Describe a meaningful improvement you would make to an immature alerting setup.

I would begin by mapping alerts to customer-impacting service-level indicators and defining clear severity, ownership, and runbooks. I would remove noisy symptom alerts, add actionable alerts for error-budget burn, availability, latency, and saturation, and ensure each notification provides diagnostic context. The goal is to reduce alert fatigue while ensuring responders are paged only for issues requiring immediate human action.

How do Terraform and GitOps work together in a reliable infrastructure delivery workflow?

Terraform is well suited to declaratively provision cloud infrastructure such as networks, IAM, EKS, and managed services, while GitOps continuously reconciles Kubernetes workloads and configuration from version-controlled repositories. I would use pull requests, automated validation, environment promotion, reviewed plans, and controlled reconciliation to make changes auditable and reversible. State protection, least-privilege access, and drift detection are also essential.

What should a high-quality production incident postmortem include?

A strong postmortem is blameless and includes the impact, timeline, detection method, contributing technical and process factors, mitigation steps, root cause, and clearly owned follow-up actions. It should distinguish between immediate remediation and longer-term prevention work. The output should improve monitoring, runbooks, architecture, testing, or operational processes rather than merely document the event.

How would you prepare for a 24/7 on-call rotation supporting a high-load system?

I would ensure services have clear ownership, escalation paths, dashboards, runbooks, and tested rollback procedures before relying on on-call responders. I would prioritize automation for common recovery tasks and regularly review alert quality, incident trends, and operational load. During incidents, I would communicate concise status updates, stabilize the service first, and schedule follow-up work to prevent recurrence.

About the Role

We’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad – a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.

You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale.

If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions – this is the place for you.

Key Responsibilities

*

Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time

*

Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment

*

Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence

*

Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem

*

Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)

*

Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues

*

Maintain and evolve CI/CD pipelines and infrastructure-as-code practices

*

Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment

*

Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance

*

Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load

Requirements

*

Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments

*

Experience with GitOps tools such as FluxCD or ArgoCD

*

Proven experience in incident response, root cause analysis, and postmortems in production systems

*

Solid experience with AWS, Terraform, Docker, and CI/CD pipelines

*

Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch

*

Strong understanding of networking concepts and protocols

*

Proficiency in at least one scripting language (e.g. Python, Go, Node.js)

*

Experience working with version control systems (Git)

*

Familiarity with incident management tools like PagerDuty, Opsgenie, or similar

*

Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability

*

Proactive, resilient mindset with a focus on continuous improvement and system stability

What We Offer

*

Competitive Salary

*

Quarterly Bonuses

*

Unlimited Paid Time Off

*

Unlimited Paid Sick Leave

*

Remote & Flexible Working

*

Private Medical Insurance

*

Financial Support for Life Events

*

Professional Development Budget

*

International Exposure

*

Regular Company Events

*Benefits may vary depending on location and contractual agreement

Recruitment Process

1. HR Interview (30-45 min)

2. Technical interview (90 min)

4. Final Interview with C-level (60 min)

By submitting your application, you acknowledge that your personal data will be processed in accordance with our [Privacy Policy](https://playson.com/privacy).

Show more

[Apply now >](https://jobicy.com/jobs/153228-senior-site-reliability-engineer-3.md)

>  Annual salary information is not provided for this position. Explore salary ranges for similar roles in our [Salary Directory ›](https://jobicy.com/salaries.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

Matched by job category10 related opportunities[DevOps & Infrastructure](https://jobicy.com/categories/admin.md) [Browse all jobs](https://jobicy.com/jobs.md)
*
![Nebius logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/d90c0566-221.webp)
Nebius  Sep 14

### [Senior System Engineer (Virtual Private Cloud Team)](https://jobicy.com/jobs/150590-senior-system-engineer-virtual-private-cloud-team.md)

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from…

*
![Phantom logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/3d0a3a49-221.png)
Phantom  Sep 12

### [Staff Software Engineer (SRE)](https://jobicy.com/jobs/153134-staff-software-engineer-sre.md)

Phantom is on a mission to connect the world to the freedom of open markets. Tens of millions of people all over the world use Phantom to access global markets…

*
![Phantom logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/3d0a3a49-221.png)
Phantom  Sep 12

### [Staff DevOps Engineer (Platform)](https://jobicy.com/jobs/153130-staff-devops-engineer-platform.md)

Phantom is on a mission to connect the world to the freedom of open markets. Tens of millions of people all over the world use Phantom to access global markets…

*
![Nebius logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/d90c0566-221.webp)
Nebius  Sep 12

### [Senior Site Reliability Engineer (Hardware Automation)](https://jobicy.com/jobs/149057-senior-site-reliability-engineer-hardware-automation.md)

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from…

*
![Webflow logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/01/b75fa7e9-221.jpg)
Webflow  Sep 12

### [Senior Infrastructure Engineer](https://jobicy.com/jobs/153105-senior-infrastructure-engineer-3.md)

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences that drive predictable growth and strengthen brand trust. Building at…

*
![Rithum logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/fe3986d8-221-1.jpeg)
Rithum  Sep 12

### [IT Automation Engineer – Business Technology (Central/Mountain Time – US)](https://jobicy.com/jobs/153100-it-automation-engineer-business-technology-central-mountain-time-us.md)

Rithum™ is the world’s most trusted commerce network, accelerating how brands, suppliers, and retailers work together to deliver seamless e-commerce experiences. We provide an unmatched platform for brands and retailers,…

*
![Remote logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/63a88d3a-221-1.png)
Remote  Sep 11

### [Engineering Manager, SRE](https://jobicy.com/jobs/153081-engineering-manager-sre.md)

About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage…

*
![Ashby logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/63864ee6-221.png)
Ashby  Sep 11

### [Staff Platform Engineer – UK](https://jobicy.com/jobs/153062-staff-platform-engineer-uk.md)

We’re looking for a curious, rigorous, problem-hungry platform software engineer (who codes!) to carry the ball as we bring Ashby to the big leagues. Ashby builds software that lets talent…

*
![ExtraHop logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/4800f505-221.jpg)
ExtraHop  Sep 11

### [Manager, Engineering | ML Infrastructure & Tooling](https://jobicy.com/jobs/153008-manager-engineering-ml-infrastructure-tooling.md)

At ExtraHop, we’re on a mission to protect and empower the connected enterprise. We reveal what is happening in the very infrastructure that sustains businesses, lives, and communities, and ensure…

*
![Cloudbeds logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/3aef1a38-221.png)
Cloudbeds  Sep 10

### [DevOps Engineer](https://jobicy.com/jobs/152990-devops-engineer-4.md)

What Makes Cloudbeds UniqueAt Cloudbeds, we’re not just building software, we’re transforming hospitality. Our intelligently designed platform powers properties across 150 countries, processing billions in bookings annually. From independent properties…