[All remote jobs](https://jobicy.com/jobs.md)[![CentralReach logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/07/a0eac9f542eb8c3fff7c19990dd5c1bb.jpg)](https://jobicy.com/company/centralreach.md)[CentralReach](https://jobicy.com/company/centralreach.md)

# Sr. Site Reliability Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/centralreach.md)ShareRemote from[USA](https://jobicy.com/job-region/usa.md)SalaryUSD 160k–180k / yrDepartment[DevOps & Infrastructure](https://jobicy.com/categories/admin.md)EmploymentFull TimeExperienceOpen levelPublished1 Oct 2026Apply before31 Oct 2026Listing views42Application actions1Application toolkit

## Make your next move.

Prepare your resume, explore your fit, and draft a cover letter for this opportunity.

AI Summary

## The role, at a glance.

CentralReach is seeking a senior Site Reliability Engineer to own and improve reliability across its public and private cloud platforms. The role focuses on availability, latency, capacity planning, observability, incident response, SLOs/SLIs, error budgets, and automation to reduce operational toil. The engineer will partner closely with software engineering teams to improve release readiness, architecture, and operational practices. Required expertise spans AWS and cloud-native infrastructure, Kubernetes, Helm, CI/CD, Linux and Windows systems, networking, and observability tooling including Datadog, Prometheus, Grafana, Splunk, and OpenTelemetry. This is a senior, hands-on reliability role in a fast-moving platform engineering environment.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

4/5IndependentCollaborative

AI insightThis role requires deep cross-functional expertise in production operations, cloud infrastructure, automation, software development, and reliability engineering. The engineer is expected to independently lead reliability improvements and respond effectively to high-impact production incidents.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate$170,000US market range$150k–$195k0$215k

AI insightThe disclosed base salary range is $160,000 to $180,000 USD yearly, with a midpoint of $170,000. This is competitive for a senior US-based Site Reliability Engineer; a typical US market range is approximately $150,000 to $195,000 annually, depending on location, cloud expertise, on-call scope, and total compensation.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Site Reliability Engineering](https://jobicy.com/jobs?search_keywords=Site%20Reliability%20Engineering.md)[AWS](https://jobicy.com/jobs?search_keywords=AWS.md)[Kubernetes](https://jobicy.com/jobs?search_keywords=Kubernetes.md)[Helm](https://jobicy.com/jobs?search_keywords=Helm.md)[Observability](https://jobicy.com/jobs?search_keywords=Observability.md)[Datadog](https://jobicy.com/jobs?search_keywords=Datadog.md)[Prometheus](https://jobicy.com/jobs?search_keywords=Prometheus.md)[Grafana](https://jobicy.com/jobs?search_keywords=Grafana.md)[OpenTelemetry](https://jobicy.com/jobs?search_keywords=OpenTelemetry.md)[CI/CD](https://jobicy.com/jobs?search_keywords=CICD.md)

Sample interview questionsHow have you used SLOs, SLIs, and error budgets to improve service reliability?I begin by identifying user-critical journeys and defining measurable SLIs such as availability, latency, and successful request rate. I set SLO targets with product and engineering stakeholders, then use error-budget consumption to guide release decisions, prioritize reliability work, and trigger investigation before service performance materially degrades.

Describe a significant production incident you managed and the steps you took to restore service.

I first establish incident ownership, scope, customer impact, and a clear communication cadence. I use dashboards, logs, traces, and recent-change analysis to isolate the failure, apply the lowest-risk mitigation such as rollback, scaling, or traffic control, and verify recovery against service metrics. Afterward, I lead a blameless retrospective with specific corrective actions, owners, and due dates.

How would you build an observability strategy for a distributed cloud-native platform?

I would standardize instrumentation for logs, metrics, and traces through OpenTelemetry and define consistent service names, correlation IDs, and meaningful labels. I would create dashboards and actionable alerts tied to SLOs rather than noisy infrastructure thresholds, then use traces and structured logs to accelerate root-cause analysis. The approach would include retention, cardinality, cost, and access-control standards.

What steps would you take to reduce operational toil for engineering teams?

I would quantify repetitive manual work and prioritize tasks with high frequency, production risk, or developer impact. Typical improvements include self-service deployment workflows, infrastructure-as-code, automated remediation, standardized runbooks, capacity forecasting, and CI/CD guardrails. I would measure success through reduced manual effort, faster recovery, fewer failed changes, and improved deployment velocity.

How do you ensure Kubernetes-based services are ready for production?

I validate resource requests and limits, autoscaling, health probes, disruption budgets, secure configuration, and rollback procedures. I also confirm that services have ownership, SLOs, dashboards, alerts, runbooks, dependency mapping, load testing, and incident-response expectations before launch. This operational-readiness process makes reliability a shared responsibility between platform and application engineering.

Opportunity details

## About this role.

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy providers, educators, and employers to scale the way they deliver ABA and related therapies with innovative technology, market-leading industry expertise, and world-class customer satisfaction.

The Platform Engineering group at CentralReach builds the underlying technologies that power our Public and Private Cloud Platforms worldwide. The group is responsible for storage, data infrastructure, IT, observability systems, DevOps, SRE, provisioning, compute, orchestration platform, internal tools, internal platforms (laptops, networks, systems etc.) and services – all the components that make up the CentralReach Platform.

If you have a passion for the future, enjoy and thrive in an agile, fast-moving, ever-changing startup environment, welcome and take on technical challenges of all shapes and sizes, have excellent interpersonal skill and sense of humor and enjoy rolling up your sleeves and jumping in, then read on!

As a Sr. SRE, you will work closely with the key stakeholders in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts, incident retrospectives, chaos testing, and end-to-end ownership.

Key Accountabilities:

* Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments.

* Define, maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices.

* Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance.

* Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns.

* Reduce toil and increase development velocity through automation and continuous improvement.

* Provide production support, including incident, change, and problem management; root cause analysis; service restoration; runbooks; and standard operating procedures.

* Identify data-driven opportunities to improve system architecture, availability, performance, and reliability.

* Collaborate with software engineering teams on release management, roadmap planning, and operational readiness.

* Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana.

Desired Skills and Experience:

* Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry.

* Experience implementing observability strategies for logs, metrics, and traces.

* Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo.

* Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts.

* Strong understanding of containerization technologies, including Kubernetes and Helm.

* Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development.

* Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts.

* Experience using AI to improve productivity and amplify technical skills.

Base Salary Range

$160,000—$180,000 USD

Backed by Roper Technologies, Inc. (Nasdaq: ROP), CentralReach is entering an exciting phase of growth, innovation, and scale.

Recognized as one of the best places to work over 10 times by organizations such as Inc, Built In, and NJBIZ, our culture is centered around impact, inclusion, and flexibility. As a hybrid company with collaborative offices in Ft. Lauderdale, FL; Holmdel, NJ; and Verona, Italy, we foster a workplace where top talent can thrive and make a real difference in the lives of those we serve.

We offer competitive compensation, comprehensive health benefits, generous PTO, 401(k) matching, and paid parental leave to our full-time employees. Our team members also enjoy hybrid work schedules, career development support, wellness programs, and opportunities to give back through CR Cares™, our community engagement initiative.

Be part of a market leader driving the future of care. Explore opportunities at [centralreach.com/careers](https://nam12.safelinks.protection.outlook.com/?url=https%3A%2F%2Fcentralreach.com%2Fcareers%2F&data=05%7C02%7Cjamie.madarasz%40centralreach.com%7Cb7b1c0d4e2354127669a08dda36f3df3%7C515bdea7d91847b3b73f5700d21635ec%7C0%7C0%7C638846420463140236%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=WNerHqhY5w8%2Fh6TrRUbAh6uyPQJS8ymWAWaKuyJfaHo%3D&reserved=0).

Protecting your information is important to us. Please take a moment to review our Notice of Privacy Practices for Job Applicants. [Applicant-Privacy-Notice-110725](https://centralreach365.sharepoint.com/:b:/s/ApplicantPrivacyNotice/IQBl-WqOdm5oQI9tZ2avE2_7AYJGJ5xXuv-ubvZ989Cf9tk?e=KdZc3m) to understand how we collect, use, store, and protect your personal information during the recruitment process.

Show more

[Apply now >](https://jobicy.com/jobs/154372-sr-site-reliability-engineer.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

*
![CentralReach logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/07/a0eac9f542eb8c3fff7c19990dd5c1bb.jpg)
CentralReach Oct 1  New

### [Principal Platform Engineer](https://jobicy.com/jobs/154365-principal-platform-engineer.md)

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy…

*
![CentralReach logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/07/a0eac9f542eb8c3fff7c19990dd5c1bb.jpg)
CentralReach Oct 1  New

### [Site Reliability Engineer (SRE), Data Products](https://jobicy.com/jobs/154368-site-reliability-engineer-sre-data-products.md)

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy…

*
![Datadog logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/c1654fc7-221.jpg)
Datadog Oct 1  New

### [Developer Advocate – Service Management EMEA](https://jobicy.com/jobs/154340-developer-advocate-service-management-emea.md)

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run…

*
![Mural logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/01/9b6f1383-221.jpeg)
Mural Oct 1  New

### [Software Engineer, Platform Engineering](https://jobicy.com/jobs/154333-software-engineer-platform-engineering.md)

ABOUT THE TEAM The Platform Engineering team builds the tools and services that accelerate the software development life cycle of the software engineers at Mural. We work on tools that…

*
![GT logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/6f6103a3-221.jpeg)
GT Oct 1  New

### [Azure DevOps Engineer | KD Pharma](https://jobicy.com/jobs/152146-azure-devops-engineer-kd-pharma.md)

GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in…

*
![TITAN logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/81b82aeb-221.png)
TITAN Sep 30  New

### [Senior Network Engineer](https://jobicy.com/jobs/154242-senior-network-engineer.md)

Available shifts: Sat–Wed, 3 PM – 12 AM PHT Sat–Wed, 11 PM – 8 AM PHT Senior Network Engineer The Senior Network Engineer will interact directly with our high-profile, elite…

*
![Legion logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/99d2bb45-221-1.jpg)
Legion Sep 30  New

### [Principal Software Engineer, DevOps](https://jobicy.com/jobs/154248-principal-software-engineer-devops.md)

Principal Software Engineer, DevOps Remote, United States ABOUT POSITION: Are you passionate about automation, DevOps, public cloud services, Kubernetes, and observability? As a site reliability engineer at Legion, you will…

*
![Okta logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/d54f5eb5-221.png)
Okta Sep 30  New

### [Manager, Site Reliability Engineering (Auth0)](https://jobicy.com/jobs/152068-manager-site-reliability-engineering-auth0.md)

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to…

*
![Kraken logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/03/Jobicy-210307081105-509016.png)
Kraken Sep 30  New

### [Senior Database Administrator – Core Infrastructure](https://jobicy.com/jobs/152041-senior-database-administrator-core-infrastructure.md)

Building the Future of Open FinancePayward – the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services and CF Benchmarks – has spent the last 15 years building one of…

*
![Truelogic logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/e7ae6cb6-221-1.png)
Truelogic Sep 30  New

### [Adobe Workfront Implementation System Architect – Advertising](https://jobicy.com/jobs/152013-adobe-workfront-implementation-system-architect-advertising.md)

About TruelogicAt Truelogic we are a leading provider of nearshore staff augmentation services headquartered in New York. For over two decades, we’ve been delivering top-tier technology solutions to companies of…

[Browse all jobs](https://jobicy.com/jobs.md)