We’ve launched our self-serve ads platform — use promo code HELLO10 and get a free $10 credit ›

Senior Site Reliability Engineer

Remote from
Poland flag
Poland
Annual salary
Undisclosed
Salary information is not provided for this position. Check our Salary Directory to estimate the average compensation for similar roles.
Employment type
Full Time,
Job posted
Apply before
17 Jun 2026
Experience level
Senior
Views / Applies
14 / 5

About Cribl

Cribl is the observability pipeline company, enabling IT and Security teams to parse, route, and shape any data, before it’s analyzed, at any scale.

Verified job posting
This job post has been manually reviewed for authenticity and compliance.

AI Summary

Cribl, a fast-growing company building telemetry infrastructure for AI, seeks a Senior Site Reliability Engineer to join its remote team in Poland. The role involves ensuring reliability, scalability, and observability of cloud-based platforms, with involvement from design to production. Candidates need experience with observability tools, cloud platforms (AWS/Azure), IaC (Terraform), and JavaScript/Node.js. The position requires on-call duties and offers a chance to impact a transformative technology.

Job Complexity

Easy Hard
AI Insight The role requires deep expertise in cloud infrastructure, observability, and automation, plus on-call responsibilities, making it challenging but not the highest level.

Salary Analysis

Median
$160,000
US Market
$130,000 – $200,000
AI Insight The offered salary is not specified, but for a Senior SRE role in Poland, the market range is $130k-$200k USD (adjusted for location). The median of $160k is competitive for the role's seniority and technical demands.

Key Skills

Site Reliability Engineering Cloud Infrastructure Observability Terraform AWS Azure Kubernetes Prometheus Grafana JavaScript/Node.js

Dear Hiring Manager,

I am excited to apply for the Senior Site Reliability Engineer position at Cribl. With extensive experience in designing and operating observability systems for cloud platforms, I am passionate about ensuring reliability and performance at scale. My background includes deep expertise in AWS, Azure, Terraform, and monitoring tools like Prometheus and Grafana, aligning perfectly with your requirements.

I thrive in collaborative, blameless environments and have a strong drive to automate and reduce toil. I am eager to contribute to Cribl's mission of empowering customers with control over their telemetry data. Thank you for considering my application.

Sincerely,
[Your Name]

Describe your experience designing and implementing observability systems for a complex cloud-based platform. What tools and strategies did you use?
I worked on a multi-cloud platform where we used Prometheus for metrics aggregation, Grafana for dashboards, and Elasticsearch for logging. We implemented distributed tracing with Jaeger. Key strategies included defining SLOs, setting up alerting based on error budgets, and using Terraform to manage infrastructure as code.
How do you approach incident response in a blameless environment? Can you give an example?
In a prior role, we had a major outage due to a misconfigured load balancer. We conducted a blameless postmortem focusing on system improvements rather than individual mistakes. We automated deployment checks and added canary deployments to catch similar issues early. This culture encourages learning and reduces fear of making mistakes.
Explain how you would reduce toil in a production environment. Provide a specific automation example.
To reduce toil, I automated database failover procedures using a combination of health checks and Ansible playbooks. Previously, manual failover took 30 minutes; after automation, it took under 2 minutes. I also implemented self-healing scripts for common issues like disk space alerts.
What is your experience with container orchestration and how do you ensure high availability?
I have extensive experience with Kubernetes, including managing clusters on AWS EKS. For high availability, I use pod anti-affinity, horizontal pod autoscaling, and multi-AZ deployments. I also implement readiness and liveness probes to ensure traffic is only routed to healthy pods.
How do you stay current with emerging technologies in cloud and observability?
I regularly follow industry blogs like The New Stack and attend conferences like KubeCon. I also contribute to open-source projects like OpenTelemetry and experiment with new tools in my home lab. Continuous learning is key in this fast-evolving field.

Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the world’s biggest enterprises, including half of the Fortune 100, to bridge the gap between AI ambition and infrastructure reality. As the AI Platform for Telemetry, we give customers the choice, control, and flexibility to manage and analyze telemetry for both humans and agents, so they can build what’s next.

We’re one of the fastest‑growing private companies and a leading player in a massive, fast‑moving market. With a global workforce, we’re remote‑first and grounded in a simple idea: software is a people business. Cribl is the place where curious, collaborative people can do their best work, grow fast, and bring their full selves to the herd.

Why You’ll Love This Role

Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in Poland. Cribl provides users a new level of observability, intelligence and control over their real-time data. You will join a team of technical engineers who are committed to shipping only high-quality software and enjoying all the goat gifs the internet has to offer. This role is remote within Poland, and you will be part of the engineering organization where you will contribute in our efforts to envision, create, deploy, test, and ship Cribl products. 

Not often do you get to be part of something that is fundamentally changing a technology. But here at Cribl we are building the next generation of software that puts our customers in full control of their observability data. If this is something that interests you, and you want to be truly at the center of the wheel helping make this work better every day, then this opportunity might be something you have been waiting for to be a part of making a real impact. 

Fixing things at the operational side should always be the last resort, so our SREs are involved from conception to design to development and all the way through production and beyond. You provide your creative input into all things Cloud, Scaling, Reliability, High Availability and much more.

If reliability is your passion, and you have always had strong opinions on how to make things better and have the desire to build consensus around ideas, then let’s talk!

As An Active Member Of Our Team, You Will…

  • Engage with teams and improve service delivery and reliability across their entire lifecycle
  • Measure and monitor all production systems with an eye towards availability, latency and overall system health
  • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
  • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability 
  • Help identify and drive down toil with creative innovation and automation
  • This position will require stand-by, on-call, or off-hours duties

If You’ve Got It – We Want It

  • Proven experience designing, implementing, and operating observability systems for complex cloud-based platforms, with deep knowledge of best practices and a strong drive to implement them leveraging Cribl products.
  • Experience with Configuration Management and Infrastructure as a Code Tools like Terraform (preferred) or Ansible. Experience working with Cloud SDKs is also a plus.
  • Knowledge of cloud platforms (prefer AWS and Azure) and container + orchestration technologies.
  • Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
  • Extensive experience with enterprise scale continuous delivery environments.
  • Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment.
  • Experience with sustainable incident response in a blameless environment.
  • Background in Linux Systems Engineering.
  • Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.
  • Comfortable with a high level of autonomy and working with a distributed team.
  • Knowledge of Cloud and application security best practices.
  • Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
  • A love for high quality and a knack for testing.
  • Opinions about business metrics, and SLOs.

#LI-GV1
#LI-Remote

Bring Your Whole Self

Diversity drives innovation, enables better decisions to support our customers, and inspires change for the better. We’re building a culture where differences are valued and welcomed, and we work together to bring out the best in each other. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which the candidate is applying.

Interested in joining the Cribl herd? Learn more about the smartest, funniest, most passionate goats you’ll ever meet at cribl.io/about-us

Apply now >

Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

How to apply

Did you apply? Let us know, and we’ll help you track your application.

See a few more

Similar DevOps & Infrastructure remote jobs

Job Search Safety Tips

Here are some tips to help you search and apply for jobs safely:
Watch out for suspicious jobs Don't apply for jobs that offer high pay for little work or offer to hire you without an interview. Read more ›
Check the employer's profile Make sure you're applying for a trustworthy job by visiting the employer's profile and learning more about them. Read more ›
Protect your information Don't share personal details like your bank account or government-issued ID on suspicious websites or messengers. Read more ›
Report jobs that feel unsafe If you see a job that seems misleading, inappropriate or discriminatory, report it for going against our policies and we'll review it.

Share this job

Jobicy+ Subscription

Jobicy

614 professionals pay to access exclusive and experimental features on Jobicy

Free

USD $0/month

For people just getting started

  • • Unlimited applies and searches
  • • Access on web and mobile apps
  • • Weekly job alerts
  • • Access to additional tools like Bookmarks, Applications, and more

Plus

USD $8/month

Everything in Free, and:

  • • Ad-free experience
  • • Daily job alerts
  • • Personal career consultant
  • • AI-powered job advice
  • • Featured & Pinned Resume
  • • Custom Resume URL
Go to account ›