All remote jobs
Open role
Remote opportunity atGT

Site Reliability Engineer (SRE) | Feeld

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
51Listing views
3Application actions
30 Sep 2026Apply before
Opportunity details

About this role.

AI Summary

This is a senior Site Reliability Engineer role embedded in a small, autonomous product squad supporting Feeld's consumer mobile platform. The engineer will own observability for critical user journeys, develop SLIs/SLOs, improve monitoring and alerting, and lead initial response to severe production incidents. The role requires a strong backend foundation in Node.js and TypeScript alongside hands-on AWS, incident-management, logging, metrics, and tracing experience. It is a distributed remote position available from Poland, Spain, or the United Kingdom, with participation in out-of-hours incident coverage.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

4/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThis is a high-seniority reliability role because it combines deep application-level backend expertise with production operations ownership. The engineer must independently interpret signals, mitigate P0/P1 incidents, and coordinate effectively across distributed teams during potentially urgent situations.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$180,000
US market range$150k–$215k
AI insightNo actual salary was disclosed in the posting. This is an estimated US-market annual base-salary range in USD for a senior Site Reliability Engineer with Node.js/TypeScript, AWS, observability, and incident-leadership responsibilities; actual compensation may vary by location, employment arrangement, and total-rewards structure.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
How would you define SLIs and SLOs for a critical user journey in a consumer mobile application?

I would first map the journey and identify the user-visible outcomes, such as successful authentication, message delivery, or profile loading. I would select measurable indicators including availability, latency, error rate, and completion rate, establish a baseline from production data, and set SLOs that balance user expectations with engineering capacity. I would then connect alerts and error-budget reviews to those objectives so the team can act before reliability materially affects users.

Describe your approach to responding to a P1 production incident when the root cause is initially unclear.

I begin by confirming impact, scope, and severity through dashboards, logs, traces, and recent deployment history. I communicate a concise status, assign or coordinate investigation streams when needed, and prioritize safe mitigation such as rollback, traffic control, feature disablement, or capacity adjustment. After recovery, I document the timeline, root cause or contributing factors, and concrete preventive actions in a blameless postmortem.

What makes an alert actionable, and how would you reduce alert fatigue?

An actionable alert identifies a meaningful user or service impact, has a clear severity, links to relevant runbooks and dashboards, and calls for a specific response. I reduce alert fatigue by removing duplicate or low-signal alerts, using symptom-based alerts rather than every underlying cause, tuning thresholds with historical data, and regularly reviewing alert outcomes with the on-call team.

How do you use your backend engineering experience to improve reliability beyond infrastructure changes?

Application-level knowledge lets me trace failures through request paths, identify inefficient queries or dependency behavior, improve retries and timeouts, and add meaningful structured logs and telemetry. I work with product engineers to make failure modes explicit, introduce safe degradation paths, and address code-level causes instead of treating infrastructure as the only source of reliability issues.

How would you collaborate with a distributed product squad during normal operations and an out-of-hours incident?

During normal operations, I would make reliability work visible through shared dashboards, concise documentation, planning discussions, and regular reviews of incidents and error budgets. During an incident, I would establish a clear communication channel, provide time-stamped impact and mitigation updates, involve the appropriate owners quickly, and keep decisions focused on restoration. Once resolved, I would ensure follow-up items have accountable owners and measurable completion criteria.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.

On behalf of Feeld, GT is looking for a Site Reliability Engineer (SRE) to join a fast-growing consumer mobile product in the online dating space.

About the Client

Founded in 2014 as a dating app, Feeld gathered millions of users in one place to create a safer and more inclusive space online for everyone open to experiencing people and relationships in a new way. Their mission is to elevate the human experience of sexuality and relationships and create a world where everyone is more intimately connected to each other and themselves.

About the Project

You’ll join a consumer mobile product with an engineering and product organization of around 50 people distributed across Europe and the US.

The team works in small, autonomous product squads, each responsible for a specific area of the product and critical user journeys. Because the team operates across multiple regions without a full follow-the-sun model, strong observability, monitoring and reliable incident response are essential.

From a technical perspective, the team is focused on building reliable, observable systems that allow engineers to identify issues early, understand their impact and respond quickly when incidents occur.

  • Technology stack: Node.js, TypeScript, AWS, Cloudflare, CloudWatch, Sentry. React Native is used on the mobile side.

  • Team: Cross-functional product squads of approximately 6–8 people, distributed across Europe, the US and LATAM.

About the Role

We are looking for an experienced Site Reliability Engineer with a strong backend engineering background in Node.js and TypeScript.

The ideal profile is someone who started in backend/software engineering and has moved into SRE or reliability-focused work, combining a strong understanding of application code with hands-on experience in observability, monitoring and production incident management.

You will be embedded within a product squad and take ownership of the reliability and observability of critical user journeys. An important part of the role is being able to interpret production signals, identify when something is going wrong and begin mitigating incidents independently while bringing in the wider engineering team when needed.

Responsibilities:

  • Own observability for critical product and user journeys within your squad.

  • Define, build and maintain meaningful metrics, dashboards and alerts.

  • Define and maintain SLIs/SLOs for key services and product-level metrics.

  • Improve monitoring, logging, tracing and alerting across the squad’s systems.

  • Act as the first responder for critical P0/P1 production incidents, including out-of-hours incidents.

  • Investigate production signals, identify potential root causes and begin mitigating issues independently.

  • Coordinate with other engineers when broader support or escalation is required.

  • Participate in incident triage, mitigation and postmortems.

  • Identify recurring reliability issues and drive improvements to infrastructure, tooling and incident-response processes.

  • Work closely with backend and product engineers in a distributed, autonomous squad.

Essential knowledge, skills & experience:

  • Strong previous experience as a Backend / Software Engineer, with senior-level hands-on experience in Node.js and TypeScript.

  • Hands-on experience working in an SRE, Production Engineering or similar reliability-focused role.

  • Strong production experience with AWS

  • Experience with monitoring and observability across metrics, logging, tracing and alerting.

  • Practical experience responding to production incidents, including triage, mitigation and postmortems.

  • Ability to interpret monitoring signals and independently investigate and begin resolving production issues.

  • Understanding of both the application and infrastructure layers rather than infrastructure-only experience.

  • Strong communication skills and the ability to work autonomously within a distributed engineering team.

  • Comfortable participating in out-of-hours incident response as part of the team’s coverage model.

Nice-to-have

  • Experience with observability tools such as Sentry.

  • Experience with Cloudflare and CloudWatch.

  • Experience defining SLIs and SLOs for product-level metrics.

  • Experience with React Native or exposure to mobile application environments.

  • Previous experience with consumer mobile products or high-traffic B2C systems.

Interview Steps

  1. GT interview with Recruiter

  2. Technical interview

  3. Final interview

  4. Reference Check

We go beyond usual perks… By working with us, you will get:

  • Health insurance.

  • Wellbeing budget.

  • Sport coverage.

  • Learning budget.

  • 18 business days of paid vacation days per year

  • Paid sick leaves.

  • All public holidays are paid days off.

GT working model:

You will work directly with a client through our Extended Team model. We try to do things differently and put our efforts into integrating you as deeply as possible into the client’s team. You work with the same tools and technologies as they do and are managed directly by the client without any intermediary in between. We help you build relationships and create an environment where you genuinely feel like a member of the client’s team. We also encourage trips to a client and join teambuilding and after-work activities. Our Extended Team model is focused on long-term projects that last over several years.

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu