All remote jobs

DevOps Engineer (Observability)

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Remote from
Ireland
Salary
Undisclosed
Employment
Full Time
Experience
Senior
Published
Apply before
8 Nov 2026
Listing views
26
Application actions
0
Application toolkit

Make your next move.

Prepare your resume, explore your fit, and draft a cover letter for this opportunity.

AI Summary

The role, at a glance.

Twilio is hiring a senior-level DevOps/Platform Observability Engineer to help rebuild its observability ecosystem around an OpenTelemetry-first architecture. The role leads the design and delivery of scalable telemetry pipelines, data lakes, query systems, APIs, and developer tooling for logs, metrics, traces, and profiling. It requires deep distributed-systems expertise and hands-on experience with AWS, Kubernetes, infrastructure as code, and large-scale tools such as Prometheus, Kafka, ClickHouse, or equivalents. The engineer will provide technical leadership across Platform Engineering and R&D, mentor peers, and balance reliability, cost, performance, and usability. The position is remote from Ireland with occasional travel for team or project meetings.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

4/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

5/5
IndependentCollaborative
AI insightThis is a high-complexity platform role involving an organization-wide observability transformation, high-cardinality telemetry, distributed systems, and cost-sensitive data architecture. Success requires both deep technical execution and the ability to influence architecture and standards across multiple engineering teams.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$185,000
US market range$155k–$220k
AI insightNo salary was disclosed in the posting, so these figures are estimated US-market annual base-salary benchmarks in USD for a senior/staff-level DevOps or platform observability engineer. Actual compensation for this Ireland-based remote role may differ materially based on local market, level, and equity or bonus components.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
How would you design a scalable telemetry platform that supports logs, metrics, traces, and profiling across thousands of microservices?

I would define a common OpenTelemetry-based collection and semantic-conventions layer, then build independently scalable ingestion, enrichment, storage, and query paths for each signal. I would use durable buffering, tenant-aware controls, schema governance, and sampling or aggregation strategies to control volume. The design would include reliability SLOs, cost attribution, and a phased migration path so teams can adopt the platform without disrupting production services.

How do you address high-cardinality metrics without losing valuable diagnostic context?

I start by identifying the cardinality source and the decisions the metric is intended to support. I use bounded labels, aggregation, exemplars that link metrics to traces, and controlled sampling for detailed dimensions. I also establish telemetry standards and dashboards that expose cardinality and cost, allowing teams to correct inefficient instrumentation before it affects platform stability.

Describe your approach to correlating logs, metrics, and traces during an incident.

I ensure trace and span identifiers are consistently propagated through service boundaries and included in structured logs. Metrics should expose useful service and operation dimensions, with exemplars or links to representative traces where supported. The resulting workflow lets responders move from an alert to a relevant metric view, trace, and related structured logs quickly, reducing time to isolate the failing dependency or code path.

What trade-offs would you consider when using an S3-based data lake with a ClickHouse-backed query layer for observability data?

I would evaluate ingestion latency, query latency, retention, indexing strategy, data compaction, and operational cost. S3 provides economical durable storage, while ClickHouse can provide fast analytical queries, but the architecture needs clear hot-versus-cold data policies and efficient partitioning. I would benchmark representative incident and investigative queries, define SLAs, and make data lifecycle policies visible to users so cost and performance remain predictable.

How would you drive adoption of new telemetry standards across teams with existing fragmented tooling?

I would begin with a clear target architecture, migration guidance, and opinionated libraries or templates that make the correct approach easier than the legacy one. I would partner with early-adopter teams to demonstrate measurable gains in incident response, data quality, and cost, then use those results to refine the rollout. Regular technical forums, documentation, compatibility support, and leadership alignment would help establish the standards as shared engineering practice rather than a centrally imposed mandate.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.
Opportunity details

About this role.

Who we are

At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences.

Our dedication to remote-first work, and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands.
.

Hiring and how we work

We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions!

Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings.
.

See yourself at Twilio

Join the team as our next Software Engineer on Twilio’s platform engineering observability team.

About the job

This position is needed to help our platform engineering observability team. Twilio is undergoing a large-scale observability transformation—and you can help shape the foundation. Observability is a strategic pillar and a key enabler for faster incident response, deeper customer-centric insights, and more cost-effective platform operations.

As a Software Engineer on the Platform Observability team, you’ll play a critical role in re-architecting how telemetry flows and is utilized through Twilio—making it structured, accessible, affordable, and actionable. Over the next 3 years, Twilio is rebuilding nearly every component of our observability platform, from data collection to real-time analytics. You will drive core initiatives that shift Twilio from fragmented tooling and wasteful data sprawl to a unified, OpenTelemetry-first observability stack built for scale.

You’ll lead technically and strategically—designing platform components, influencing org-wide architectural decisions, mentoring engineers, and engaging directly with teams across Platform Engineering and R&D

Responsibilities

In this role, you’ll:

  • Lead the end-to-end architecture and delivery of key observability platform components, with a focus on reliability, scalability, and usability.
  • Drive consistency and quality across all observability signals—logs, metrics, traces, and continuous profiling—building intuitive workflows for engineers.
  • Serve as a technical advisor and mentor across the platform org, guiding design decisions and aligning cross-team efforts with long-term architectural goals.
  • Go deep in one or more problem areas (e.g., high-cardinality telemetry, distributed tracing correlation, compute cost insights), while ensuring the platform scales horizontally.
  • Collaborate with product teams, SREs, and developer experience groups to deeply understand telemetry needs and integrate observability into core engineering workflows.
  • Design and build developer-friendly tooling and APIs to support incident response, performance analysis, and platform debugging at scale.
  • Leverage (and optionally contribute to) open-source standards like OpenTelemetry to ensure interoperability and extensibility.
  • Champion a pragmatic approach to observability—balancing performance, cost, and user value across diverse engineering teams.

Qualifications

Twilio values diverse experiences from all kinds of industries, and we encourage everyone who meets the required qualifications to apply. If your career is just starting or hasn’t followed a traditional path, don’t let that stop you from considering Twilio. We are always looking for people who will bring something new to the table!

*Required:

  • Proven expertise in building and scaling observability systems (e.g., logging platforms, metrics pipelines, tracing infrastructure, or profiling tools).
  • Lead technical execution for major components of Twilio’s observability overhaul, including our shift to centralized S3-based data lakes, OpenTelemetry instrumentation, and ClickHouse-backed query engines.
  • Proficiency in at least one modern programming language (e.g., Go, Python, Java).
  • Familiarity with high-cardinality data challenges and telemetry correlation techniques.
  • Experience designing high-scale telemetry systems (e.g., Prometheus, ClickHouse, OpenTelemetry, Kafka, or equivalent).
  • Solid understanding of distributed systems and the challenges of observability in complex, microservice-based environments.
  • Experience with AWS, Kubernetes, and infrastructure-as-code tools.
  • Provide architectural guidance and thought leadership across teams, helping to establish clear telemetry standards, efficient usage patterns, and scalable platform abstractions.
  • Ability to make forward-looking technical decisions and lead others through ambiguity and chan

Desired:

  • Familiarity with ClickHouse, Grafana Mimir, Athena, or equivalent systems for log and metrics querying.
  • Contributions to open-source observability tools or communities.
  • Experience building cost visibility or FinOps tooling for cloud compute and telemetry pipelines.

Location

This role will be remote from Ireland.

Travel

We prioritize connection and opportunities to build relationships with our customers and each other. For this role, you may be required to travel occasionally to participate in project or team in-person meetings.

What We Offer

Working at Twilio offers many benefits, including competitive pay, generous time off, ample parental and wellness leave, healthcare, a retirement savings program, and much more. Offerings vary by location.

Twilio thinks big. Do you?

We like to solve problems, take initiative, pitch in when needed, and are always up for trying new things. That’s why we seek out colleagues who embody our values — something we call Twilio Magic. Additionally, we empower employees to build positive change in their communities by supporting their volunteering and donation efforts.

So, if you’re ready to unleash your full potential, do your best work, and be the best version of yourself, apply now! If this role isn’t what you’re looking for, please consider other open positions.

.

Stay alert to recruitment fraud

We care about your safety. Scammers sometimes impersonate Twilio recruiters through fake job postings, emails, websites, or messages. Please ensure you are engaging with an official @twilio.com email address. We will never ask for payment, gift cards, cryptocurrency, or banking information during the recruiting process. We do not make job offers without a formal interview process or conduct interviews exclusively through text-based messaging apps.

.

Twilio is proud to be an equal opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Qualified applicants with arrest or conviction records will be considered for employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Additionally, Twilio participates in the E-Verify program in certain locations, as required by law.

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. .

Log in to save
One quick step before you apply

Sign in to continue.

Sign in or create a free account to continue to the employer's application.

Applying is free. After signing in, return to this job and select Apply Now.
Add alert
Jobs Talent AI Tools Salaries
Menu