All remote jobs
Open role
Remote opportunity atGremlin

Data Scientist, AI/ML

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
40Listing views
1Application actions
7 Oct 2026Apply before
Opportunity details

About this role.

AI Summary

Gremlin is seeking a senior Data Scientist, AI/ML to turn chaos-engineering experiment data into automated failure analysis, remediation recommendations, and product capabilities. The role combines applied machine learning research with production engineering, including scalable pipelines, feature stores, real-time inference, and MLOps. Core technical methods include causal inference, graph ML, time-series modeling, and reinforcement learning for distributed-system reliability use cases. The successful candidate will partner closely with platform engineers and SREs in a remote-first, high-ownership environment focused on quality, delivery, and measurable customer impact.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

4/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThis is a highly senior applied ML role requiring at least five years of production ML experience plus strong software engineering and distributed-systems knowledge. Building trustworthy, actionable AI systems from complex experiment data requires advanced modeling judgment, rigorous evaluation, and close cross-functional execution.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianHighly competitive
$255,000
US market range$220k–$290k
AI insightThe posting explicitly discloses a US salary range of $220,000 to $290,000 per year, with a midpoint of $255,000. This range is consistent with the upper end of the US market for a senior/principal-level applied ML data scientist working on production infrastructure reliability, causal AI, and MLOps.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
Describe a production ML system you built for a complex operational or infrastructure problem.

I would explain the business objective, data sources, modeling approach, deployment architecture, and operational metrics. I would also describe how I monitored model quality, handled data drift, and partnered with engineering stakeholders to ensure the system delivered measurable outcomes.

How would you use causal inference to distinguish a true root cause from correlation in chaos-engineering experiment data?

I would define the treatment, outcomes, confounders, and causal graph before selecting an estimation strategy. Depending on experiment design and data quality, I would use randomized interventions where possible, propensity methods or doubly robust estimation for observational signals, and validate conclusions with controlled follow-up experiments.

What would you consider when designing a feature store for both offline training and real-time inference?

I would prioritize point-in-time correctness, feature definition consistency, lineage, freshness requirements, access controls, and online/offline parity. I would also establish validation, versioning, backfill processes, and monitoring so model behavior remains reproducible and reliable in production.

How would you evaluate an automated remediation recommendation system before allowing it to orchestrate changes?

I would begin with offline evaluation against historical outcomes, emphasizing precision, safety, expected impact, and calibrated uncertainty. I would then use staged rollout through human review, shadow mode, and limited-scope experiments, with clear rollback mechanisms and guardrails before enabling autonomous execution.

Tell us about a time you converted an ambiguous research problem into a shipped product capability.

I would outline how I clarified the customer problem and success metrics, broke the work into testable milestones, and delivered an initial solution quickly. I would highlight iterative experimentation, stakeholder communication, and how production feedback changed the final design.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

Data Scientist, AI/ML

Job Description:

Today’s complex, fast-paced systems have become a minefield of reliability risks, any of which could cause an outage that costs millions and destroys customer confidence. That’s why high-availability teams use Gremlin to find and fix reliability risks before they become incidents.

Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide. As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world’s largest organizations where high availability is non-negotiable.

About the Role of the Data Scientist, AI/ML

As a Data Scientist, AI/ML at Gremlin, you will have the opportunity to improve the reliability of the internet at large by turning millions of chaos engineering experiments into automated failure analysis and remediation. You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations). You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.

In this role, you’ll get to:

  • Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems
  • Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments
  • Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior
  • Develop scalable data pipelines and feature stores to process, enrich, and serve large volumes of experiment data for both model training and real-time inference
  • Collaborate closely with platform engineers and SREs to integrate AI-driven failure analysis and remediation capabilities directly into Gremlin’s core product
  • Apply advanced techniques, including causal inference, graph ML, time-series modeling, and reinforcement learning, to continuously improve the accuracy and actionability of automated failure analysis
  • Translate insights from millions of chaos experiments into AI-powered features that help customers automatically understand blast radius, pinpoint root causes, and accelerate recovery
  • Research and productionize novel ML approaches, including causal AI and agentic systems, that turn raw chaos experiment data into automated, reliable remediation strategies

We’ll expect you to have:

  • Experience as a self-driven and collaborative problem solver with strong communication skills
  • 5+ years professional experience building and productionizing machine learning, ideally for distributed systems, infrastructure, or DevOps and SRE use cases with more overall years of experience in software development.
  • Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning
  • Experience building data pipelines and feature stores that support both offline training and real-time inference
  • Experience with agile development environments and practices
  • Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices
  • Comfort partnering with platform engineers and SREs to turn research into shipped product features
  • Strong at breaking down ambiguous problems into concrete actions and milestones

Bonus Experience:

  • Experience with chaos engineering, site reliability engineering, or distributed systems
  • Background in agentic AI systems or large-scale causal inference in production
  • Experience standing up MLOps tooling such as model serving, monitoring, or feature store infrastructure
  • Working in Remote first environments
  • Has been on-call and participated in an incident management program

*The role does not offer sponsorship employment benefits.

**If you don’t think you meet all of the criteria above but still are interested in the job, please apply. Nobody checks every box, we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

Compensation

We expect the salary range for this role to be $220,000 – $290,000. We recognize that salary varies from person to person depending on level of experience and we welcome direct conversations about it. The final offer will vary based on assessment of a candidate’s skills and ability and our budget and market data.

Gremlin offers competitive total compensation packages including 401k Matching, Equity and other benefits such as flexible time off and paid company holidays.

About Gremlin:

Gremlin is a team of industry veterans and people eager to learn from one another. We set the standard for reliability and equip leading organizations with the mindset and expertise needed to drive reliability improvements that move the world forward. We’re backed by top-tier investors Index Ventures, Amplify Partners, and Redpoint Ventures. Our customers love us, and we’re thrilled to be a partner in their success.

What Do We Care About:

We Care about our People

People are our critical differentiators. The company strives to treat our people with respect, empathy, and dignity. We expect that our people will treat each other similarly. In both cases, we will assume good intent. All are welcome at Gremlin. We know our differences make us stronger and that our best ideas and contributions can come from anyone at any level.

We Care about Collaboration

Gremlin is strongest when we come together as one team with shared goals. Be the glue, not the glitter. But as a remote company, teamwork and collaboration won’t happen by accident. We approach every challenge as a shared challenge. We rely on each other for diverse perspectives and creative ideas. We celebrate our wins as a team.

We Care about Results

Be high productivity, low drama. Results matter. To keep our pace, everyone owns the outcomes of their actions and takes action when needed. We reward speed over perfection. We empower each other to iterate and experiment. You are welcome at Gremlin for who you are. The more voices and ideas we have represented in our business, the more we will all flourish, contribute, and build a more reliable internet.

Gremlin is a place where everyone can grow and is encouraged. However you identify and whatever background you bring with you, please apply if this sounds like a role that would make you excited to come into work everyday. It’s in our differences that we will find the power to keep building a more reliable internet by building and designing tools used by the best companies in the world.

Visit our website to learn more – https://www.gremlin.com/about

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu