[All remote jobs](https://jobicy.com/jobs.md)[![hims & hers logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2023/11/e7b9b920-221.jpeg)](https://jobicy.com/company/hims-hers.md)[hims & hers](https://jobicy.com/company/hims-hers.md)

# Sr. Staff Machine Learning Systems Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/hims-hers.md)ShareRemote from[USA](https://jobicy.com/job-region/usa.md)SalaryUSD 240k–265k / yrDepartment[Software Engineering](https://jobicy.com/categories/engineering.md)EmploymentFull TimeExperienceSeniorPublished9 Oct 2026Apply before8 Nov 2026Listing views54Application actions3Application toolkit

## Make your next move.

Prepare your resume, explore your fit, and draft a cover letter for this opportunity.

AI Summary

## The role, at a glance.

This Senior Staff Machine Learning Systems Engineer role leads the end-to-end technical strategy for trustworthy AI evaluation in a regulated healthcare environment. The position combines ML infrastructure, data pipeline engineering, LLM evaluation, statistical regression testing, and adversarial/red-team testing. It requires setting reusable organizational standards, delivering multi-quarter cross-functional programs, and aligning engineering, data science, clinical, legal, and product stakeholders. The successful candidate will mentor experienced engineers while building production systems that determine whether AI services are safe and effective to ship.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

5/5IndependentCollaborative

AI insightThis is a senior staff-level role requiring more than 10 years of relevant experience and ownership of ambiguous, organization-wide AI safety and evaluation challenges. It demands deep technical judgment in statistical methodology, ML data systems, regulated-product risk, and cross-team leadership.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate$252,500US market range$210k–$310k0$341k

AI insightThe disclosed base salary range is $240,000-$265,000 USD yearly, with a midpoint of $252,500. For a US-based Senior Staff ML Systems Engineer focused on evaluation infrastructure, AI safety, and regulated healthcare, an estimated market base-salary range is approximately $210,000-$310,000 USD yearly; equity and benefits may add materially to total compensation.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Machine Learning Infrastructure](https://jobicy.com/jobs?search_keywords=Machine%20Learning%20Infrastructure.md)[LLM Evaluation](https://jobicy.com/jobs?search_keywords=LLM%20Evaluation.md)[AI Safety](https://jobicy.com/jobs?search_keywords=AI%20Safety.md)[Data Engineering](https://jobicy.com/jobs?search_keywords=Data%20Engineering.md)[Python](https://jobicy.com/jobs?search_keywords=Python.md)[Statistical Testing](https://jobicy.com/jobs?search_keywords=Statistical%20Testing.md)[Regression Testing](https://jobicy.com/jobs?search_keywords=Regression%20Testing.md)[Red Team Testing](https://jobicy.com/jobs?search_keywords=Red%20Team%20Testing.md)[Dataset Versioning](https://jobicy.com/jobs?search_keywords=Dataset%20Versioning.md)[Healthcare Technology](https://jobicy.com/jobs?search_keywords=Healthcare%20Technology.md)

Sample interview questionsHow would you design a statistically sound regression gate for an LLM-based clinical workflow?I would first define the unit of evaluation, risk-weighted failure taxonomy, representative benchmark slices, and human-label ground truth. I would use paired comparisons against the current production baseline, select appropriate significance tests, apply multiple-comparison correction where metrics or slices are evaluated jointly, and establish practical-effect thresholds in addition to statistical significance. The gate would include minimum performance requirements for high-risk cohorts, confidence intervals, rollback criteria, and auditable versioning of data, prompts, models, and scoring logic.

Describe how you would calibrate and validate an LLM judge used in release decisions.

I would build a stratified calibration set with expert human labels, including common, difficult, and safety-critical cases. I would measure agreement with human reviewers using appropriate reliability metrics, inspect systematic disagreements by failure category and cohort, and refine the rubric, prompt, and judge configuration without leaking test data into development. Before release, I would maintain an independent holdout set and continuously monitor judge drift and disagreement rates after deployment.

How would you turn a manual and inconsistent review process into an automated evaluation platform?

I would begin by mapping the existing workflow, decision criteria, data sources, reviewer disagreement patterns, and operational pain points. I would define a versioned evaluation contract covering datasets, labels, metrics, thresholds, ownership, and escalation paths, then implement the smallest reliable automated path alongside the manual process for validation. After demonstrating agreement and time savings, I would expand coverage, publish reporting surfaces for non-engineering stakeholders, and retire manual steps only when controls and exception handling are mature.

How do you approach adversarial and red-team testing for a healthcare AI product?

I would use a risk-based threat model covering harmful advice, missed escalation, hallucinations, privacy exposure, bias, prompt injection, and edge cases related to patient context. Test suites would combine curated scenarios, expert-generated attacks, historical failures, synthetic variations, and targeted testing across demographic and clinical cohorts. Every discovered failure would be categorized, prioritized by severity and likelihood, linked to mitigation owners, and converted into a regression test so the program compounds over time.

Tell us about handling a major technical disagreement across engineering, data science, and product.

I would make the disagreement concrete by documenting the decision, assumptions, success metrics, risks, and reversible versus irreversible consequences. I would invite the relevant experts to review evidence, run a focused experiment or prototype where uncertainty remains, and communicate trade-offs in language appropriate to each stakeholder group. Once a decision is made, I would document the rationale, establish measurable follow-up criteria, and ensure the team remains aligned on execution even if individuals preferred a different option.

Opportunity details

## About this role.

Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve.

Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit [hims.com/about](http://hims.com/about) and [hims.com/how-it-works](http://hims.com/how-it-works) . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit [www.hims.com/careers-professionals](http://www.hims.com/careers-professionals).

### About the Role:

How do we make advanced AI/ML not just powerful but trustworthy enough to run in a regulated healthcare environment? We’re looking for a Senior Staff engineer who can own that question end to end: the data pipelines that feed our models and evaluations, and the evaluation infrastructure — judges, scorers, statistical regression gates, red-team testing — that decides whether AI models are effective and safe to ship to patients.

This is a leadership role for someone who works comfortably across disciplines. You’ll set technical direction for how we build, version, and trust the data and judgments that our AI products are evaluated against. Your scope will expand from raw data ingestion and feature/dataset pipelines, through evaluation methodology and statistical rigor, to the reporting surfaces that let clinical and product teams act on what we learn.

You’ll spend most of your time on problems that don’t have an existing playbook: ambiguous, cross-team, and genuinely hard to reason about. The job is to bring clarity to that ambiguity, chart a path the rest of the team and organization can follow, and see it through from idea to production, building the relationships and buy-in along the way to make it stick.

### You Will:

### Own the evaluation as a whole, not just a slice of it

*

Set the technical direction for our evaluation systems; metric, judge and scorer design, the statistical methodology behind regression decisions, and the infrastructure that tracks and categorizes failures over time.

*

Design and scale the data pipelines — ingestion, transformation, dataset versioning, labeling and calibration workflows — that both evaluation and downstream data science work depend on.

*

Proactively address challenges in scaling and complexity AI evaluation.

Lead projects that span teams and quarters

*

Define how we evaluate any new AI service from scratch. Drive multi-team initiatives like replacing manual, inconsistent review processes with statistically sound, automated gates.

*

Own our approach to adversarial and red-team evaluation as a risk-reduction program, designing the test suites and failure taxonomies that catch safety and edge-case issues before they reach patients.

*

Work through complex, cross-team technical disagreements and drive alignment across engineering, product and AI leaders.

Turn hard, ambiguous problems into solutions other teams can build on

*

Originate new approaches and methodology that becomes a reusable standard rather than a one-off fix.

*

Take vague, cross-team pain points (“we do this manually and it’s inconsistent”) all the way from a rough idea to a fully-specified, shipped system, without needing to hand off any part of the journey.

*

Lead major platform improvements; re-architecting core systems, removing brittle logic, modernizing how things run with impact that’s felt org-wide, not just on your own team.

Grow the people and the network around you

*

Build real working relationships across ML engineering, data science, platform engineering, clinical, legal, and product, the kind of trust that gets you looped in early, before decisions are locked in.

*

Become a go-to voice on evaluation methodology and data pipeline design: share what you’ve learned in internal talks, write things up so other teams can use them, and expect your ideas to shape how others approach similar problems.

*

Mentor other engineers, including experienced ones, and help raise the technical and statistical bar of the teams you work with.

### You Have:

*

10+ years of experience in ML infrastructure, data engineering, or evaluation/testing systems, with a track record of impact that reaches beyond a single team or project.

*

Hands-on depth in evaluation systems: designing and calibrating LLM judges/scorers, building statistically sound regression-testing methodology (e.g., paired significance testing with proper correction for multiple comparisons), measuring agreement against human labels, and designing adversarial/red-team evaluation approaches.

*

Hands-on depth in data pipeline engineering: dataset versioning, feature and benchmark pipelines, labeling and calibration workflows, and high-throughput ingestion and transformation systems.

*

A history of building things that became the standard approach for others — not just solving your own problem, but changing how a broader group of people tackle a category of problem.

*

Experience leading multi-team projects to completion, including navigating and resolving genuine technical disagreement along the way.

*

A track record of mentoring other engineers, including senior ones, and visibly raising the bar for the teams around you.

*

Excellent communication — comfortable adapting the same idea for different audiences, and confident building support for it well before launch.

*

Strong Python, and enough statistical fluency to design and defend a testing framework that real production decisions ride on.

### Nice to Have:

*

Experience with Databricks, MLflow, Unity Catalog, or similar data/eval platforms.

*

Experience building reporting tools for people without direct engineering access (e.g., automated Slack digests, spreadsheet reports for non-technical teams).

*

Prior experience in a regulated industry (healthcare, fintech, life sciences).

*

A track record of company-wide talks or write-ups that changed how other teams approached a problem.

### Our Benefits (there are more but here are some highlights):

*

Competitive salary & equity compensation for full-time roles

*

Unlimited PTO, company holidays, and quarterly mental health days

*

Comprehensive health benefits including medical, dental & vision, and parental leave

*

Employee Stock Purchase Program (ESPP)

*

401k benefits with employer matching contribution

*

Offsite team retreats

We are committed to building a workforce that reflects diverse perspectives and prioritizes ethics, wellness, and a strong sense of belonging. If you’re excited about this role, we encourage you to apply—even if you’re not sure if your background or experience is a perfect match.

Hims considers all qualified applicants for employment, including applicants with arrest or conviction records, in accordance with the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance, the California Fair Chance Act, and any similar state or local fair chance laws.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Hims & Hers is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at [accommodations@forhims.com](mailto:accommodations@forhims.com) and describe the needed accommodation. Your privacy is important to us, and any information you share will only be used for the legitimate purpose of considering your request for accommodation. Hims & Hers gives consideration to all qualified applicants without regard to any protected status, including disability. Please do not send resumes to this email address.

To learn more about how we collect, use, retain, and disclose Personal Information, please visit our [Global Candidate Privacy Statement](https://www.hims.com/global-candidate-privacy-statement).

Show more

[Apply now >](https://jobicy.com/jobs/154909-sr-staff-machine-learning-systems-engineer.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

*
![Spotify logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/03/Jobicy-210328105941-289246.jpg)
Spotify Oct 9  New

### [Engineering Manager – Experimentation](https://jobicy.com/jobs/154910-engineering-manager-experimentation.md)

Think it, build it, ship it, tweak it” is at the core of Spotify’s product development approach. In the Experimentation product area we are providing Spotify and external customers with…

*
![ClickUp logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/cd5be0ef-221.jpg)
ClickUp Oct 9  New

### [Business Systems Engineer](https://jobicy.com/jobs/152814-business-systems-engineer.md)

At ClickUp, we’re building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an…

*
![Bayesian Health logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/f74593d8-221.png)
Bayesian Health Oct 9  New

### [Software Engineer, Data Integration](https://jobicy.com/jobs/152819-software-engineer-data-integration.md)

Software Engineer, Data Integration In Brief We’re a rapidly growing startup on a mission to make healthcare proactive by empowering physicians, nurses, and care team members with real-time data to…

*
![Reddit logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/10/Reddit.jpg)
Reddit Oct 9  New

### [Ads Conversion Modeling, Machine Learning Engineering Manager](https://jobicy.com/jobs/152810-ads-conversion-modeling-machine-learning-engineering-manager.md)

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit…

*
![CentralReach logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/07/a0eac9f542eb8c3fff7c19990dd5c1bb.jpg)
CentralReach Oct 9  New

### [Sr. Software Engineer, Ruby on Rails](https://jobicy.com/jobs/152795-sr-software-engineer-ruby-on-rails.md)

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy…

*
![Bloomreach logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/60790cd5-221.png)
Bloomreach Oct 8  New

### [Quality Assurance Engineer II](https://jobicy.com/jobs/154875-quality-assurance-engineer-ii.md)

Bloomreach is building the world’s premier agentic platform for personalization.We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We’re taking…

*
![Bayesian Health logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/f74593d8-221.png)
Bayesian Health Oct 8  New

### [Senior Design Quality Engineer](https://jobicy.com/jobs/154879-senior-design-quality-engineer.md)

Senior Design Quality Engineer In Brief We’re a rapidly growing startup on a mission to make healthcare proactive by empowering physicians, nurses, and care team members with real-time data to…

*
![StackAdapt logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/01/bf5f2bb689c6f35f6cab5459f3eff869.jpg)
StackAdapt Oct 8  New

### [Senior Quality Engineer](https://jobicy.com/jobs/154871-senior-quality-engineer-3.md)

StackAdapt is the leading technology company that empowers marketers to reach, engage, and convert audiences with precision. With 465 billion automated optimizations per second, the AI-powered StackAdapt Marketing Platform seamlessly…

*
![Cribl logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/65d82c04-221.jpeg)
Cribl Oct 8  New

### [Sr. Engineering Manager, SDET](https://jobicy.com/jobs/154877-sr-engineering-manager-sdet.md)

Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the world’s biggest enterprises, including half…

*
![PadSplit logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/06/da525286873130ed0b95212896ef55d4.jpg)
PadSplit Oct 8  New

### [QA Engineer (Anywhere in Europe)](https://jobicy.com/jobs/154874-qa-engineer-anywhere-in-europe.md)

The Role We Need: PadSplit is hiring a QA Engineer to strengthen quality across our Marketplace pod, spanning Mobile (iOS/Android), Web, and backend Platform systems. As our product complexity increases,…

[Browse all jobs](https://jobicy.com/jobs.md)