All remote jobs
Open role
Remote opportunity atSWORD Health

Senior ML Engineer (Europe-based/Remote)

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
64Listing views
5Application actions
21 Sep 2026Apply before
Opportunity details

About this role.

AI Summary

Sword is seeking a Europe-based Senior ML Engineer to build and operate production AI systems for a clinical-care platform. The role centers on reliable agentic LLM workflows, including retrieval, tool use, orchestration, evaluation harnesses, monitoring, and continuous improvement. The engineer will own projects from ambiguous exploration through deployment, while selecting techniques such as prompting, RAG, distillation, or fine-tuning based on evidence. Success requires strong production engineering judgment and close collaboration with product, clinical, and engineering stakeholders in a high-stakes healthcare environment. The position also includes code review, knowledge sharing, and mentoring of less-experienced engineers.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

5/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

5/5
IndependentCollaborative
AI insightThis is a senior, end-to-end role requiring deep ML and software engineering capability, particularly in production LLM systems and rigorous evaluation. Clinical use cases raise the quality, safety, reliability, and stakeholder-management expectations beyond those of a typical ML engineering position.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$165,000
US market range$125k–$210k
AI insightNo salary range was provided. For a senior ML engineer building production LLM and healthcare AI systems remotely across Europe, an estimated total annual compensation range is USD 125,000 to USD 210,000, with a midpoint of USD 165,000; actual pay should vary materially by country, local employment structure, experience, and equity valuation.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
Describe an LLM system you shipped to production and how you measured whether it was successful.

I would explain the user problem, system architecture, model and retrieval choices, launch criteria, and production metrics such as task success, groundedness, latency, cost, escalation rate, and user outcomes. I would also describe the monitoring and iteration process used after launch.

How would you design an evaluation framework for a clinical-facing agentic LLM workflow?

I would begin with a representative, versioned evaluation set covering routine, edge-case, adversarial, and safety-critical scenarios. The framework would combine deterministic checks, task-specific quality metrics, structured LLM-as-judge scoring, expert clinical review, and regression gates that prevent releases when quality or safety declines.

When would you choose prompting, RAG, fine-tuning, or distillation for an ML problem?

I would start with clear error analysis and baseline measurements. Prompting is appropriate for fast iteration and well-supported general capabilities; RAG is useful when answers require current or proprietary evidence; fine-tuning can improve repeatable behavior or specialized tasks with sufficient quality data; and distillation is useful when lower cost or latency is needed after establishing a stronger teacher system.

How do you manage safety and reliability when an AI system can call tools or take multi-step actions?

I would constrain tool permissions, validate inputs and outputs, use structured schemas, require confirmation for consequential actions, and implement timeouts, retries, idempotency, and audit logging. I would also use policy checks, human escalation paths, simulation testing, and production monitoring to detect unsafe behavior early.

How would you communicate an ML tradeoff to clinical and product stakeholders?

I would frame the decision in terms of user and clinical impact, expected benefits, known failure modes, uncertainty, and operational cost rather than model terminology alone. I would present measurable options, recommend a safe rollout plan, and make explicit which risks require stakeholder or clinical-owner approval.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.
At Sword, we’re building AI to heal billions and unlock humanity’s full potential. In doing so, we’re pioneering AI Care, a fundamentally new approach to healthcare built for medical reasoning, safety, and real-time treatment, not generic technology applied after the fact. As both a clinical-centric frontier AI lab and an applied AI platform, Sword is reimagining how care is delivered at scale, removing traditional barriers like appointments, waiting rooms, and stigma so more people can access the care they need – and ultimately get back to lives lived in full.
Since 2020, Sword has expanded across Musculoskeletal, Women’s Health, Cardiometabolic, and Mental Health, and is now moving beyond the session to a fully AI-native, 24/7 care program that brings physical activity, therapeutic exercise, psychotherapy, nutrition, and behavior change into one connected experience. More than 1 million members across three continents have completed over 15 million AI sessions, helping 2,000+ enterprise clients avoid more than $1 billion in unnecessary healthcare costs. Backed by 59 clinical studies, 43 patents, and more than $500 million raised from leading investors including Khosla Ventures, General Catalyst, and Founders Fund, Sword is defining a new standard for healthcare.
 
This position is open to candidates based anywhere in Europe.

AI Proficiency at Sword

    AI fluency is a core expectation at Sword. Every candidate is assessed against our three-level framework — be ready to share real examples of how AI is already part of how you work.

  • Explorer (Level 1) — Uses AI daily to boost personal productivity

  • Builder (Level 2) — Creates workflows and tools that elevate the whole team

  • Integrator (Level 3) — Embeds AI into products and processes at scale

  • Every hire must demonstrate at least Level 1. The expected level will vary depending on the seniority of the role.

What you’ll be doing:

  • Own ML projects end-to-end: take problems from exploration through production deployment and keep iterating once real users are on them;
  • Build agentic LLM systems: design multi-step workflows with tool use, retrieval, and orchestration, and make them reliable enough for clinical settings;
  • Treat evaluation as core engineering work: build eval sets, offline and online harnesses, LLM-as-judge pipelines with human review, and regression tests that catch quality drops before they reach users;
  • Improve model quality with whatever fits the problem: prompting, retrieval, distillation, or fine-tuning, chosen on evidence rather than habit;
  • Work across the full AI stack: data prep, model adaptation, serving, monitoring, and the feedback loops that keep systems improving in production;
  • Partner with Product, Clinical, and Engineering: translate clinical requirements into technical decisions and surface tradeoffs early;
  • Help the team get better: review code, share what you learn, and mentor engineers earlier in their careers.

What you need to have:

    • Experience shipping ML systems to production that people actually depend on;
    • Hands-on LLM work in production: prompting, retrieval, tool calling, and agent-style workflows;
    • A rigorous approach to evaluation: you’ve built eval datasets and frameworks, and you can tell a real improvement from noise;
    • Strong ML fundamentals: you know which approach fits which problem and can reason clearly about tradeoffs;
    • Comfort with ambiguity: you’ve taken loosely defined problems and turned them into something running in production;
    • Solid engineering skills: production-quality code, familiarity with distributed systems, and the patience to debug messy ML pipelines;
    • Clear communication with both technical and clinical stakeholders.
    • Bonus points

      • Experience with fine-tuning or preference optimization (RLHF, DPO, or similar);
      • Healthcare AI, or other high-stakes domains where errors carry real cost;
      • Built agent frameworks or evaluation tooling from scratch;
      • Open source contributions, technical writing, or other knowledge sharing.
*This range includes base, variable and equity
 
As this role may be filled in different European countries, the compensation range will be adjusted to reflect local market conditions and cost of living in the country of hire.
These compensation bands are just the starting point. Once someone joins and proves they’re outlier talent, we adjust quickly to ensure their compensation aligns with their impact.
Our job titles may span more than one career level. Actual pay is determined by skills, qualifications, experience, location, market demand, and other factors. Compensation details listed in this posting reflect the base salary and any potential variable, bonus or sales incentives, and the Company’s estimation of the value of private company stock options, if applicable. The pay range is subject to change, future value of company stock options is not guaranteed, and compensation may be modified in the future. In addition to our total compensation, Sword offers a number of benefits as listed below.
 
 
Sword Benefits & Perks
 
At Sword, we are committed to supporting our employees in every location. The benefits and perks applicable to this role are those standard to the country where the candidate will be based, in line with local regulations and market practices.
 
SWORD Health, which includes SWORD Health, Inc. and Sword Health Professionals (consisting of Sword Health Care Providers, P.A., SWORD Health Care Providers of NJ, P.C., SWORD Health Care Physical Therapy Providers of CA, P.C.*) complies with applicable Federal and State civil rights laws and does not discriminate on the basis of Age, Ancestry, Color, Citizenship, Gender, Gender expression, Gender identity, Gender information, Marital status, Medical condition, National origin, Physical or mental disability, Pregnancy, Race, Religion, Caste, Sexual orientation, and Veteran status.

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.
Application method

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
or continue without an account
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu