About this role.
Sword Health is seeking an AI Research Scientist to design and execute research on LLM fine-tuning, alignment, and post-training methods for clinical and therapeutic domains. The role involves developing foundational AI models, contributing to the full model development cycle, and collaborating across teams to translate research into production systems. Candidates should have a PhD in a relevant field, hands-on experience with LLMs, and a strong publication record. This remote position offers an opportunity to work on cutting-edge AI in healthcare, with a focus on long-term ambitious goals like clinical memory and safety validation.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
5/5Autonomy Level
4/5Communication Load
5/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Manager,
I am writing to express my strong interest in the AI Research Scientist position at Sword Health. With a PhD in Computer Science and extensive experience in fine-tuning large language models, I am excited about the opportunity to apply my skills to revolutionize healthcare through AI.
My research has focused on post-training methods such as SFT and RLHF, and I have a track record of publishing in top-tier AI conferences. I am particularly drawn to Sword's mission of building AI-native care programs that remove barriers to access.
I am confident that my technical expertise and collaborative mindset would contribute to Sword's ambitious goals. Thank you for considering my application.
Sincerely, [Your Name]
Sample interview questions
In my PhD research, I implemented RLHF to align a 7B parameter language model with clinical guidelines. I collected human feedback from medical professionals, trained a reward model, and used PPO to fine-tune the policy. The project improved factual accuracy in medical responses by 30%.
I start by defining key metrics such as accuracy, safety, and fairness. For healthcare, I also include domain-specific evaluations like clinical relevance and adherence to guidelines. I use a combination of automated benchmarks and human evaluation, ensuring rigorous statistical testing to validate results.
During my internship at a leading AI lab, I worked with engineering and product teams to deploy a dialogue system for customer support. I handled model optimization and evaluation, while engineers integrated it into the platform. The system reduced response times by 40% and was well-received by users.
I have worked on vision-language models for medical imaging. For instance, I developed a model that combines chest X-rays with patient histories to generate diagnostic reports. This improves accuracy and provides explainable AI, which is crucial in clinical settings.
I regularly read papers from top conferences like NeurIPS and ICML, and participate in online communities. I prioritize approaches that have strong theoretical foundations and empirical results. For new techniques, I run small-scale experiments to validate their potential before integrating them into larger projects.
Since 2020, Sword has expanded across physical therapy, women’s health, cardiometabolic, and mental health, and is now moving beyond the session to a fully AI-native, 24/7 care program that brings physical activity, therapeutic exercise, psychotherapy, nutrition, and behavior change into one connected experience. More than 700,000 members across three continents have completed over 10 million AI sessions, helping 1,000+ enterprise clients avoid more than $1 billion in unnecessary healthcare costs. Backed by 42 clinical studies, 44+ patents, and more than $500 million raised from leading investors including Khosla Ventures, General Catalyst, and Founders Fund, Sword is defining a new standard for healthcare.
Explorer (Level 1) — Uses AI daily to boost personal productivity
Builder (Level 2) — Creates workflows and tools that elevate the whole team
Integrator (Level 3) — Embeds AI into products and processes at scale
AI fluency is a core expectation at Sword Health. Every candidate is assessed against our three-level framework — be ready to share real examples of how AI is already part of how you work.
Every hire must demonstrate at least Level 1. The expected level will vary depending on the seniority of the role.
What you’ll be doing:
- Design and execute research on LLM fine-tuning, alignment, and post-training methods (SFT, RLHF) tailored for clinical and therapeutic domains;
- Develop and improve foundational AI models that power our AI agents, spanning language, vision, speech, and multimodal systems;
- Contribute to the full model development cycle: dataset curation and annotation, architecture design, training, evaluation, and iteration;
- Collaborate across AI Engineering, Product, and Clinical teams to translate research breakthroughs into production systems that deliver patient care;
- Work towards long-term ambitious research goals, such as clinical memory, long-horizon planning, and safety validation, while identifying and delivering immediate milestones;
- Advance the field by publishing in top-tier AI venues and clinical journals, contributing to Sword’s growing body of peer-reviewed research.
What you need to have:
- A PhD in Computer Science, Machine Learning, Natural Language Processing, or a closely related AI field;
- Hands-on experience fine-tuning large language models (pre-training, SFT, RLHF, or related post-training techniques);
- A strong publication track record in peer-reviewed AI conferences or journals;
- Proficiency in Python and deep experience with modern ML frameworks (e.g., PyTorch, JAX);
- Demonstrated ability to design rigorous experiments and interpret their results.
What we would love to see:
- First-author publications in top-tier AI conferences (e.g., NeurIPS, ICML, ICLR, ACL, EMNLP, COLM, CVPR);
- Deep expertise in one or more of: large language models, reinforcement learning from human feedback, multimodal learning (vision, speech), or agentic AI systems;
- Experience building or contributing to LLM-based agents, including prompt engineering, memory orchestration, or agentic workflows;
- A track record of taking research ideas from conception to working systems, including developing and debugging complex ML pipelines;
- Industry experience during or after the PhD (e.g., research internships at leading AI labs);
- Comfort with ambiguity and a track record of delivering results in fast-moving, high-uncertainty environments where research and product development happen in parallel;
- Strong communication skills and a history of effective cross-functional collaboration;
- A broader record of research excellence demonstrated through grants, fellowships, patents, or impactful open-source contributions.
These compensation bands are just the starting point. Once someone joins and proves they’re outlier talent, we adjust quickly to ensure their compensation aligns with their impact.
Our job titles may span more than one career level. Actual pay is determined by skills, qualifications, experience, location, market demand, and other factors. Compensation details listed in this posting reflect the base salary and any potential variable, bonus or sales incentives, and the Company’s estimation of the value of private company stock options, if applicable. The pay range is subject to change, future value of company stock options is not guaranteed, and compensation may be modified in the future. In addition to our total compensation, Sword offers a number of benefits as listed below.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.










