All remote jobs
Open role
Remote opportunity atSporty Group

Senior AI Scientist

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
21Listing views
2Application actions
18 Sep 2026Apply before
Opportunity details

About this role.

AI Summary

Sporty is seeking a Senior AI Scientist to improve the quality, accuracy, safety, and effectiveness of AI-powered conversational products. The role centers on analyzing AI interactions, designing prompt strategies, optimizing RAG retrieval, evaluating LLMs, and measuring outcomes through structured experiments and quality metrics. The successful candidate will work closely with AI Engineers to move validated improvements into production while maintaining benchmarks, evaluation frameworks, and documentation. This remote-first position requires at least five years of relevant AI, ML, NLP, or conversational AI experience, with strong Python, SQL, stakeholder communication, and ownership skills.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

4/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThis is a senior, technically demanding role requiring deep practical expertise across LLM evaluation, prompt optimization, RAG systems, experimentation, Python, SQL, and production AI operations. The role also requires independent judgment to translate interaction data and user feedback into measurable product improvements.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianHighly competitive
$210,000
US market range$170k–$250k
AI insightNo salary range was provided. For the U.S. market, a Senior AI Scientist focused on LLMs, RAG, evaluation, and production AI commonly commands an estimated base salary range of $170,000 to $250,000 annually; a $210,000 annual median is a reasonable market estimate before bonus, equity, or location adjustments.

Core skills

Skills and capabilities most closely associated with this opportunity.

Cover letter sample

Dear Hiring Team,

I am excited to apply for the Senior AI Scientist role at Sporty. My background in machine learning, conversational AI, and production evaluation systems aligns well with your focus on improving LLM response quality, retrieval performance, and user satisfaction.

I bring strong experience designing prompt experiments, analyzing interaction data, defining measurable quality metrics, and partnering with engineering teams to deploy validated improvements. I am particularly motivated by the opportunity to apply responsible AI practices and rigorous benchmarking to scalable AI-powered products.

I would welcome the opportunity to contribute an analytical, ownership-driven approach to Sporty’s remote-first team.

Sample interview questions
How would you build an evaluation framework for a production conversational AI system?

I would begin by defining the target user outcomes and a baseline across quality dimensions such as factual accuracy, relevance, task completion, safety, latency, and user satisfaction. I would create a representative evaluation dataset from real interactions, combine automated metrics with expert review, run controlled experiments, and monitor production results after deployment.

How would you diagnose and improve poor retrieval quality in a RAG pipeline?

I would inspect failed queries by intent, domain, language, and retrieval score to determine whether the issue is document coverage, chunking, embedding quality, metadata filtering, ranking, or answer-generation behavior. I would then test targeted changes such as revised chunking, hybrid retrieval, reranking, query rewriting, and better grounding prompts against a fixed benchmark before releasing them incrementally.

Describe how you would evaluate a new prompt-engineering strategy.

I would state a clear hypothesis, define primary and guardrail metrics, select a representative sample, and compare the candidate prompt with the baseline through offline evaluation and, where appropriate, an online A/B test. I would assess quality gains alongside latency, token cost, safety, and consistency, then document results and deploy only if the improvement is statistically and operationally meaningful.

Which metrics would you use to assess LLM quality and safety?

I would use a combination of automated checks, rubric-based human evaluation, adversarial test cases, and production feedback. Important measures would include factual grounding, instruction adherence, harmful-content rates, hallucination frequency, escalation behavior, and performance across different user segments and edge cases.

How would you recommend an LLM for a specific product use case?

I would first establish the business objective and constraints, including quality requirements, latency, cost, data privacy, deployment environment, and governance needs. I would evaluate shortlisted models on a domain-specific benchmark, review error patterns rather than relying solely on aggregate scores, and recommend the model that offers the strongest overall trade-off for the use case.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

About the role

We are looking for a Senior AI Scientist to continuously improve the quality, accuracy and effectiveness of our AI-powered products. You will analyse AI interactions, identify improvement opportunities and optimise conversational behaviour, prompting strategies and retrieval mechanisms. Working closely with AI Engineers, you will drive measurable improvements in AI performance and user satisfaction.

What you’ll be doing

  • Analyse AI conversations and identify opportunities to improve response quality and user experience.
  • Design, implement and evaluate prompt engineering strategies.
  • Optimise RAG pipelines and retrieval quality.
  • Evaluate different LLMs and recommend the most effective models for specific use cases. Potentially train customer personal models.
  • Define and monitor AI quality metrics, including accuracy, relevance, latency and user satisfaction.
  • Design and execute experiments to validate AI improvements.
  • Analyse user feedback and interaction data to continuously improve AI behaviour.
  • Collaborate with AI Engineers to deploy validated improvements into production.
  • Maintain evaluation frameworks, benchmarks and documentation.
  • Stay current with advances in generative AI and recommend practical improvements to existing products.

What you’ll bring

  • 5+ years of experience in AI, ML, NLP or a related field.
  • Strong understanding of LLMs and conversational AI.
  • Experience with prompt engineering, prompt evaluation and LLM optimisation.
  • Experience evaluating conversational AI systems using qualitative and quantitative methods.
  • Strong Python and SQL skills.
  • Strong analytical and problem-solving skills.
  • Excellent communication and stakeholder management skills.
  • Proactive mindset with strong ownership and accountability.
  • Experience with speech and voice AI technologies.
  • Experience with AI evaluation frameworks and benchmarking.
  • Experience with vector databases and embedding models.
  • Knowledge of AI governance, safety and responsible AI practices.
  • Experience working with production AI systems at scale.

What’s in it for you

  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities

Interview Process

  • Remote video screening with our Talent Acquisition Team 
  • Online home assignment
  • Remote video interview with Team Members (3×45 Mins)

If you’re interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.
Application method

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
or continue without an account
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent Salaries
Menu