About this role.
Sporty is seeking a Senior AI Scientist to improve the quality, accuracy, safety, and effectiveness of AI-powered conversational products. The role centers on analyzing AI interactions, designing prompt strategies, optimizing RAG retrieval, evaluating LLMs, and measuring outcomes through structured experiments and quality metrics. The successful candidate will work closely with AI Engineers to move validated improvements into production while maintaining benchmarks, evaluation frameworks, and documentation. This remote-first position requires at least five years of relevant AI, ML, NLP, or conversational AI experience, with strong Python, SQL, stakeholder communication, and ownership skills.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
4/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Team,
I am excited to apply for the Senior AI Scientist role at Sporty. My background in machine learning, conversational AI, and production evaluation systems aligns well with your focus on improving LLM response quality, retrieval performance, and user satisfaction.
I bring strong experience designing prompt experiments, analyzing interaction data, defining measurable quality metrics, and partnering with engineering teams to deploy validated improvements. I am particularly motivated by the opportunity to apply responsible AI practices and rigorous benchmarking to scalable AI-powered products.
I would welcome the opportunity to contribute an analytical, ownership-driven approach to Sporty’s remote-first team.
Sample interview questions
I would begin by defining the target user outcomes and a baseline across quality dimensions such as factual accuracy, relevance, task completion, safety, latency, and user satisfaction. I would create a representative evaluation dataset from real interactions, combine automated metrics with expert review, run controlled experiments, and monitor production results after deployment.
I would inspect failed queries by intent, domain, language, and retrieval score to determine whether the issue is document coverage, chunking, embedding quality, metadata filtering, ranking, or answer-generation behavior. I would then test targeted changes such as revised chunking, hybrid retrieval, reranking, query rewriting, and better grounding prompts against a fixed benchmark before releasing them incrementally.
I would state a clear hypothesis, define primary and guardrail metrics, select a representative sample, and compare the candidate prompt with the baseline through offline evaluation and, where appropriate, an online A/B test. I would assess quality gains alongside latency, token cost, safety, and consistency, then document results and deploy only if the improvement is statistically and operationally meaningful.
I would use a combination of automated checks, rubric-based human evaluation, adversarial test cases, and production feedback. Important measures would include factual grounding, instruction adherence, harmful-content rates, hallucination frequency, escalation behavior, and performance across different user segments and edge cases.
I would first establish the business objective and constraints, including quality requirements, latency, cost, data privacy, deployment environment, and governance needs. I would evaluate shortlisted models on a domain-specific benchmark, review error patterns rather than relying solely on aggregate scores, and recommend the model that offers the strongest overall trade-off for the use case.
About the role
We are looking for a Senior AI Scientist to continuously improve the quality, accuracy and effectiveness of our AI-powered products. You will analyse AI interactions, identify improvement opportunities and optimise conversational behaviour, prompting strategies and retrieval mechanisms. Working closely with AI Engineers, you will drive measurable improvements in AI performance and user satisfaction.
What you’ll be doing
- Analyse AI conversations and identify opportunities to improve response quality and user experience.
- Design, implement and evaluate prompt engineering strategies.
- Optimise RAG pipelines and retrieval quality.
- Evaluate different LLMs and recommend the most effective models for specific use cases. Potentially train customer personal models.
- Define and monitor AI quality metrics, including accuracy, relevance, latency and user satisfaction.
- Design and execute experiments to validate AI improvements.
- Analyse user feedback and interaction data to continuously improve AI behaviour.
- Collaborate with AI Engineers to deploy validated improvements into production.
- Maintain evaluation frameworks, benchmarks and documentation.
- Stay current with advances in generative AI and recommend practical improvements to existing products.
What you’ll bring
- 5+ years of experience in AI, ML, NLP or a related field.
- Strong understanding of LLMs and conversational AI.
- Experience with prompt engineering, prompt evaluation and LLM optimisation.
- Experience evaluating conversational AI systems using qualitative and quantitative methods.
- Strong Python and SQL skills.
- Strong analytical and problem-solving skills.
- Excellent communication and stakeholder management skills.
- Proactive mindset with strong ownership and accountability.
- Experience with speech and voice AI technologies.
- Experience with AI evaluation frameworks and benchmarking.
- Experience with vector databases and embedding models.
- Knowledge of AI governance, safety and responsible AI practices.
- Experience working with production AI systems at scale.
What’s in it for you
- Sporty is a remote first company in pursuit of sustainability
- A competitive salary + individual performance based bonuses every quarter
- 28 days paid annual leave
- Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
- Referral bonuses & flash bonuses
- Top of the line equipment
- Annual company retreats to provide great internal networking opportunities
Interview Process
- Remote video screening with our Talent Acquisition Team
- Online home assignment
- Remote video interview with Team Members (3×45 Mins)
If you’re interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.






