[All remote jobs](https://jobicy.com/jobs.md)Open role[![Sonatype logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/12/WRILS-201223232219-694679.jpg)](https://jobicy.com/company/sonatype.md)Remote opportunity at[Sonatype](https://jobicy.com/company/sonatype.md)

# Data Scientist

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/sonatype.md)Share30 Aug 2026Published18Listing views2Application actions29 Sep 2026Apply before  Opportunity details

## About this role.

AI SummarySonatype is seeking a senior Data Scientist to lead applied machine-learning and generative-AI initiatives across product, engineering, security, and research. The role covers the full lifecycle from ambiguous problem definition and experimentation through model evaluation, deployment enablement, and production reliability. Core work includes anomaly, malicious-behavior, and fraud detection, as well as LLM, RAG, embedding, and agentic-workflow applications. The candidate will act as an internal AI consultant, translating technical tradeoffs for varied stakeholders and helping improve organizational AI capability. Strong Python, ML engineering, LLM application design, evaluation, and collaborative software-development practices are required.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

4/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

5/5IndependentCollaborative

AI insightThis is a senior, cross-functional technical-lead role requiring 5+ years of applied AI experience and the ability to deliver reliable ML and GenAI systems across several problem domains. Success requires independent judgment in ambiguous settings, rigorous evaluation, and production-minded engineering in a security-sensitive environment.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate$175,000US market range$145k–$210k0$231k

AI insightNo actual salary range is disclosed in the posting. This is an estimated US-market yearly base-salary range for a senior Data Scientist / applied AI technical lead with ML, LLM, MLOps, and security-domain responsibilities; actual compensation may vary by location, level, and total-rewards structure.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Python](https://jobicy.com/jobs?search_keywords=Python.md)[Machine Learning](https://jobicy.com/jobs?search_keywords=Machine%20Learning.md)[Generative AI](https://jobicy.com/jobs?search_keywords=Generative%20AI.md)[Large Language Models](https://jobicy.com/jobs?search_keywords=Large%20Language%20Models.md)[Retrieval-Augmented Generation](https://jobicy.com/jobs?search_keywords=Retrieval-Augmented%20Generation.md)[Anomaly Detection](https://jobicy.com/jobs?search_keywords=Anomaly%20Detection.md)[MLOps](https://jobicy.com/jobs?search_keywords=MLOps.md)[Model Evaluation](https://jobicy.com/jobs?search_keywords=Model%20Evaluation.md)[Databricks](https://jobicy.com/jobs?search_keywords=Databricks.md)[Cybersecurity](https://jobicy.com/jobs?search_keywords=Cybersecurity.md)

Sample interview questionsDescribe an ML or GenAI application you took from an early concept to a production workflow. How did you measure success?I would begin by defining the user decision or workflow to improve and establishing measurable success criteria, such as precision at a review threshold, task-completion quality, latency, and adoption. I would build a small prototype, evaluate it against representative labeled cases and failure scenarios, then productionize it with versioning, monitoring, and a feedback loop. After launch, I would compare the system against the baseline and iterate based on both quantitative performance and user feedback.

How would you approach malicious-behavior or anomaly detection when labeled examples are limited?

I would combine domain-driven features with unsupervised or semi-supervised approaches such as isolation forests, clustering, autoencoders, or one-class methods. I would work with security experts to create high-value review queues, use analyst feedback to develop labels over time, and evaluate precision at operationally useful alert volumes. I would also monitor feature and data drift because adversarial behavior evolves.

How do you evaluate whether an LLM or RAG application is reliable enough to deploy?

I would define a task-specific evaluation set that includes normal cases, edge cases, adversarial prompts, and known failure patterns. Metrics would cover groundedness, factual accuracy, retrieval quality, structured-output validity, latency, cost, and safety behavior. I would use automated evaluations alongside expert review, establish release gates, and instrument production tracing and feedback so performance can be monitored after deployment.

How do you communicate AI tradeoffs to non-technical stakeholders?

I frame discussions around the business decision, expected benefit, risk, cost, and confidence level rather than model internals alone. I use concrete examples and clearly distinguish what the system can automate, what needs human review, and how outcomes will be measured. This helps stakeholders make informed prioritization decisions while maintaining realistic expectations.

What practices would you use to make an AI system maintainable and compliant when working with customer data?

I would apply data minimization, access controls, retention rules, audit logging, and clear lineage for datasets, prompts, models, and outputs. I would partner early with governance and security teams to identify privacy constraints, establish approved data flows, and define human-review or escalation paths for sensitive outcomes. Reproducible pipelines, model versioning, testing, monitoring, and incident-response procedures would support safe long-term operation.

Sonatype is the software supply chain management company that invented componentized software development and pioneered the software supply chain category. As leaders in the open-source community and the DevSecOps industry, we run the world’s largest repository of Java open-source components—Maven Central.

Our groundbreaking, full-spectrum platform empowers customers to rapidly create, deploy, and maintain innovative software at scale, all while aligning directly to their business needs. Trusted by more than 2,000 organizations—including 70% of the Fortune 100—and over 15 million software developers, Sonatype’s tools and guidance help deliver exceptional, secure software.

From inventing modern artifact management with Nexus Repository to introducing the world’s only solution that halts malicious open-source malware in its tracks, we’re committed to constant innovation. We leverage AI/ML to give our clients, developers, and the industry complete confidence in the quality, automation, and security of their software.

Learn more at[www.sonatype.com](http://www.sonatype.com)

### The Role

We’re looking for a Data Scientist to join our growing AI & Data Science team. You’ll operate as an internal AI consultant and technical lead, helping multiple teams across Sonatype apply machine learning and generative AI to real-world problems — from malicious-behavior and anomaly detection in our security data, to developer- and analyst-facing GenAI experiences.

You’ll explore complex datasets, design experiments, build and validate models, and collaborate closely with product, engineering, and security experts to turn research ideas into practical, scalable solutions. We have a mature data engineering team, so you can focus on doing what you do best — building and shipping models.

This role is ideal for someone who thrives on autonomy, loves translating ambiguous ideas into working systems, and enjoys working across boundaries rather than staying in a single product lane.

### What you’ll do:

*

Lead applied AI projects from concept to impact — prototype, validate, and help teams deploy practical ML and GenAI solutions.

*

Act as an internal consultant across product, engineering, security, and research teams: scope problems, evaluate approaches, and advise on ML/AI best practices and productive use of generative technologies.

*

Lead the research, development, and deployment of models for use cases such as malicious behavior detection, anomaly detection, and fraud analysis — using techniques ranging from classical ML to LLMs, embeddings, retrieval-augmented generation, and agentic workflows.

*

Design robust experiments and establish evaluation pipelines for model reliability, accuracy, and business impact (cross-validation, drift monitoring, ground-truth evaluation).

*

Bridge research and production: translate research insights into scalable APIs, tools, or workflows that enable other teams to adopt AI effectively.

*

Explore new techniques (LLMs, embeddings models, RAG, agentic workflows) to enhance developer and security experiences.

*

Communicate technical concepts, tradeoffs, and recommendations clearly to both technical and non-technical stakeholders through presentations, documentation, and collaboration; mentor peers and help elevate the organization’s AI literacy and capabilities.

*

Partner with our data governance team to ensure compliance with data-privacy regulations and ethical considerations when working with customer data.

### What you bring:

*

5+ years of hands-on experience in applied data science, machine learning, AI engineering, or AI research.

*

Computer Science or equivalent technical degree strongly preferred

*

Strong Python skills and practical experience with data and AI libraries/platforms such as Databricks, and LLM APIs, scikit-learn

*

Experience building and shipping ML or GenAI applications—from early prototype through usable internal or customer-facing workflows.

*

Deep familiarity with modern LLM ecosystems, including OpenAI, Anthropic/Claude, Hugging Face, and open-weight models.

*

Ability to select models and design effective LLM applications using prompting, context management, structured outputs, retrieval, and tool use.

*

Experience building agentic or multi-step AI workflows with LangGraph, LangChain, Semantic Kernel, or similar orchestration frameworks.

*

Strong evaluation mindset: defining useful quality metrics, building representative evaluation datasets, assessing reliability, and making data-driven tradeoffs.

*

Comfortable working with large, messy, structured, and unstructured data to produce features, insights, and clear visualizations.

*

Proficiency with Git, testing, code review, and collaborative software-development practices.

*

Practical, balanced judgment: comfortable exploring emerging AI capabilities while building maintainable, secure, dependable systems.

*

Proactive and accountable, with strong written and verbal communication skills across technical and non-technical partners.

### It’d be great if you had:

*

Strong MLOps experience, including MLflow or comparable tooling, experiment tracking, reproducible pipelines, model/application versioning, CI/CD, serving, and production monitoring.

*

Experience operating ML or GenAI systems at scale, including observability, tracing, incident response, and data or model-drift detection.

*

Experience with Databricks ML, AWS SageMaker, Azure ML, or similar managed ML platforms.

*

Familiarity with MCP, agent-tool integrations, LLM guardrails, and production safety practices.

*

Experience with AI-assisted development tools such as Copilot, Claude Code, or Codex.

*

Exposure to cybersecurity, fraud detection, anomaly detection, code analysis, or software supply-chain security.

*

Experience with PySpark and production data pipelines.

*

Experience working within a software product company or SaaS.

### Things we’re proud of:

* 2026 Gartner® Magic Quadrant™ Leader for Software Supply Chain Security
* 2026 Celebrating 15 Years of Sonatype Research Labs – Industry-leading software supply chain and open source security research
* 2026 Founding Member of the Linux Foundation Initiative for Open Source Sustainability
* 2026 State of the Software Supply Chain® Report – Continuing industry leadership in software supply chain security and AI security research
* 2025 Visionary in Gartner® Magic Quadrant™ for Application Security Testing!
* 2025 AI Compliance Solution of the Year – AI Breakthrough Awards
* 2025 DEVIES Award to our SBOM Manager for a new product for its innovation and impact in developer technology
* 2024 Industry Leader in Forrester-Wave for Software Composition Analysis (2024 Q4 report)
* Constellation AST Shortlist: Sonatype has been listed on the Constellation ShortList™ for Application Security Testing for 2024
* Data Breakthrough Awards: Sonatype was announced as a 2024 winner in the “Open Source Data Solution of the Year.”
* SD Times: Best in Show Security
* Fast Company Best Workplaces for Innovators 2024
* The Herd Top 100 Private Software Companies 2024
* Diversity & Inclusion Working Groups
* Parental Leave Policy
* Paid Volunteer Time Off (VTO)

### Compensation

### Additional Information

At Sonatype, we value diversity and inclusivity. We offer perks such as parental leave, diversity and inclusion working groups, and flexible working practices to allow our employees to show up as their whole selves. We are an equal-opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. If you have a disability or special need that requires accommodation, please do not hesitate to let us know.

Show more

[Apply now >](https://jobicy.com/jobs/152101-data-scientist-2.md)

>  Annual salary information is not provided for this position. Explore salary ranges for similar roles in our [Salary Directory ›](https://jobicy.com/salaries.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

Matched by job category10 related opportunities[Data Science & Analytics](https://jobicy.com/categories/data-science.md) [Browse all jobs](https://jobicy.com/jobs.md)
*
![Instacart logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2023/03/fe39c2bc9bb5ebb0b5e24318b1f3b60d.jpeg)
Instacart  Aug 30

### [Ads AI Analytics Lead II](https://jobicy.com/jobs/152090-ads-ai-analytics-lead-ii.md)

We’re transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more…

*
![Quantum Metric logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/52363ead-221.jpg)
Quantum Metric  Aug 30

### [Senior/Strategic Digital Analytics Consultant – Spain](https://jobicy.com/jobs/152083-senior-strategic-digital-analytics-consultant-spain.md)

😎 Our Culture Quantum Metric’s number one objective is happy people, diverse and inclusive culture. We’re passionate about empowering our people to become the best version of themselves, offering coaching…

*
![Customer.io logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/b82146da-221.jpg)
Customer.io  Aug 30

### [Senior Data Scientist](https://jobicy.com/jobs/144856-senior-data-scientist-4.md)

About Customer.ioOver 8,000 companies — from scrappy startups to global brands — use our platform to send billions of emails, push notifications, in-app messages, and SMS every day. Customer.io powers…

*
![Pindrop logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/497cbe3e-221.jpeg)
Pindrop  Aug 30

### [Research Scientist (Fraud)](https://jobicy.com/jobs/147988-research-scientist-ii.md)

Who We Are Pindrop is the Real Human + Right Human® Identity Trust Platform for the AI era. As AI-driven fraud and deepfakes erode trust in digital communication, Pindrop delivers…

*
![Zartis logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/9bd0869f-221.webp)
Zartis  Aug 29

### [Lead Data Architect](https://jobicy.com/jobs/152045-lead-data-architect.md)

The company and our mission: Zartis is a global AI transformation and technology consulting partner where talented engineers and technologists work on cutting edge innovation. We partner with ambitious organizations to…

*
![Quora logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/09/WRILS-200916172339-629302.jpg)
Quora  Aug 29

### [Data Scientist – Quora](https://jobicy.com/jobs/151982-data-scientist-quora.md)

[Quora is a privately held, “remote-first” company. This position can be performed remotely from multiple countries around the world. Please visit careers.quora.com/eligible-countries for details regarding employment eligibility by country.] About…

*
![Liberty Mutual logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/08/a0f68d769896c0f098621799a6396382.jpeg)
Liberty Mutual  Aug 29

### [Director I, Data Science, Enterprise Data & Data Science](https://jobicy.com/jobs/148333-director-i-data-science-enterprise-data-data-science.md)

Description We’re seeking an exceptional, hands-on Data Scientist with deep expertise in data science, MLOps, and building GenAI solutions to join our Enterprise Data & Data Science team. In this…

*
![Meta logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/750f86a9-221.jpeg)
Meta  Aug 29

### [Data Scientist, Analytics (Technical Leadership)](https://jobicy.com/jobs/141971-data-scientist-analytics-technical-leadership.md)

We are seeking experienced Data Scientists to join our team and drive impact across various product areas. As a Data Scientist, you will collaborate with cross-functional partners to identify and…

*
![H1 logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/8f84b1ed-221.jpg)
H1  Aug 29

### [Principal Analyst, Data Integration](https://jobicy.com/jobs/147852-principal-analyst-data-integration.md)

At H1, we believe access to the best healthcare information is a basic human right. Our mission is to provide a platform that can optimally inform every doctor interaction globally….

*
![RevenueCat logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/09/e9b2999e3e7950f153097909d027dda2.jpg)
RevenueCat  Aug 28

### [Senior Data Scientist, Fraud & Risk](https://jobicy.com/jobs/151942-senior-data-scientist-fraud-risk.md)

RevenueCat removes the headaches of building and scaling in‑app subscriptions. Since graduating from YC’s S18 batch we’ve grown into the default monetization platform for mobile: we’re in >40% of newly…