[![Image]() Meet Jobicy Copilot — free AI autofill for job applications + remote job alerts ›](#)   [All remote jobs](https://jobicy.com/jobs.md)Open role[![Reka logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/4c8a0366-221.png)](https://jobicy.com/company/reka.md)Remote opportunity at[Reka](https://jobicy.com/company/reka.md)

# Member of Technical Staff (Data Intelligence)

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/reka.md)Share14 Sep 2026Published31Listing views1Application actions14 Oct 2026Apply before  Opportunity details

## About this role.

AI SummaryThis Member of Technical Staff role focuses on data intelligence for large-scale multimodal foundation-model training. The hire will define data-quality standards with researchers, curate and assess datasets, and build automated methods for quality evaluation, domain mixtures, and synthetic-to-real adaptation. They will also own reproducibility, metadata and provenance tracking, CI/CD, and operational tooling for petabyte-scale data and compute pipelines. The position combines applied ML research with production data infrastructure, distributed processing, and cross-functional collaboration in a remote-first startup environment.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

4/5IndependentCollaborative

AI insightThe role requires deep ML knowledge alongside hands-on ownership of reliable, petabyte-scale data systems and experimental methodology. Success depends on independently translating research needs into scalable tooling while balancing quality, compute, storage, reproducibility, and delivery speed.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate$220,000US market range$180k–$280k0$308k

AI insightNo actual salary was disclosed, so these are estimated US-market annual base-salary figures in USD for a senior-level ML/data infrastructure Member of Technical Staff role at a foundation-model startup. Equity, bonuses, and location-specific compensation could materially change total compensation.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Machine Learning](https://jobicy.com/jobs?search_keywords=Machine%20Learning.md)[Deep Learning](https://jobicy.com/jobs?search_keywords=Deep%20Learning.md)[Data Engineering](https://jobicy.com/jobs?search_keywords=Data%20Engineering.md)[Data Quality](https://jobicy.com/jobs?search_keywords=Data%20Quality.md)[Distributed Processing](https://jobicy.com/jobs?search_keywords=Distributed%20Processing.md)[Python](https://jobicy.com/jobs?search_keywords=Python.md)[PyTorch](https://jobicy.com/jobs?search_keywords=PyTorch.md)[Spark](https://jobicy.com/jobs?search_keywords=Spark.md)[Ray](https://jobicy.com/jobs?search_keywords=Ray.md)[Airflow](https://jobicy.com/jobs?search_keywords=Airflow.md)

Sample interview questionsHow would you define and operationalize data quality for a new foundation-model training dataset?I would begin by aligning quality criteria to intended model behaviors and downstream evaluations. I would implement measurable checks for integrity, duplication, relevance, distribution coverage, safety, labeling quality, and modality-specific properties, then set acceptance thresholds and monitoring dashboards. Finally, I would validate that the metrics correlate with controlled training outcomes rather than treating proxy metrics as sufficient evidence.

Describe how you would make a petabyte-scale data pipeline reproducible.

I would version immutable dataset manifests, transformation code, container environments, configurations, and model-training inputs. Each pipeline run would record lineage from raw sources through filtering and sharding to final artifacts, including checksums, metadata, and execution parameters. Orchestration would produce auditable run records so any training or evaluation dataset can be reconstructed or precisely identified.

How would you evaluate whether a synthetic-to-real data adaptation method improves model performance?

I would formulate a hypothesis and create controlled experiments that hold model architecture, compute budget, and most training conditions constant. I would compare well-defined mixture ratios against real-only and synthetic-only baselines across representative, held-out evaluations. I would report uncertainty, investigate subgroup regressions, and use ablations to determine whether the observed gain comes from the adaptation method rather than confounding factors.

What approaches would you use to optimize throughput and cost in a distributed data-processing workflow?

I would first instrument the pipeline to locate bottlenecks in I/O, serialization, skew, compute utilization, storage layout, and scheduling. I would then improve partitioning and sharding, batch operations, cache reusable artifacts, reduce unnecessary data movement, and select resource configurations based on observed workload characteristics. Changes would be validated with benchmarks that measure both end-to-end throughput and cost per usable training example.

How would you work with model researchers when their requested dataset is expensive or difficult to produce at scale?

I would clarify the research objective and identify the smallest experiment that can test the underlying assumption. I would propose staged prototypes, estimate quality and infrastructure tradeoffs, and communicate the expected cost, latency, and risks of each option. After selecting an approach, I would deliver measurable milestones and use experiment results to decide whether further scaling is justified.

In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get.

### What you’ll do

*

Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds

*

Explore open source datasets and create internal ones most suitable to build fundamental World Models

*

Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data.

*

Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs

*

Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction

*

Track and optimize throughput, storage, and compute utilization across pipelines and related assets

### What we’re looking for

*

Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems

*

Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems

*

Demonstrated research experience with data compositions, quality, and dataset releases

*

Ability to design and execute experiments with convincing unbiased outcomes

*

Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)

*

Solid Python skills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)

*

Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale

*

Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers

*

Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments

### Reka’s Mission

Reka’s mission is to build useful multimodal artificial intelligence and use it to empower organisations and businesses. We are a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world. Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI over the past decade.

###
Why Reka?

*

An Elite Team: Collaborate with top-tier engineers, researchers, operators from renowned organizations like Google DeepMind and Facebook AI Research (FAIR) and successful startups, driving innovation in cutting-edge AI technology.

*

Massive Market Opportunity: Be part of a rapidly growing industry poised to transform multiple sectors globally, offering the chance to make a significant impact.

*

Mission-Driven Environment: Work alongside a collaborative, mission-focused team dedicated to advancing AI for meaningful applications.

*

Inclusive and Open Culture: Thrive in an open and inclusive work environment that values diverse perspectives and fosters creativity.

*

Generous Benefits: Enjoy 5 weeks of paid leave to recharge, comprehensive healthcare benefits including vision and dental, and additional perks that support your well-being.

*

Visa Support: We provide visa assistance, including H1B and OPT transfers, for US employees to ensure a smooth transition and support your career with us.

Show more

[Apply now >](https://jobicy.com/jobs/153245-member-of-technical-staff-data-intelligence.md)

>  Annual salary information is not provided for this position. Explore salary ranges for similar roles in our [Salary Directory ›](https://jobicy.com/salaries.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

Matched by job category10 related opportunities[Data Science & Analytics](https://jobicy.com/categories/data-science.md) [Browse all jobs](https://jobicy.com/jobs.md)
*
![UpGuard logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/474a12c7-221.png)
UpGuard  Sep 14

### [Product Data Science Lead](https://jobicy.com/jobs/153276-product-data-science-lead.md)

Who are we? At UpGuard, we are replacing manual security bottlenecks with AI-driven precision. Fresh off a US$75M Series C, we are scaling our infrastructure to process 100 billion risk…

*
![Affirm logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/fac7714c-221-1.jpg)
Affirm  Sep 14

### [Analytics Engineer II](https://jobicy.com/jobs/153263-analytics-engineer-ii.md)

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters…

*
![Affirm logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/fac7714c-221-1.jpg)
Affirm  Sep 14

### [Analyst II, Full Stack (Revenue Analytics)](https://jobicy.com/jobs/153258-analyst-ii-full-stack-revenue-analytics.md)

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters…

*
![Affirm logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/fac7714c-221-1.jpg)
Affirm  Sep 14

### [Analytics Lead, Full Stack](https://jobicy.com/jobs/153266-analytics-lead-full-stack.md)

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters…

*
![Playson logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/6f266d9c-221-1.png)
Playson  Sep 14

### [Senior Game Mathematician](https://jobicy.com/jobs/153229-senior-game-mathematician.md)

About the Role Playson is a leading online gaming supplier with worldwide recognition which was founded in 2012. We offer complete gaming solutions based on the latest technologies and detailed…

*
![Nebius logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/d90c0566-221.webp)
Nebius  Sep 14

### [Senior ML Engineer (AI Research, Physical AI)](https://jobicy.com/jobs/150592-senior-ml-engineer-ai-research-physical-ai.md)

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from…

*
![Ruby Labs logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/0799a7b7-221-1.jpeg)
Ruby Labs  Sep 14

### [Data Engineer (Marketing Systems)](https://jobicy.com/jobs/150584-data-engineer-marketing-systems.md)

About usRuby Labs is a leading tech company that creates and operates innovative consumer products. We offer a diverse range of opportunities across the health, education, and entertainment industries. Our…

*
![Invisible Technologies logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/5d284c35-221.jpeg)
Invisible Technologies  Sep 14

### [Video Collection (LATAM) – Freelance AI Trainer Project](https://jobicy.com/jobs/149160-video-collection-latam-freelance-ai-trainer-project.md)

We’re looking for experts in LATAM to record short, first-person videos of everyday manual tasks using a head-mounted smartphone. Tasks include household activities such as cleaning, organizing, and laundry, as…

*
![Dutchie logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/02/305a6ae1871314b06b00082374e7ba74.jpeg)
Dutchie  Sep 13

### [Brands Analyst](https://jobicy.com/jobs/153178-brands-analyst.md)

About Dutchie Founded in 2017, Dutchie is a comprehensive technology platform powering dispensary operations, while providing consumers with safe and easy access to cannabis. Dutchie aims to further support the…

*
![Nebius logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/d90c0566-221.webp)
Nebius  Sep 13

### [Senior Research Scientist (Architectures Research)](https://jobicy.com/jobs/150577-senior-research-scientist-architectures-research.md)

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from…