All remote jobs
Open role
Remote opportunity atSynthesia

Research Engineer in Data

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
29Listing views
3Application actions
10 Oct 2026Apply before
Opportunity details

About this role.

AI Summary

Synthesia is seeking a Research Engineer in Data to build and improve the data layer supporting generative AI research and model training. The role combines applied machine learning, data engineering, and ML infrastructure, with a focus on sourcing, curating, annotating, processing, and evaluating large-scale video and audio datasets. The engineer will collaborate closely with model-training teams to translate research needs into higher-quality datasets and features that improve model performance. Strong Python engineering practices are essential, while experience with workflow orchestration and large-scale processing systems is advantageous. The position is full time and supports hybrid work from select European hubs or remote work within Europe.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

5/5
RelaxedFast-paced

Autonomy Level

4/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThis is a technically demanding role involving data-centric ML for generative video and audio at million-hour scale. Success requires strong judgment across data quality, scalable pipeline design, experimentation, and collaboration with research and model-training stakeholders.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$170,000
US market range$145k–$200k
AI insightNo numerical salary was disclosed, so these are estimated US-market annual base-salary figures in USD for a senior-level research/data engineer working on large-scale generative-AI data systems. Actual compensation may vary materially by country, level, cash-versus-equity mix, and location; the posting also mentions stock options and a bonus without amounts.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
Describe a project where improving data quality produced a measurable improvement in model performance.

I would explain the baseline model metric, identify the data issue through slice analysis or error review, and describe the intervention, such as relabeling, deduplication, rebalancing, or filtering. I would then quantify the improvement using a controlled evaluation and explain how I operationalized the quality checks to prevent regression.

How would you design a pipeline for processing and curating millions of hours of video and audio data?

I would use modular, versioned stages for ingestion, validation, metadata extraction, deduplication, feature generation, annotation, and publishing. The system should use distributed compute, idempotent jobs, data lineage, observability, and quality gates, with reproducible dataset versions that model teams can reliably consume.

How do you decide which new annotations or features are worth adding to a training dataset?

I would start with model-team hypotheses and error slices, then assess whether a proposed annotation is predictive, reliable to generate, and actionable in training or evaluation. I would run a small pilot, measure annotation agreement and model lift against a baseline, and only scale features that show meaningful value relative to cost and operational complexity.

What practices do you use to ensure Python data-processing code remains maintainable and trustworthy?

I use clear interfaces, typed and documented code where appropriate, unit and integration tests, deterministic transformations, structured logging, and code review. For production pipelines, I also add data contracts, schema validation, monitoring for volume and distribution shifts, and reproducible environments.

How would you collaborate with researchers when their data requirements are incomplete or change frequently?

I would turn broad requests into a short written specification covering the target model behavior, required slices, dataset definitions, quality thresholds, and evaluation metrics. I would deliver an early sample or prototype, review findings together, and iterate through versioned datasets so changes remain traceable and do not disrupt existing experiments.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US.

As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations.

Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia’s VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow.

About the role

The Data team manages the complete lifecycle of data for researchers – from sourcing and large-scale processing to delivering datasets that power our models. Data sits at the heart of our Research efforts and enables all other teams. As part of the Data team, you’ll work with over a million hours of video and audio data.

This role exists at the intersection of applied research, data engineering, and ML infrastructure rather than being a traditional research position.

You’ll build the world’s best human-centric data lake by collaborating closely with our model training teams. By understanding their requirements, you’ll extract new features and annotations that elevate our datasets. You should be passionate about enhancing model performance through high-quality, accurate datasets. Our infrastructure and pipelines are in great shape, and this role provides room to not only enhance them but also influence the team’s longer-term strategy.

What we’re looking for:

  • A strong background in data-centric, applied Machine Learning, with hands-on experience improving model performance through data quality, curation, labeling, and evaluation rather than model architecture alone

  • Experience working on the data layer of Generative AI products, particularly involving images, video, or audio

  • Excellent Python skills, with a strong focus on writing clean, maintainable, and well-tested code

  • Nice to have if you have experience designing, building, and operating workflow orchestration systems and large-scale data processing pipelines

Why join us?

We’re living the golden age of AI. The next decade will yield the next iconic companies, and we dare to say we have what it takes to become one. Here’s why,

Our culture

At Synthesia we’re passionate about building, not talking, planning or politicising. We strive to hire the smartest, kindest and most unrelenting people and let them do their best work without distractions. Our work principles serve as our charter for how we make decisions, give feedback and structure our work to empower everyone to go as fast as possible. You can find out more about these principles here.

Serving 50,000+ customers (and 50% of the Fortune 500)

We’re trusted by leading brands such as Heineken, Zoom, Xerox, McDonald’s and more. Read stories from happy customers and what 1,200+ people say on G2.

Proprietary AI technology

Since 2017, we’ve been pioneering advancements in Generative AI. Our AI technology is built in-house, by a team of world-class AI researchers and engineers. Learn more about our AI Research Lab and the team behind.

AI Safety, Ethics and Security

AI safety, ethics, and security are fundamental to our mission. While the full scope of Artificial Intelligence’s impact on our society is still unfolding, our position is clear: People first. Always. Learn more about our commitments to AI Ethics, Safety & Security.

The good stuff…

  • Competitive compensation (salary + stock options + bonus)

  • Hybrid work setting with an office in London, Amsterdam, Zurich, Munich, or remote in Europe.

  • 25 days of annual leave + public holidays

  • Great company culture with the option to join regular planning and socials at our hubs

  • + other benefits depending on your location

#LI-MD1

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu