[![Image]() Meet Jobicy Copilot — free AI autofill for job applications + remote job alerts ›](#)   [All remote jobs](https://jobicy.com/jobs.md)Open role[![Serve Robotics logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/8ad805d9-221.png)](https://jobicy.com/company/serve-robotics.md)Remote opportunity at[Serve Robotics](https://jobicy.com/company/serve-robotics.md)

# Lead Machine Learning Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/serve-robotics.md)Share23 Sep 2026Published36Listing views3Application actions23 Oct 2026Apply before  Opportunity details

## About this role.

AI SummaryServe Robotics is seeking a senior machine learning engineer to build and optimize distributed training systems for autonomy models used in sidewalk-delivery robots. The role centers on petabyte-scale multimodal robotics data, including video, point clouds, and other sensor inputs, processed across large GPU clusters. Key work includes eliminating training bottlenecks, improving data pipelines and GPU utilization, refining model architectures and loss functions, and operating reliable distributed jobs. The engineer will partner closely with ML researchers and infrastructure teams to productionize experiments and accelerate model iteration. Candidates need at least five years of production ML experience, strong Python skills, and deep familiarity with distributed training and large-scale datasets.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

4/5GuidedFull ownership

### Communication Load

4/5IndependentCollaborative

AI insightThis is a highly technical senior role combining distributed systems, deep learning optimization, large-scale data engineering, and robotics perception. Success requires independently diagnosing complex performance bottlenecks while collaborating effectively across research and infrastructure functions.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianHighly competitive$242,500US market range$190k–$280k0$308k

AI insightThe disclosed US yearly base-salary range is $225,000–$260,000 USD, with a midpoint of $242,500. This is competitive for a lead-level machine learning engineer specializing in distributed training, robotics, and multimodal autonomy; the estimated US market range is $190,000–$280,000 USD annually, varying by location, scope, and equity or bonus structure.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Machine Learning](https://jobicy.com/jobs?search_keywords=Machine%20Learning.md)[Distributed Training](https://jobicy.com/jobs?search_keywords=Distributed%20Training.md)[PyTorch](https://jobicy.com/jobs?search_keywords=PyTorch.md)[Python](https://jobicy.com/jobs?search_keywords=Python.md)[GPU Clusters](https://jobicy.com/jobs?search_keywords=GPU%20Clusters.md)[Multimodal Learning](https://jobicy.com/jobs?search_keywords=Multimodal%20Learning.md)[Computer Vision](https://jobicy.com/jobs?search_keywords=Computer%20Vision.md)[Robotics](https://jobicy.com/jobs?search_keywords=Robotics.md)[Data Pipelines](https://jobicy.com/jobs?search_keywords=Data%20Pipelines.md)[MLOps](https://jobicy.com/jobs?search_keywords=MLOps.md)

Sample interview questionsHow would you identify and prioritize bottlenecks in a distributed training pipeline?I would establish baseline metrics for GPU utilization, data-loader wait time, host CPU and memory usage, network throughput, communication overhead, and step-time breakdowns. I would then profile the highest-impact constraint, validate a targeted change with controlled experiments, and monitor both throughput and model-quality effects before rolling it out.

Describe an approach for training on multimodal robotics data such as video, LiDAR, and telemetry.

I would first define robust time synchronization, calibration, and quality checks across modalities. The pipeline would use scalable sharding and preprocessing, modality-specific encoders, and a fusion architecture appropriate to the downstream task, while tracking performance by modality availability and environmental scenario.

What techniques would you use to improve GPU utilization in multi-node training?

I would address input-pipeline parallelism, storage locality, batching strategy, mixed precision, gradient accumulation, communication overlap, and efficient distributed data parallel configuration. I would also inspect load imbalance between workers and ensure checkpointing and evaluation do not unnecessarily interrupt training.

How do you decide whether a model-performance issue is caused by data, optimization, or architecture?

I would use structured ablations: verify data integrity and label quality, compare training and validation curves, inspect representative failures, and hold the dataset constant while changing optimization or architecture choices. This makes it possible to isolate whether the issue is poor signal, unstable training dynamics, capacity limitations, or generalization.

How would you make research experiments reproducible and production-ready?

I would version datasets, code, configurations, and model artifacts; record environment and hardware details; and use experiment tracking with clear lineage. For production, I would add automated validation, fault-tolerant orchestration, checkpoint recovery, monitoring, and documented runbooks so that successful experiments can be repeated and scaled reliably.

At Serve Robotics, we’re reimagining how things move in cities. Our personable sidewalk robot is our vision for the future. It’s designed to take deliveries away from congested streets, make deliveries available to more people, and benefit local businesses.

The Serve fleet has been delighting merchants, customers, and pedestrians along the way in Los Angeles, Miami, Dallas, Atlanta and Chicago while doing commercial deliveries. We’re looking for talented individuals who will grow robotic deliveries from surprising novelty to efficient ubiquity.

### Who We Are

We are tech industry veterans in software, hardware, and design who are pooling our skills to build the future we want to live in. We are solving real-world problems leveraging robotics, machine learning and computer vision, among other disciplines, with a mindful eye towards the end-to-end user experience. Our team is agile, diverse, and driven. We believe that the best way to solve complicated dynamic problems is collaboratively and respectfully.

This role develops and scales large-scale machine learning training systems for multimodal robotics data, enabling the creation of high-performance autonomy models. By optimizing distributed training pipelines, neural network architectures, and data processing workflows, the position improves training efficiency, accelerates model iteration, and maximizes GPU utilization. The role collaborates closely with ML researchers and infrastructure teams, influencing the design, deployment, and performance of end-to-end autonomy models and the large-scale data pipelines that support them.

Responsibilities

*

Design and maintain training systems that can process and learn from petabyte-scale multimodal datasets (e.g., video and point cloud data). This includes ensuring data is efficiently loaded, distributed, and processed across large GPU clusters.

*

Identify and resolve bottlenecks in the training pipeline, including data loading, preprocessing, model computation, and inter-node communication, to maximize GPU utilization and reduce training time.

*

Work with the ML team to develop and refine neural network architectures suitable for autonomy tasks, particularly those handling high-dimensional and sequential sensor data.

*

Create and adjust loss functions and training strategies that help the model learn effectively from complex multimodal inputs and improve autonomy performance.

*

Configure, monitor, and maintain large-scale distributed training jobs across multiple machines and GPUs, ensuring stability, fault tolerance, and efficient resource usage.

*

Implement scalable systems to preprocess, transform, and augment large robotics datasets so that they are suitable for model training.

*

Work closely with ML scientists and other engineers to integrate new models, experiments, and training approaches into the production training pipeline.

*

Analyze training metrics, model outputs, and experiment logs to assess model performance and guide improvements in architecture, data usage, or training strategies.

*

Develop tools and workflows that allow teams to run experiments, track results, and iterate quickly on new model ideas or training approaches.

Qualifications

*

Master’s or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a closely related technical discipline.

*

Minimum of 5 years of professional experience developing, training, and deploying machine learning models in production environments.

*

Hands-on experience training machine learning models across multiple GPUs or compute nodes, including familiarity with distributed training frameworks and large dataset handling.

*

Strong programming skills in Python for implementing machine learning models, data pipelines, and training workflows.

*

Solid knowledge of core concepts such as neural networks, optimization algorithms, loss functions, model evaluation, and training methodologies.

What Makes You Stand out

*

Experience identifying and resolving training bottlenecks related to compute utilization, memory usage, and data throughput in machine learning systems.

*

Experience training machine learning models on robotics or autonomous driving datasets involving multimodal sensor inputs such as camera video, LiDAR point clouds, radar, or telemetry data.

*

Experience developing models that combine multiple data modalities (e.g., images, point clouds, and structured sensor data) into a unified learning system.

*

Peer-reviewed publications or significant research contributions in machine learning, robotics, or related areas.

*Please note: The listed base salary range applies to candidates based in the US. Compensation may vary depending on location, experience, and role alignment. We are open to qualified candidates working remotely in Canada

*

Canada – ALL: $177k – $215k CAD

Show more

[Apply now >](https://jobicy.com/jobs/153948-lead-machine-learning-engineer-3.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

Matched by job category10 related opportunities[Data Science & Analytics](https://jobicy.com/categories/data-science.md) [Browse all jobs](https://jobicy.com/jobs.md)
*
![Tenable logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/01879e02-221.jpg)
Tenable  Sep 23

### [Data Analyst – SQL / Databricks](https://jobicy.com/jobs/153978-data-analyst-sql-databricks.md)

Who is Tenable? Tenable® is the Exposure Management company. Over 40,000 organizations around the globe rely on Tenable to understand and reduce cyber risk. Our global employees support 65 percent…

*
![Quora logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/09/WRILS-200916172339-629302.jpg)
Quora  Sep 23

### [Senior Machine Learning Engineer, Ads – Quora](https://jobicy.com/jobs/153939-senior-machine-learning-engineer-ads-quora.md)

[Quora is a privately held, “remote-first” company. This position can be performed remotely from multiple countries around the world. Please visit careers.quora.com/eligible-countries for details regarding employment eligibility by country.] About…

*
![Quora logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/09/WRILS-200916172339-629302.jpg)
Quora  Sep 23

### [Staff Data Scientist – Quora](https://jobicy.com/jobs/153935-staff-data-scientist-quora.md)

[Quora is a privately held, “remote-first” company. This position can be performed remotely from multiple countries around the world. Please visit careers.quora.com/eligible-countries for details regarding employment eligibility by country.] About…

*
![Luxury Presence logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/05/f2b241fb769140fff2f70e8efaf5c0ff.jpg)
Luxury Presence  Sep 23

### [Senior Data Analyst, GTM Analytics – US](https://jobicy.com/jobs/153937-senior-data-analyst-gtm-analytics-us.md)

Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners and other top investors, we’re a Series C company that has hit $100M in…

*
![Clover Health logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/8d4e74b3-221.jpg)
Clover Health  Sep 23

### [Data Analyst, Clinical Data Effectiveness](https://jobicy.com/jobs/150923-data-analyst-clinical-data-effectiveness.md)

Counterpart Health is an AI‑powered physician enablement platform that delivers clinical insights to providers at the point of care. Our flagship product, Counterpart Assistant, is embedded into clinicians’ workflows and…

*
![Pleo logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/b27fbed2-221.jpg)
Pleo  Sep 22

### [Senior Data Engineer (Commercial Analytics)](https://jobicy.com/jobs/153896-senior-data-engineer-commercial-analytics.md)

About Pleo Messy spend management is tricky business. And tedious processes are a lose-lose situation for all involved, not just finance. At Pleo, we’re changing that. We build spend solutions…

*
![Socure logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/02/4dd877ff6ea919681d6e790a57ab639c.jpeg)
Socure  Sep 22

### [Senior Data Engineer](https://jobicy.com/jobs/153886-senior-data-engineer-4.md)

Why Socure? Socure is building the identity trust infrastructure for the digital economy — verifying 100% of good identities in real time and stopping fraud before it starts. The mission…

*
![Oddball logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/02/8c034287ffd7b6474f90645b1c72e60a.jpeg)
Oddball  Sep 22

### [Cloud Data Architect](https://jobicy.com/jobs/151390-cloud-data-architect.md)

Oddball believes that we can bring change and improve the daily lives of millions by bringing quality software to the federal space. Our team is full of experienced engineering, product,…

*
![Thermo Fisher Scientific logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/fbe52b8e-221.webp)
Thermo Fisher Scientific  Sep 22

### [Clinical Data Team Lead](https://jobicy.com/jobs/149471-clinical-data-team-lead.md)

Work ScheduleStandard (Mon-Fri)Environmental ConditionsOfficeJob DescriptionJoin Us as a Clinical Data Team Lead – Make an Impact at the Forefront of InnovationWe have successfully supported the top 50 pharmaceutical companies and…

*
![Ruby Labs logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/0799a7b7-221-1.jpeg)
Ruby Labs  Sep 22

### [Data Analytics Engineer](https://jobicy.com/jobs/151318-data-analytics-engineer.md)

About usRuby Labs is a leading tech company that creates and operates innovative consumer products. We offer a diverse range of opportunities across the health, education, and entertainment industries. Our…